Application programming interface for releasing a data structure
Patent Information
- Application Number
- DE102025101968
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS REFERENCE TO RELATED APPLICATIONSThis application claims priority to provisional U.S. Application No. 63 / 625,278 (Attorney Docket No. 0112912-A27PR0), entitled "APPLICATION PROGRAMMING INTERFACE TO MANAGEMENT RESOURCES," filed January 25, 2024, the entire contents of which are incorporated herein by reference. This application also includes for all purposes the complete disclosure of co-pending U.S. Patent Application No. 18 / 593,578, entitled "APPLICATION PROGRAMMING INTERFACE TO ALLOCATE A DATA STRUCTURE" (Attorney docket No. 0112912-A27US0), co-pending U.S. Patent Application No. 18 / 593,593, entitled "APPLICATION PROGRAMMING INTERFACE TO STORE AN IDENTIFIER OF A DATA STRUCTURE" (Attorney docket No. 0112912-C04US0), U.S. Patent Application Serial No. 18 / 593,710, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO INDICATE MULTIPROCESSOR AVIALLABILEITY" (Attorney docket No. 0112912-C05US0), U.S. Patent Application Serial No. 18 / 593,716, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO STORE IDENTIFIERS OF MULTIPROCESSOR GROUPS" (Attorney docket No. 0112912-C06US0), U.S. Patent Application Serial No. 18 / 593,720, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO INDICATE MULTIPROCESSOR GROUPS" (Attorney Docket No. 0112912-C07US0), U.S. Patent Application Serial No. 18 / 593,722, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO READ FROM A DATA STRUCTURE" (Attorney Docket No. 0112912-C08US0), U.S. Patent Application Serial No. 18 / 593,726, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO COMMUNICATION CONTEXT" (Attorney Docket No. 0112912-C09US0) and U.S. Patent Application Serial No. 18 / 593,731, filed concurrently herewith, entitled "APPLICATION PROGRAMMING INTERFACE TO WAIT FOR CONTEXT" (Attorney Docket No. 0112912-C10US0).REGIONAt least one embodiment relates to processing resources used to execute one or more software programs on a graphics processing unit ("GPU"). For example, at least one embodiment relates to performing application programming interfaces to manage resources of software programs.BACKGROUNDThe execution of computer programs can require considerable memory, time or resources. The amount of memory, time, and / or resources used to execute computer programs can be improved. Despite advances that speed up or otherwise support the execution of various components of a computer program, there are still challenges in executing computer programs with improved use of memory, time, and / or resources.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 is a block diagram illustrating a computer system for executing applications of computing resources via application programming interfaces (APIs), according to at least one embodiment; FIG. 2 is a block diagram illustrating contexts and sub-contexts according to at least one embodiment; FIG. 3 is a block diagram illustrating a software program executed by one or more processors according to at least one embodiment; FIG. 4 is a block diagram illustrating a process for executing one or more application programming interfaces (APIs), in accordance with at least one embodiment; FIG. 5 is a block diagram illustrating an application programming interface (API) to create a subcontext according to at least one embodiment; FIG. 6 is a block diagram illustrating a process for performing an application programming interface (API) to create a subcontext, in accordance with at least one embodiment; FIG. 7 is a block diagram illustrating an application programming interface (API) to destroy a subcontext, in accordance with at least one embodiment; FIG. 8 is a block diagram illustrating a process for performing an application programming interface (API) to destroy a sub-context, in accordance with at least one embodiment; FIG. 9 is a block diagram illustrating an application programming interface (API) to obtain a subcontext from a stream, in accordance with at least one embodiment; FIG. 10 is a block diagram illustrating a process for performing an application programming interface (API) to obtain a subcontext from a data stream, in accordance with at least one embodiment; FIG. 11 is a block diagram illustrating an application programming interface (API) to obtain resources associated with a context, in accordance with at least one embodiment; FIG. 12 is a block diagram illustrating a process for performing an application programming interface (API) to obtain resources associated with a context, in accordance with at least one embodiment; FIG. 13 is a block diagram illustrating an application programming interface (API) for partitioning context resources, in accordance with at least one embodiment; FIG. 14 is a block diagram illustrating a process for performing an application programming interface (API) to partition context resources, according to at least one embodiment; FIG. 15 is a block diagram illustrating an application programming interface (API) for generating a resource descriptor in accordance with at least one embodiment; FIG. 16 is a block diagram illustrating a process for performing an application programming interface (API) to generate a resource descriptor, in accordance with at least one embodiment; FIG. 17 is a block diagram illustrating an application programming interface (API) for retrieving device resources of a context, in accordance with at least one embodiment; FIG. 18 is a block diagram illustrating a process for performing an application programming interface (API) to obtain device resources of a context, in accordance with at least one embodiment; FIG. 19 is a block diagram illustrating an application programming interface (API) for recording a context event, according to at least one embodiment; FIG. 20 is a block diagram illustrating a process for performing an application programming interface (API) to record a context event, according to at least one embodiment; FIG. 21 is a block diagram illustrating an application programming interface (API) to wait for a context event, in accordance with at least one embodiment; FIG. 22 is a block diagram illustrating a process for performing an application programming interface (API) to wait for a context event, in accordance with at least one embodiment; FIG. 23 is a block diagram illustrating an example software stack in which application programming interfaces (API) are processed, in accordance with at least one embodiment; FIG. 24 is a block diagram illustrating a processor and modules according to at least one embodiment; FIG. 25 is a block diagram illustrating a driver and / or runtime that includes one or more libraries to provide one or more application programming interfaces (APIs), according to at least one embodiment; FIG. 26 illustrates an example data center, according to at least one embodiment; FIG. 27 illustrates a processing system according to at least one embodiment; FIG. 28 illustrates a computer system according to at least one embodiment; FIG. 29 illustrates a system according to at least one embodiment; FIG. 30 illustrates an example integrated circuit according to at least one embodiment; FIG. 31 illustrates a computer system according to at least one embodiment; FIG. 32 illustrates an APU according to at least one embodiment; FIG. 33 illustrates a CPU according to at least one embodiment; FIG. 34 illustrates an example accelerator integration slice, according to at least one embodiment; FIGS. 35A and 35B illustrate example graphics processors, in accordance with at least one embodiment; FIG. 36A illustrates a graphics core according to at least one embodiment; FIG. 36B illustrates a GPGPU according to at least one embodiment; FIG. 37A illustrates a parallel processor according to at least one embodiment; FIG. 37B illustrates a processing cluster in accordance with at least one embodiment; FIG. 37C illustrates a graphics multiprocessor in accordance with at least one embodiment; FIG. 38 illustrates a graphics processor according to at least one embodiment; FIG. 39 illustrates a processor according to at least one embodiment; FIG. 40 illustrates a processor according to at least one embodiment; FIG. 41 illustrates a graphics processor core in accordance with at least one embodiment; FIG. 42 illustrates a PPU according to at least one embodiment; FIG. 43 illustrates a GPC according to at least one embodiment; FIG. 44 illustrates a streaming multiprocessor in accordance with at least one embodiment; FIG. 45 illustrates a software stack of a programming platform according to at least one embodiment; FIG. 46 illustrates a CUDA implementation of a software stack of FIG. 45 according to at least one embodiment; FIG. 47 illustrates an ROCm implementation of a software stack of FIG. 45, in accordance with at least one embodiment; FIG. 48 illustrates an OpenCL implementation of a software stack of FIG. 45 in accordance with at least one embodiment; FIG. 49 illustrates software supported by a programming platform, in accordance with at least one embodiment; FIG. 50 illustrates compilation of code for execution on the programming platforms of FIGS. 45-48, according to at least one embodiment; FIG. 51 illustrates in greater detail the compilation of code for execution on the programming platforms of FIGS. 45-48, in accordance with at least one embodiment; FIG. 52 illustrates translating source code prior to compiling source code, in accordance with at least one embodiment; FIG. 53A illustrates a system configured to compile and execute CUDA source code using different types of processing units, in accordance with at least one embodiment; FIG. 53B illustrates a system configured to compile and execute CUDA source code of FIG. 53A using a CPU and a CUDA-enabled graphics processor, in accordance with at least one embodiment; FIG. 53C illustrates a system configured to compile and execute the CUDA source code of FIG. 53A using a CPU and a non-CUDA-capable GPU, in accordance with at least one embodiment; FIG. 54 illustrates an example kernel translated by the CUDA to HIP translation tool of FIG. 53C, in accordance with at least one embodiment; FIG. 55 shows a non-CUDA-capable GPU of FIG. 53C in greater detail, in accordance with at least one embodiment; FIG. 56 illustrates how threads of an example CUDA grid are mapped to different compute units of FIG. 55, in accordance with at least one embodiment; FIG. 57 illustrates how existing CUDA code may be migrated to Data Parallel C++ code, in accordance with at least one embodiment; and FIG. 58 illustrates components of a system for accessing a large language model, in accordance with at least one embodiment.DETAILED DESCRIPTIONIn at least one embodiment, software (e.g., a computer program) executed by one or more processors causes software executed by other processors (e.g., in a data center) to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, software executed by one or more processors executes one or more application programming interfaces (APIs) to create a subcontext that manages a subset of resources that can be used to execute software programs. In at least one embodiment, software executed by one or more processors executes one or more APIs to create a subcontext that manages a subset of resources that can be used to execute software programs, thereby causing software executed by other processors to perform operations to manage resources associated with software executed by other processors.In at least one embodiment, software executed by one or more processors executes one or more APIs to destroy a subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to destroy a subcontext, thereby causing software executed by other processors to perform operations to manage resources associated with software executed by other processors. In at least one embodiment, a subcontext that is destroyed by an application programming interface (API) to destroy a subcontext becomes a subcontext that is created by an API to create a subcontext as described above.In at least one embodiment, software executed by one or more processors executes one or more APIs to create a subcontext from a data stream (e.g., a series of operations performed in a particular order by one or more other processors). In at least one embodiment, software executed by one or more processors executes one or more APIs to create a subcontext from a data stream, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, a subcontext created by an API to create a subcontext functions as a subcontext created by an API to create a subcontext as described above, and is a subcontext that can be destroyed with an API to destroy a subcontext as also described above.In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain resources associated with a context or subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain resources associated with a context or subcontext, thereby causing software executed by other processors to perform operations to manage resources associated with software executed by other processors. In at least one embodiment, resources obtained when executing an API for creating resources associated with a context or subcontext are resources associated with a context or subcontext created as described above (e.g., using an API for creating a subcontext or an API for creating a subcontext from a stream).In at least one embodiment, software executed by one or more processors executes one or more APIs to subdivide context resources according to one or more criteria. In at least one embodiment, software executed by one or more processors executes one or more APIs to partition context resources, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, context resources partitioned by a context resource partitioning API are resources associated with a context or subcontext created as described above (e.g., using a subcontext creation API or a subcontext creation API from a stream). In at least one embodiment, resource partitions (e.g., obtained by a context resource partition API) are used to generate sub-contexts as described herein.In at least one embodiment, software executed by one or more processors executes one or more APIs to generate a resource descriptor from a list of resources. In at least one embodiment, software executed by one or more processors executes one or more APIs to generate a resource descriptor, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, API generated list of resource descriptors is used to generate a resource descriptor to create and / or manage subcontext as described above.In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain device resources associated with a context or subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain device resources, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, device resources obtained by execution of an API for retrieving device resources are usable to create a subcontext and / or to further subdivide resources, as described above.In at least one embodiment, software executed by one or more processors executes one or more APIs to record a context event (e.g., to issue an event associated with another context so that synchronization between contexts may be performed). In at least one embodiment, software executed by one or more processors executes one or more APIs to record a context event, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, a context event recorded by an API to record a context event is used by a context connected to an API to wait for a context event (see below).In at least one embodiment, software executed by one or more processors executes one or more APIs to wait for a context event (e.g., to wait for an event associated with another context so that synchronization between contexts may be performed). In at least one embodiment, software executed by one or more processors executes one or more APIs to wait for a context event, thereby causing software executed by other processors to perform operations for managing resources associated with software executed by other processors. In at least one embodiment, a context event of an API waiting for a context event is used by a context connected to an API to record a context event, as described above.FIG. 1 is a block diagram 100 illustrating a computer system for executing applications of computing resources via application programming interfaces (APIs), according to at least one embodiment. In at least one embodiment, a processor 102 executes one or more software programs 104. In at least one embodiment, processor 102 is a processor as described below. In at least one embodiment, processor 102 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general graphics processing unit (GPGPU), a compute cluster, and / or a combination of these and / or other such processors. In at least one embodiment, processor 102 is part of a computer system as described herein. In at least one embodiment, not shown in FIG. 1, processor 102 is a processor of a client computer system. In at least one embodiment, a client computer system includes one or more client devices as described herein. In at least one embodiment, a client computing system includes one or more devices that are clients of a cloud computing environment, as described herein.In at least one embodiment, software programs 104 include one or more computational operations, such as, for example, computational operations for training a neural network, executing a neural network, executing a compute uniform device architecture (CUDA) program, executing a large language model, executing a rendering operation, executing data analysis, and / or executing other operations, including, but not limited to, those described herein. In at least one embodiment, software programs 104 include software as described herein.In at least one embodiment, software programs 104 perform resource operations 106 using resource operation APls 108. In at least one embodiment, resource operations 106 are operations to manage resources associated with processor 110. In at least one embodiment, processor 110 is a processor as described below. In at least one embodiment, processor 110 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general graphics processing unit (GPGPU), a compute cluster, and / or a combination of these and / or other such processors. In at least one embodiment, processor 110 is part of a computer system as described herein.In at least one embodiment, not shown in FIG. 1, processor 110 is one of a plurality of processors of a high-performance computer system. In at least one embodiment, a high-performance computer system is a computer system that includes a plurality of processors to perform computational operations such as those described herein. In at least one embodiment, a high-performance computer system is a distributed computer system. In at least one embodiment, a high-performance computer system is a deep learning computer system. In at least one embodiment, operations for performing computational operations are performed by a high-performance computer system using systems, methods, operations, and / or techniques described herein. In at least one embodiment, a high-performance computer system includes a cloud computing environment as described herein. In at least one embodiment, a high performance computer system includes one or more processors, such as processor 110. In at least one embodiment, a high-performance computer system includes one or more graphics processors as described herein. In at least one embodiment, processors of a high performance computer system include one or more central processing units (CPUs), graphics processing units (GPUs), parallel processing units (PPUs), general graphics processing units (GPGPUs), compute clusters, and / or a combination of these and / or other such processors as described herein.In at least one embodiment, a high-performance computer system includes a plurality of processors, such as processor 110, having an associated set of resources that may be used by a processor, such as processor 102, to perform computational operations, including one or more resource operations 106 using one or more resource operation APls 108. In at least one embodiment, resource operations 106 are compute operations for managing resources of processes performed by one or more processors using systems, methods, and / or operations as described herein.In at least one embodiment, resource operations APIs 108 include one or more APIs for managing resources of processors, such as processor 110, that may be used to execute software programs 104. In at least one embodiment, resource operations APIs 108 include one or more APIs such as subcontext of stream obtained API 502 (described herein at least in connection with FIGS. 5 and 6 ), subcontext destructive API 702 (described herein at least in connection with FIGS. 7 and 8 ), subcontext of stream obtained API 902 (described herein at least in connection with FIGS. 9 and 10 ), obtained context resource API 1102 (described herein at least in connection with FIGS. 11 and 12 ), context resource sub-APl 1302 (described herein at least in connection with FIGS. 13 and 14 ), Obtain device resource API 1502 (described herein at least in connection with FIGS. 15 and 16 ), obtain device resource API 1702 (described herein at least in connection with FIGS. 17 and 18 ), context event inclusion API 1902 (described herein at least in connection with FIGS. 19 and 20 ), wait for context event API 2102 (described herein at least in connection with FIGS. 21 and 22 ), and / or other such resource operation APIs.In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more contexts 112 that include context resources 114. In at least one embodiment, context 112 includes a description of resources available to a processor such as processor 110, as described below at least in connection with FIG. 2. In at least one embodiment, processor resources 114 include resources available to processor 110 for execution of one or more workloads (see below). In at least one embodiment, a context 112 is a primary context of processor 110, which is a pre-created and / or default context of processor 110.In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more sub-contexts (e.g., sub-context 118A- 118N) that may be used to execute one or more workloads (e.g., workload 120A- 120N). In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more sub-contexts based on one or more subsets of processor resources 114 (e.g., subset 116A- 116N). In at least one embodiment, subcontexts (e.g., subcontext 118A- 118N) are used to execute workloads (e.g., workload 120A- 120N). In at least one embodiment, a subcontext is derived from context 112 that uses a subset of resources available to processor 110 (e.g., as described herein at least in connection with FIG. 2 ). In at least one embodiment, a subcontext is referred to as a green context. In at least one embodiment, a subset of processor resources 114 (e.g., subset 116A- 116N) includes some or all of processor resources 114 (e.g., a first subset of processor resources 114 may include a first portion of said processor resources, a second subset of processor resources 114 may include a second portion of said processor resources, etc.). In at least one embodiment, a subset of processor resources are not empty (e.g., it includes at least a portion of processor resources 114). In at least one embodiment, a workload (e.g., workload 120A- 120N) includes at least a portion of computational operations of software programs 104 to be executed by one or more processors, such as processor 110, using systems, methods, and operations as described herein. In at least one embodiment, a workload (e.g., workload 120A- 120N) is also referred to as a software workload. In at least one embodiment, a workload (e.g., workload 120A- 120N) is also referred to as a kernel. In at least one embodiment, a workload (e.g., workload 120A- 120N) is also referred to as a software kernel.In at least one embodiment, not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to create a subcontext using subcontext create API 502, as described herein in connection with at least FIGS. 5 and 6. In at least one embodiment not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to destroy a subcontext using subcontext destroying API 702 described herein in connection with at least FIGS. 7 and 8. In at least one embodiment not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to obtain a subcontext from a stream by using subcontext from stream obtaining API 902 described herein in connection with at least FIGS. 9 and 10. In at least one embodiment, not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to obtain resources associated with a context by using obtain context resource APl 1102 described herein in connection with at least FIGS. 11 and 12. In at least one embodiment, not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more context resource partitioning operations using context resource partitioning APl 1302 described herein at least in connection with FIGS. 13 and 14. In at least one embodiment, not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to generate a resource descriptor using resource descriptor generation API 1502 described herein in connection with at least FIGS. 15 and 16. In at least one embodiment not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to retrieve resources of a device using obtain device resource API 1702 (described herein at least in connection with FIGS. 17 and 18 ). In at least one embodiment not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more context event recording operations using context event recording API 1902 described herein in connection with at least FIGS. 19 and 20. In at least one embodiment not shown in FIG. 1, processor 110 receives resource operations APIs 108 and performs one or more operations to wait for a context event using wait for context event API 2102, which is described herein at least in connection with FIGS. 21 and 22.In at least one embodiment, a plurality of subsets of processor resources 114 (e.g., a plurality of subsets of subset 116A- 116N) are used by and / or associated with a subcontext (e.g., one of subcontext 118A- 118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A- 116N) is used by and / or is associated with a sub-context (e.g., one of sub-contexts 118A- 118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A- 116N) is used by and / or is associated with a plurality of sub-contexts (e.g., a plurality of sub-contexts of sub-context 118A- 118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A- 116N) is different than a number of subcontexts (e.g., subcontext 118A- 118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A- 116N) is identical to a number of subcontexts (e.g., subcontext 118A- 118N).In at least one embodiment, a plurality of sub-contexts (e.g., a plurality of sub-contexts of sub-context 118A- 118N) are used by and / or associated with a sub-context (e.g., one of sub-context 118A- 118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A- 116N) is used by and / or is associated with a sub-context (e.g., one of sub-contexts 118A- 118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A- 116N) is used by and / or is associated with a plurality of sub-contexts (e.g., a plurality of sub-contexts of sub-context 118A- 118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A- 116N) is different than a number of subcontexts (e.g., subcontext 118A- 118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A- 116N) is identical to a number of subcontexts (e.g., subcontext 118A- 118N).In at least one embodiment, as used in each implementation described herein, terms such as "module" and nominally verbs (e.g., subcontext, workload, context, and / or other terms) each refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide functionality described herein, unless context dictates otherwise or expressly state otherwise. In at least one embodiment, software may be embodied as a software package, code and / or instruction set or instructions, and "hardware" as used in any implementation described herein may include, for example, singly or in any combination, hardwired circuits, programmable circuits, state machine circuits, fixed function circuits, execution unit circuits, and / or firmware storing instructions executed by programmable circuits. In at least one embodiment, modules may be collectively or individually embodied as circuits that are part of a larger system, such as integrated circuits (ICs), system-on-a-chip (SoC), etc.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use. 1, such as one or more circuits for executing an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to execute one or more software threads and / or otherwise to perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to generate one or more masks that indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more software threads and / or to otherwise perform operations described herein. 1, such as one or more circuits for executing an application programming interface (API) to generate one or more masks that indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs), which may be used to perform one or more kernels and / or otherwise perform operations described herein. In at least one embodiment, systems, methods, operations, and / or instructions described herein in connection with FIG. 1 are included and / or otherwise included to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use for executing one or more software threads and / or otherwise for executing operations described herein. In at least one embodiment, components described in connection with FIG. 1 execute one or more of processes described in connection with FIGS. 1-25 to execute an application programming interface (API) to associate one or more data structures indicating which of one or more streaming multiprocessors (SMs) of one or more processors to use for executing one or more software threads and / or otherwise executing operations described herein. In at least one embodiment not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more of components described in connection with FIGS. 26-58 to execute an application programming interface (API) to associate one or more data structures including indicating which of one or more streaming multiprocessors (SMs) of one or more processors to use for executing one or more software threads and / or otherwise executing operations described herein.In at least one embodiment not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, performs operations described herein in connection with at least FIGS. 1-25, such as operations for executing an application programming interface (API) to assign one or more data structures indicating which of one or more streaming multiprocessors (SMs) of one or more processors to use for executing one or more software threads and / or for otherwise executing operations described herein. 1-25, such as operations for executing an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to execute one or more software threads and / or otherwise to execute the operations described herein. In at least one embodiment not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, performs operations described herein in connection with at least FIGS. 1-25, such as operations to execute an application programming interface (API) to generate one or more masks to display one or more subsets of streaming multiprocessors of one or more graphics processor units (GPUs) that may be used to perform one or more operations. 1-25, such as operations for executing an application programming interface (API) to generate one or more masks indicating one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels and / or otherwise perform the operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use. 1, such as one or more circuits for executing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use for executing one or more software threads and / or for otherwise executing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to cause deactivation of one or more masks, where one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processors. 1, such as one or more circuitry for executing an application programming interface (API) to cause deactivation of one or more masks, the one or more masks indicating one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels and / or otherwise to perform the operations described herein. In at least one embodiment, systems, methods, operations, and / or instructions described herein in connection with FIG. 1 are included and / or otherwise included to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) to use by one or more processors to execute one or more software threads and / or otherwise execute operations described herein. In at least one embodiment, components described in connection with FIG. 1 execute one or more of processes described in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) to use by one or more processors to execute one or more software threads and / or otherwise execute operations described herein. In at least one embodiment not shown in FIG. 1, one or more components, herein in conjunction with FIG. 1, include one or more components described herein in conjunction with FIGS. 26-58 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) is to be used by one or more processors to execute one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, not shown in FIG. 1, a set of instructions stored on a non-transitory machine-readable medium are to perform, when executed by one or more processors, operations described herein in at least connection with FIGS. 1-25, such as operations to perform an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a set of instructions stored on a non-transitory machine-readable medium are to perform operations, when executed by one or more processors, that perform operations described herein in at least one of FIGS. 1-25, such as operations to perform an application programming interface (API) to cause one or more masks to be deactivated, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable to perform one or more kernels and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more interfaces to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to indicate one or more identifiers of one or more masks, where one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable to perform one or more kernels and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed.In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors to perform operations described herein in connection with at least FIGS. 1-25, such as operations for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a set of instructions stored on a non-transitory machine-readable medium are to perform operations, when executed by one or more processors, that perform operations described herein at least in connection with FIGS. 1-25, such as operations to perform an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable to perform one or more kernels and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing applications and / or instructions described herein in connection with FIG. 1, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for use in performing one or more software threads and / or for otherwise performing operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 1, and / or instructions, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in executing one or more software kernels and / or for performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for execution of one or more software threads and / or otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available to be used to execute one or more software threads and / or other operations described herein.In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes a set of instructions that, when executed by one or more processors, are operable to perform operations described herein, at least in conjunction with FIGS. 1-25, such as operations for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or for otherwise executing operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium stores an instruction set that, when executed by one or more processors to perform operations described herein in connection with at least FIGS. 1-25, such as operations to perform an application programming interface (API) to display one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in performing one or more software kernels and / or otherwise performing operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to perform an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs is to be used and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more interfaces for performing operations described herein in connection with FIG. 1 and / or instructions, such as one or more circuits for performing an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available for use in performing one or more software kernels and / or for otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs and / or otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs and / or otherwise perform operations described herein.In at least one embodiment, not shown in FIG. 1, a set of instructions stored on a non-transitory machine-readable medium are to perform, when executed by one or more processors, operations described herein in at least connection with FIGS. 1-25, such as operations to perform an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, is operable to perform operations described herein at least in connection with FIGS. 1-25, such as operations for executing an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available to be used for executing one or more software kernels and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits for performing an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 1, and / or instructions, such as one or more circuits for performing an application programming interface (API) to generate a data structure containing information to indicate one or more streaming multiprocessors (SMs) of one or more GPUs as usable for performing one or more software kernels and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations described otherwise herein are performed. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations described otherwise herein are performed.In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, performs operations described herein in connection with at least FIGS. 1-25, such as operations for performing an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or otherwise performs operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium stores an instruction set that, when executed by one or more processors, performs operations described herein in connection with at least FIGS. 1-25, such as operations for executing an application programming interface (API) to generate a data structure including information indicating that one or more streaming multiprocessors (SMs) of one or more GPUs may be used to execute one or more software cores and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 1, and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause a number of streaming multiprocessors indicated by one or more masks to be used for performing one or more software kernels to be indicated, and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.In at least one embodiment, not shown in FIG. 1, a set of instructions is stored on a non-transitory machine-readable medium that, when executed by one or more processors, are to perform operations described herein in at least connection with FIGS. 1-25, such as operations to perform an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, is operable to perform operations described herein, at least in connection with FIGS. 1-25, such as operations for executing an application programming interface (API) to cause a number of streaming multiprocessors indicated by one or more masks to be usable to execute one or more software cores to be indicated and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits for performing an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions and / or operations otherwise described herein to be performed. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits for performing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of processor to be indicated to one or more resources of a second context of processor and / or otherwise operations described herein to be performed. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause context of one or more first software programs to be communicated to one or more second software programs, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause context of one or more first software commands to be communicated to one or more second software commands, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause context of one or more first software commands to be communicated to one or more second software commands, and / or to otherwise perform operations described herein.In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium stores an instruction set that, when executed by one or more processors to perform operations described herein in connection with at least FIGS. 1-25, such as operations to perform an application programming interface (API), to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes an instruction set that, when executed by one or more processors, is operable to perform operations described herein, at least in connection with FIGS. 1-25, such as operations for executing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of processor to be indicated to one or more resources of a second context of processor, and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 1, such as one or more circuits to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing applications and / or instructions described herein in connection with FIG. 1, such as one or more circuits for performing an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of processor until one or more indications of one or more events of a second context are generated and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 1 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software instructions and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 1 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software programs and / or otherwise execute operations described herein. In at least one embodiment, not shown in FIG. 1, one or more components described herein in connection with FIG. 1 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more second commands to wait until the one or more second commands obtain a context corresponding to one or more first software commands and / or otherwise execute operations described herein.In at least one embodiment, not shown in FIG. 1, a set of instructions stored on a non-transitory machine-readable medium that, when executed by one or more processors to perform operations described herein in connection with at least FIGS. 1-25, such as operations for executing an application programming interface (API), cause one or more second instructions to wait until the one or more second instructions obtain a context corresponding to one or more first software instructions, and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 1, a non-transitory machine-readable medium includes a set of instructions that, when executed by one or more processors to perform operations described herein in connection with at least FIGS. 1-25, such as operations to perform an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of processor until one or more indications of one or more events of a second context are generated and / or otherwise perform operations described herein.FIG. 2 is a block diagram 200 illustrating contexts and sub-contexts in accordance with at least one embodiment. In at least one embodiment, a device 202 is used to carry out one or more workloads (e.g., one or more of workloads 120A- 120N described herein in at least connection with FIG. 1 ). In at least one embodiment, device 202 includes one or more processors, such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, device 202 includes one or more channels used to execute streams of workloads (e.g., a first workload followed by a second workload, etc.). In at least one embodiment, not shown in FIG. 2, device 202 includes a hardware context (e.g., a set of hardware resources). In at least one embodiment, device 202 includes a plurality of hardware contexts. In at least one embodiment, performance of device 202 is improved if device 202 has only one hardware context. In at least one embodiment, a hardware context of device 202 includes one or more primary contexts associated with that hardware context. In at least one embodiment, a primary context includes a description of and / or references to resources of device 202 (e.g., hardware context). In at least one embodiment, a primary context is a one-to-one mapping with a hardware context of device 202 (e.g., it includes only references to hardware resources of device 202).In at least one embodiment, one or more channels of device 202 are primary context channels 212 (e.g., C1 204, C2 206, C3 208, and C4 210) used by a primary context to perform software workloads using device 202. In at least one embodiment, primary context channels 212 are used exclusively by a primary context for executing software workloads. In at least one embodiment, one or more channels of device 202 are preferred selection channels 232 (e.g., C5 214, C6 220, C7 226, and C8 230). In at least one embodiment, preferred sub-context selection channels 232 are used to execute software workloads. In at least one embodiment, sub-contexts are dynamically assigned to one or more preferred selection channels 232. In at least one embodiment, a subcontext (also referred to as green context) is a subset of a context (e.g., a subset of resources of device 202) that does not require a context switch (e.g., a hardware context switch). In at least one embodiment, a subcontext is a one-to-one mapping of a subset of resources of a hardware context. In at least one embodiment, software programs, such as software programs 104 described herein in at least connection with FIG. 1, perform computational operations to transition between sub-contexts when workloads, such as workload 120A- 120N, are executed using device 202. In at least one embodiment, one or more sub-contexts are managed by a single thread executing on device 202. In at least one embodiment, only one subcontext managed by a thread may be current. In at least one embodiment, a thread performs operations to switch between subcontexts managed by that thread to keep a subcontext current at a time.In at least one embodiment, a sub-context is associated with a plurality of channels of preferred selection channels 232 (e.g., sub-context 216 associated with both C5 214 and C6 220 and sub-context 218 also associated with both C5 214 and C6 220). In at least one embodiment, a subcontext is assigned to a single channel (e.g., subcontext 222 assigned to C7 226 and subcontext 224 also assigned to C7 226). In at least one embodiment, a subcontext is dedicated to only one channel (e.g., subcontext 228 dedicated to C8 230).In at least one embodiment, a primary context is referred to as a full context. In at least one embodiment, a full context includes access to all execution resources of device 202. In at least one embodiment, a primary context is a default context that a CUDA runtime uses to manage resources of device 202. In at least one embodiment, a primary context is thread-safe. In at least one embodiment, a subcontext is mapped to a primary context and may use a subset of resources of device 202. In at least one embodiment, a subcontext may also use other resources (e.g., resources from other devices and / or contexts).In at least one embodiment, a subcontext includes a resource group for managing a subset of resources used by subcontext. In at least one embodiment, a subcontext may implicitly specify a set of resources that a software library or framework (e.g., as described herein) may use. In at least one embodiment, a resource group represents one or more resources that may use a particular context, subcontext, or API. In at least one embodiment, a resource group is a descriptor (e.g., a description of resources) that may be partitioned by one or more APIs, such as those described herein, to manage resources (e.g., in a hierarchical manner). In at least one embodiment, a resource group includes descriptions or listings of a plurality of different resources of device 202, including, but not limited to, streaming multiprocessors (SMs), device connections (e.g., copy and compute hardware channels), processor bindings, software schedulers, work distribution, etc.). In at least one embodiment, other APIs (e.g., CUDA APls) may be assigned to a subcontext (e.g., a current subcontext) and utilize resources of that subcontext. In at least one embodiment, other APIs (e.g., CUDA APls) may be directly assigned to a resource group (e.g., without using a sub-context).In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 2, such as one or more circuits for performing an application programming interface (API) to assign one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 2 and / or instructions, such as one or more circuits for performing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 2, such as one or more circuits for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing applications and / or instructions described herein in connection with FIG. 2, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for use in performing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API), to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API), to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used to perform one or more software threads, and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 2 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of streaming multiprocessor (SM) identifiers of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or operations described otherwise herein are to be performed. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 2, such as one or more circuits for performing an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or for otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 2, such as one or more circuits to execute an application programming interface (API), to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 2 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause context of one or more first software commands to be communicated to one or more second software commands, and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 2, and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 2 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software instructions, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 2 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software programs and / or otherwise execute operations described herein. In at least one embodiment, not shown in FIG. 2, one or more components described herein in connection with FIG. 2 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more second instructions to wait to execute until the one or more second instructions obtain a context corresponding to one or more first software instructions and / or otherwise execute operations described herein.FIG. 3 is a block diagram 300 illustrating a software program executed by one or more processors, in accordance with at least one embodiment. In at least one embodiment, block diagram 300 illustrates a software program 304 executed by a processor, such as a central processing unit (CPU) 302, as well as a graphics processing unit (GPU) 310, and an accelerator 314 within a heterogeneous processor. In at least one embodiment, CPU 302 is a processor such as processor 102 described herein at least in connection with FIG. 1. In at least one embodiment, CPU 302 is a processor such as processor 110 described herein in at least conjunction with FIG. 1. In at least one embodiment, a CPU 302 is any processor having an architecture further described herein. In at least one embodiment, a CPU 302 is any general processor having any architecture further described herein. In at least one embodiment, a processor, such as a CPU 302, includes circuitry for performing one or more computational operations. In at least one embodiment, a processor, such as a CPU 302, includes any configuration of circuitry to perform one or more computational operations further described herein.In at least one embodiment, a processor, such as a central processing unit (CPU) 302, executes a parallel computing environment 308. In at least one embodiment, a processor, such as a CPU 302, executes a parallel computing environment 308, such as Compute Uniform Device Architecture (CUDA). In at least one embodiment, parallel computing environment 308 includes instructions that, when executed by one or more processors, such as CPUs 302, facilitate execution of one or more software programs by one or more CPUs 302, one or more parallel processing units (PPUs), such as GPUs 310, and / or one or more accelerators 314 within a heterogeneous processor.In at least one embodiment, one or more PPUs are processors that include one or more circuits for performing parallel computing operations, such as GPUs 310 and any other parallel processor further described herein. In at least one embodiment, GPU 310 is hardware that includes circuitry for performing one or more computational operations, as described further below in connection with various embodiments. In at least one embodiment, a GPU 310 includes one or more processing cores, each of which performs one or more computational operations. In at least one embodiment, a GPU 310 includes one or more processing cores to perform one or more parallel computational operations. In at least one embodiment, a GPU 310 is packaged together with a CPU 302 or other processors as a system-on-a-chip (SoC). In at least one embodiment, a GPU 310 is packaged on a common chip or other substrate with a CPU 302 or other system-on-a-chip (SoC) processors. In at least one embodiment, one or more accelerators 314 within heterogeneous processors are hardware that includes one or more circuits for performing specific computational operations, such as a deep learning accelerator (DLA), a programmable image processing accelerator (PVA), a field programmable gate array (FPGA), or other accelerator further described herein. In at least one embodiment, an accelerator 314 is packaged within a heterogeneous processor along with a CPU 302 or other processors as a system-on-a-chip (SoC). In at least one embodiment, an accelerator 314 is packaged within a heterogeneous processor on a common chip or other substrate with a CPU 302 or other system-on-a-chip (SoC) processors. In at least one embodiment, one or more CPUs 302, one or more GPUs 310, or other PPUs and / or accelerators 314 are housed in heterogeneous processors as a system-on-a-chip (SoC). In at least one embodiment, one or more CPUs 302, one or more GPUs 310, or other PPUs and / or accelerators 314 are packaged in heterogeneous processors on a common chip or other substrate than system-on-chip (SoC).In at least one embodiment, parallel computing environment 308 such as CUDA includes libraries and other software programs for performing one or more computational operations using one or more PPUs such as GPUs 310 and / or one or more accelerators 314 within a heterogeneous processor. In at least one embodiment, parallel computing environment 308 includes libraries and other software programs that, when executed by one or more processors, such as one or more CPUs 302, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within a heterogeneous processor to perform one or more computational operations. In at least one embodiment, parallel computing environment 308 includes libraries that, when executed, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors to perform mathematical operations. In at least one embodiment, parallel computing environment 308 includes libraries that, when executed, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors to perform any further operation described herein.In at least one embodiment, one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors perform one or more computational operations in response to one or more application programming interfaces (APIs). In at least one embodiment, an API is an instruction set that, when executed by one or more processors, such as CPUs 302, causes one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 in heterogeneous processors to perform one or more computational operations. In at least one embodiment, parallel computing environment 308 includes one or more APIs 306 that, when executed by one or more processors, such as CPUs 302, cause one or more PPUs, such as GPUs 310 and / or one or more accelerators 314 within heterogeneous processors to perform one or more computational operations. In at least one embodiment, one or more APIs 306 include one or more functions that, when executed, cause one or more processors, such as CPUs 302, to perform one or more operations, such as computational operations, error reporting, scheduling of other operations to be performed by GPUs 310 and / or accelerators 314 within heterogeneous processors, or any other operation further described herein. In at least one embodiment, one or more APIs 306 include one or more functions that, when executed, cause one or more PPUs, such as GPUs 310, to perform one or more operations, such as compute operations, error reports, or any other operation further described herein. In at least one embodiment, one or more APIs 306 include one or more functions, such as those described below in connection with FIGS. 5-22, which, when executed, cause one or more accelerators 314 in heterogeneous processors to perform one or more operations, such as computational operations, error reporting, or any other operation described herein. In at least one embodiment, one or more APIs 306 include one or more functions to cause a CPU 302 to perform one or more computing operations in response to information or events generated by one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 in heterogeneous processors. In at least one embodiment, one or more APIs 306 include one or more functions that, when invoked, cause a CPU 302 to perform one or more computational operations in response to information or events generated by one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 in heterogeneous processors.In at least one embodiment, a processor, such as a CPU 302, executes one or more software programs 304. In at least one embodiment, one or more software programs are sets of instructions that, when executed, cause one or more processors, such as CPUs 302, PPUs, such as GPUs 310, and / or accelerators 314 in heterogeneous processors to perform computational operations. In at least one embodiment, software programs 304 include instructions and / or operations executed by one or more PPUs, such as GPUs 310. In at least one embodiment, one or more software programs 304 include GPU-specific code 312 and / or accelerator-specific code 316. In at least one embodiment, instructions and / or operations to be executed by one or more PPUs, such as GPUs 310, are PPU-specific or GPU-specific code 312. In at least one embodiment, GPU-specific code 312 is a set of software instructions and / or other operations, as further described herein, to be executed by one or more GPUs 310. In at least one embodiment, software programs 304 include instructions and / or operations to be executed by one or more accelerators 314 in heterogeneous processors. In at least one embodiment, instructions and / or operations to be executed by one or more accelerators 314 in heterogeneous processors are accelerator specific code 316. In at least one embodiment, accelerator-specific code 316 is a set of software instructions and / or other operations, as further described herein, to be executed by one or more accelerators 314. In at least one embodiment, PPU-specific or GPU-specific code 312 and / or accelerator-specific code 316 is executed in response to one or more APIs 306, as described below in connection with FIGS. 5-22.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for performing an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or other operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for performing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed.In at least one embodiment, one or more processors (e.g., those described herein) include one or more interfaces for executing applications described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API), to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available to be used to execute one or more software threads and / or other operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIG. 3, such as one or more circuits for performing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of streaming multiprocessor (SM) identifiers of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or operations described otherwise herein are to be performed. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for performing an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations described otherwise herein are performed.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more interfaces to perform operations and / or instructions described herein in connection with FIG. 3, such as one or more circuits to perform an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause context of one or more first software commands to be communicated to one or more second software commands, and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 3, and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 3 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software instructions, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 3 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software programs and / or otherwise execute operations described herein. In at least one embodiment, not shown in FIG. 3, one or more components described herein in connection with FIG. 3 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more second commands to wait until the one or more second commands obtain a context corresponding to one or more first software commands and / or otherwise execute operations described herein.FIG. 4 is a block diagram 400 illustrating a process for performing one or more application programming interfaces (APIs) in accordance with at least one embodiment. In at least one embodiment, process shown in block diagram 4 for executing one or more APIs uses one or more accelerators within a heterogeneous processor through a parallel computing environment, such as parallel computing environment 308, as described herein at least in connection with FIG. 3. In at least one embodiment, process shown in block diagram 400 begins at step 404 for execution of one or more APIs 402, where one or more processors execute a software program comprising one or more instructions that, when executed, cause one or more processors and / or one or more other processors, such as graphics processing units (GPUs) and / or one or more accelerators within a heterogeneous processor or processors, to perform one or more computational operations. In at least one embodiment, a software program to be executed by one or more processors includes, at step 404, one or more instructions that, when executed, cause execution of one or more APIs 306 of a parallel computing environment 308 as described above. In at least one embodiment, after step 404, process shown in block diagram 400 for executing one or more APIs continues at step 406.In at least one embodiment, a processor performing process shown in block diagram 400 for performing one or more APls determines, at step 406 of process for performing one or more APls shown in block diagram 400, whether to perform an API such as those described herein in connection with at least FIGS. 5-22 (e.g., sub-context create API 502, sub-context destroy API 702, obtain context resource of stream API 902, obtain context of stream obtain API 1102, context resource partition API 1302, generate resource descriptor API 1502, obtain device resource API 1702, Get Context Resource Inclusion API 1902 and / or wait for Context Event API 2102). In at least one embodiment, if it is determined that an API is not being executed ("NO" branch), then the process shown in block diagram 400 for executing one or more APls continues at step 416. In at least one embodiment, if it is determined that an API is being performed ("YES" branch), then the process shown in block diagram 400 for performing one or more APls continues at step 408.In at least one embodiment, in step 408 of process for performing one or more APls shown in block diagram 400, a processor performing process for performing one or more APls shown in block diagram 400 performs an API as described herein in at least connection with FIGS. 5-22. In at least one embodiment, at step 408, one or more processors execute one or more instructions to execute one or more API calls as described herein in at least connection with FIGS. 5-22 (e.g., Create SubKontext API 502, Destroy SubKontext API 702, Obtain SubKontext of Stream API 902, Obtain Context Resource API 1102, Subdivide Context Resource API 1302, Create Resource Descriptor API 1502, Obtain Resource Descriptor API 1702, Record Context Event API 1902, and / or Wait for Context Event API 2102) by one or more processors and / or by one or more other processors, such as GPUs and / or accelerators within a heterogeneous processor as described above. In at least one embodiment, after step 408, process shown in block diagram 400 for executing one or more APls continues at step 410.In at least one embodiment, a processor executing process shown in block diagram 400 determines, at step 410 of process for executing one or more APls shown in block diagram 400, whether a return value should be returned as a result of executing one or more commands for executing one or more API calls as described herein in at least connection with FIGS. 5-22 (e.g., Create SubKontext API 502, Destroy SubKontext API 702, Obtain SubKontext of Stream API 902, Obtain Context Event API 1102, Subdivide Context Resource API 1302, Create Resource Descriptor API 1502, Obtaining resource descriptor API 1702, recording context event API 1902, and / or waiting for context event API 2102) by the one or more processors and / or by one or more other processors, such as GPUs and / or accelerators within a heterogeneous processor, as described above. In at least one embodiment, a processor performing process shown in block diagram 400 for executing one or more APls determines whether to return a return value using an API return as described herein at least in connection with FIGS. 5-22 (e.g., obtain context resource API return 520, subcon context destroy API return 720, obtain context of stream obtain API return 920, obtain context resource API return 1120, obtain context resource API return 1320, obtain context event record API return 1520, obtain resource descriptor API return 1720, in step 410, Resource descriptor inclusion API return 1920 and / or wait for context event API return 2120). In at least one embodiment, if it is determined that a return value is returned ("YES" branch), then in step 410, process for performing one or more of APIs shown in block diagram 400 continues in step 412. In at least one embodiment, if it is determined that no return value is returned ("NO" branch), then in step 410, process for performing one or more of APIs shown in block diagram 400 continues in step 414.In at least one embodiment, a processor performing process shown in block diagram 400 for executing one or more APIs sets a return value in step 412. In at least one embodiment, in step 412, a return value is set by storing return value at a storage location specified by an API as described herein at least in connection with FIGS. 5-22 (e.g., sub-context-get API 502, sub-context-destroy API 702, sub-context-of-stream-get API 902, context event-get API 1102, context resource-sub-API 1302, resource descriptor-generate API 1502, device resource API 1702, context event-capture API 1902, and / or waiting for context event API 2102). In at least one embodiment, in step 412, a return value is determined by storing return value at a storage location that includes an API return as described herein in connection with at least FIGS. 5-22 (e.g., Get Context Resource API return 520, SubKontext Destroy APl return 720, Get Context of Stream Get API return 920, Get Context Resource API return 1120, Get Context Resource API return 1320, Get Resource Descriptor Inclusion API return 1520, Get Resource Descriptor Inclusion API return 1720, Get resource descriptor inclusion API return 1920 and / or wait for context event API return 2120). In at least one embodiment, after step 412, process shown in block diagram 400 for executing one or more APIs continues at step 414.In at least one embodiment, a processor performing process illustrated in block diagram 400 for executing one or more APIs returns success or failure (e.g., an error) using an API return as described herein in at least connection with FIGS. 5-22 (e.g., obtaining sub-context API return 520, destroying sub-context API return 720, obtaining sub-context from stream API return 920, obtaining context resources API return 1120, dividing context resources API return 1320, generating resource descriptor API return 1520, obtaining device resources API return 1720, in step 414, Record Context Events API Return 1920 and / or Wait for Context Events API Return 2120). In at least one embodiment, after step 414, process shown in block diagram 400 for executing one or more APIs continues at step 416.In at least one embodiment, a processor executing process shown in block diagram 400 for executing one or more APIs determines whether execution of software program (e.g., started in step 404) is complete in step 416 of process shown in block diagram 400 for executing one or more APls. In at least one embodiment, a processor executing process shown in block diagram 400 for executing one or more APIs determines, at step 416, that execution of software program (e.g., started at step 404) is complete based at least in part on whether one or more processors execute instructions of software program (e.g., started at step 404). In at least one embodiment, if it is determined that execution of software program (e.g., initiated at step 404) is complete, the process shown in block diagram 400 for executing one or more APIs is terminated (418). In at least one embodiment, if it is determined that execution of software program (e.g., started at step 404) is not complete, process shown in block diagram 400 for executing one or more APls continues at step 404 to continue execution of one or more instructions of a software program.In at least one embodiment, operations of process shown in block diagram 400 for executing one or more APIs are performed in a different order than shown in FIG. 4. In at least one embodiment, operations of process shown in block diagram 400 for executing one or more APIs are performed simultaneously or in parallel. In at least one embodiment, operations of process shown in block diagram 400 for performing one or more APls that are not dependent on each other (e.g., are independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process for executing one or more of APIs shown in block diagram 400 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4 and / or instructions, such as one or more circuits for performing an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4 and / or instructions, such as one or more circuits for performing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or operations otherwise described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 4, such as one or more circuits to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing applications and / or instructions described herein in connection with FIG. 4, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for use in performing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API), to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used to perform one or more software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of streaming multiprocessor (SM) identifiers of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4 and / or instructions, such as one or more circuits for performing an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIG. 4, such as one or more circuits to execute an application programming interface (API), to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4 and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause context of one or more first software commands to be communicated to one or more second software commands and / or to otherwise perform operations described herein.In at least one embodiment, one or more processors (e.g., such as those described herein) include one or more circuits for performing operations described herein in connection with FIG. 4, and / or instructions, such as one or more circuits for performing an application programming interface (API), to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software programs, and / or to otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIG. 4 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software instructions and / or to otherwise execute operations described herein. In at least one embodiment, components described herein in connection with FIG. 4 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until one or more second instructions obtain a context corresponding to one or more first software programs and / or otherwise execute operations described herein. In at least one embodiment, not shown in FIG. 4, one or more components described herein in connection with FIG. 4 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to cause one or more second commands to wait to execute until the one or more second commands receive a context corresponding to one or more first software commands and / or otherwise execute operations described herein.FIG. 5 is a block diagram 500 illustrating an application programming interface (API) to create a subcontext in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a sub-context create API 502 to create a sub-context of a primary context using resources of primary context. In at least one embodiment, not shown in FIG. 5, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a sub-context create API 502, to execute an application programming interface (API), to assign one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use for execution of one or more software threads. In at least one embodiment, not shown in FIG. 5, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a sub-context build API 502 to generate an application programming interface (API) to generate one or more masks that indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment, also not shown in FIG. 5, one or more circuits of a processor, such as those described herein, executes one or more instructions to execute a sub-context create API 502, to execute an application programming interface (API), to assign one or more data structures to indicate which one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads in response to receipt of a second API, such as those described herein.In at least one embodiment, sub-context create API 502, when invoked, receives one or more arguments that indicate information about operations to be performed using techniques such as those described herein. In at least one embodiment, sub-context create API 502, when invoked, receives one or more arguments that indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, sub-context create API 502 receives as input one or more arguments including sub-context return 504. In at least one embodiment, sub-context return 504 is a data value that includes information that can be used to identify, specify, or otherwise specify a storage location to store a created sub-context (e.g., created with sub-context creation API 502). In at least one embodiment, a sub-context return location identified, specified, or otherwise specified by sub-context return 504 is one of a plurality of parameters that may be used by sub-context creation API 502 to create a sub-context. In at least one embodiment, sub-context return 504 is a data value for identifying, specifying, or otherwise specifying an API, such as a create sub-context API 502, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, sub-context create API 502 receives as input one or more arguments including resource descriptor 506. In at least one embodiment, resource descriptor 506 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location of a resource descriptor to create a subcontext (e.g., created with subcontext create API 502). In at least one embodiment, a resource descriptor identified, specified, or otherwise specified by resource descriptor API 506 is one of a plurality of parameters that can be used by sub-context create API 502 to create a sub-context. In at least one embodiment, resource descriptor 506 is a data value for identifying, specifying, or otherwise specifying a set of operations or instructions for an API, such as sub-context create API 502, to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, sub-context create API 502 receives as input one or more arguments that include a device 508. In at least one embodiment, device 508 is a data value that includes information that can be used to identify, specify, or otherwise specify a device that is associated with a subcontext (e.g., created using subcontext create API 502). In at least one embodiment, a device identified, specified, or otherwise specified by device 508 is one of a plurality of parameters that may be used by sub-context creation API 502 to create a sub-context. In at least one embodiment, device 508 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, such as described herein, to an API, such as sub-context creation API 502.In at least one embodiment, sub-context create API 502 receives as input one or more arguments including flags 510. In at least one embodiment, flags 510 is a data value that includes information used to identify, indicate, or otherwise specify flags when creating a sub-context (e.g., using sub-context creation API 502). In at least one embodiment, flag 510 is a data value that indicates which SMs of a device (e.g., indicated by device 508) are to be used for a subcontext created by subcontext creation API 502. In at least one embodiment, flags identified, indicated, or otherwise specified by flags 510 are one of a plurality of parameters that may be used by sub-context create API 502 to create a sub-context. In at least one embodiment, flags 510 is a data value that identifies, indicates, or otherwise specifies an API, such as sub-context create API 502, a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, sub-context create API 502 receives as input one or more arguments including one or more other arguments 518. In at least one embodiment, other arguments 518 are data that includes information to indicate other information that may be used to create a sub-context when executing sub-context create API 502.In at least one embodiment not shown in FIG. 5, a processor executes one or more instructions to execute one or more APIs, such as sub-context create API 502 to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to execute one or more software threads using one or more arguments including, but not limited to, sub-context return 504, resource descriptors 506, device 508, flags 510, and / or other arguments 518. In at least one embodiment, not shown in FIG. 5, a processor executes one or more instructions to execute one or more APIs, such as sub-context create API 502 to execute an application programming interface (API) to generate one or more masks to display one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels, using one or more arguments including, but not limited to, sub-context return 504, resource descriptor 506, device 508, flags 510, and / or other arguments 518.In at least one embodiment, sub-context create API 502, when invoked, causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, sub-context creation API 502, when invoked, causes one or more APIs, such as one or more APIs 306, to add, insert, or otherwise include into a stream or set of instructions to be executed by one or more accelerators within a heterogeneous processor, in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, in response to sub-context create API 502, one or more APIs 306 are to cause one or more processors to perform sub-context create API return 520, if performed. In at least one embodiment, sub-context create API return 520 is an instruction set that, when executed, generates and / or indicates one or more data values in response to sub-context create API 502. In at least one embodiment, sub-context create API return 520 indicates a success indication 522. In at least one embodiment, success indication 522 is data that includes any value to indicate success of sub-context creation API 502. In at least one embodiment, success indicator 522 includes information indicating one or more specific types of incidents generated as a result of execution of sub-context create API 502. In at least one embodiment, success indicator 522 includes information indicating one or more other data values generated as a result of sub-context create API 502.In at least one embodiment, sub-context create API return 520 indicates an error indication 524. In at least one embodiment, fault indicator 524 is data comprising any value indicative of failure of sub-context create API 502. In at least one embodiment, fault indicator 524 includes information indicative of one or more specific types of faults generated as a result of execution of sub-context create API 502. In at least one embodiment, error indicator 524 includes indicia indicating one or more other data values generated as a result of sub-context create API 502.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, sub-context create API 502 that adds various operations of various types to a data stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a semaphore operation. In at least one embodiment, stream operations include a release semaphore operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memories, such as L2 cache memories of a PPU, such as a GPU, and / or cache memories of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more operations to indicate communication of an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, example software code indicating types of stream operations is as follows: / *** types of stream operations ** / typedefenum / **< obtaining semaphores * / CUSOCKET_STREAM_OP_SEMA_ACQ / **< releasing semaphore * / CUSOCKET_STREAM_OP_SEMA_REL / **<PU L2 cache empty * / CUSOCKET_STREAM OP GPU _ L2 FLUSH, / **< GPU-L2 cache invalidate * / CUSOCKET_STREAM_OP_GPU_L2_INVALIDATE, / **<Transmission of an operation to an external device * / CUSOCKET_STREAM_OP_EXTERNAL_DEVICE_SUBMIT } cuJetStreamOpType;In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, sub-context build API 502, one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be performed. In at least one embodiment, example software code indicating a function signature for a callback function looks as follows: / *** Funktions signature of callback function for delivery to an external device. * / typedef unsigned int (*cuSocketExternalDeviceSubmitCallback)(void *submitArgs);In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by sub-context create API 502 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, example software code indicating a data structure representing a device node for one or more accelerators within heterogeneous processors is: / ***structure representing external device node that submits information * about a particular task to an external device. * / typedef structure void *submitArgs; cuSocketExternalDeviceSubmitCallback callback; } cuSocketExternesGerätKnotenParams;In at least one embodiment, to indicate type and data of one or more operations to be performed by one or more accelerators in heterogeneous processors, one or more data structures of one or more APIs 306 are to be used. In at least one embodiment, example software code indicating a data structure to specify type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors looks as follows: / *** structure that tracks type and data for stream operations. The data is filled with * with semaphore address and payload for the types *::CUSOCKET_STREAM_OP_SEMA_ACQ and *::CUSOCKET_STREAM_OP_SEMA_REL * / typedef struct / ** * Type of stream operation * / cuSockeStreamOpType; union { / ** * Parameters for semaphores * / struct { / ** * Address of the semaphores to be acquired or released. * / void *SeAdd; / ** * Payload value of the semaphores. * / unsigned int payload; } sema; / ** * The special task to be transmitted to the external device. * / cuSocketExternalDeviceNodeParams task; } data; } cuJetStreamOp;In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions executed by one or more accelerators in heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to execute in response to sub-context create API 502, as described above. In at least one embodiment, example software code indicating a stream operation API call in parallel computing environment 308, such as CUDA, looks like: / *** Submits a list of operations to a CUDA stream. * - param[in] userStream - Stream into which operations are submitted. * - param[in] streamOp - List of operations to submit. * - param[in] count - Number of operations to submit. * - Returns CUDA_SUCCESS if successful, otherwise a corresponding error is issued. * / CUres cuJetStreamOps(CUstream userStream, cujetStreamOp*streamOp, unsigned int count, unsigned int flags);In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs, similar to how one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to sub-context create API 502. In at least one embodiment, example software code indicating addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308 looks like: / *** Submit a task for an external device to a CUDA stream. * * - param[in] graphNode - Node newly created. * - param[in] graph - Graph in which that node is to be added. * - param[in] dependency - Dependencies that need to be satisfied, before this node * can be executed. * - param[in] numDe - The number of dependencies. * - param[in] nodePara - The execution parameters of the node. * - returns CUDA_SUCCESS if successful, otherwise a corresponding error is issued. CUres cuSocketAddExternalDeviceNode (CUgraphNode*graphNode, CUgraph graph, CUgraphNode*dependences, unsigned int numDependences, cuSocketExternalDeviceNodeParams* nodePara);FIG. 6 is a block diagram 600 illustrating a process for performing an application programming interface (API) to create a subcontext, in accordance with at least one embodiment. In at least one embodiment, process for performing a sub-context creation API shown in block diagram 600 is a process for performing sub-context creation API 502, described herein at least in connection with FIG. 5. In at least one embodiment, a portion or all of process shown in block diagram 600 for performing an API to create a subcontext (or other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices as described in connection with FIGS. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor such as processor 110 described herein in at least connection with FIG. 1 performs one or more steps of process to execute an API to create a subcontext shown in block diagram 600. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to execute an API to create a subcontext shown in block diagram 600.In at least one embodiment, a processor executing process performs one or more operations to receive or otherwise obtain a sub-context creation API at step 602 of process executing an API to create a sub-context shown in block diagram 600. In at least one embodiment, at step 602, an API for creating a subcontext is received or otherwise provided to a processor such as processor 110 described herein in at least connection with FIG. 1. In at least one embodiment, at step 602, an API for creating subcontext is received or otherwise provided by a processor, such as processor 102 described herein at least in connection with FIG. 1. In at least one embodiment, in step 602, an API for creating a subcontext includes one or more arguments described herein at least in connection with FIG. 5. In at least one embodiment, after step 602, process for performing an API to create a subcontext as shown in block diagram 600 continues at step 604.In at least one embodiment, a processor performing process performs one or more operations to determine whether a sub-context creation API (e.g., received in step 602) is valid in step 604 of process shown in block diagram 600 to perform a sub-context creation API. In at least one embodiment, operations to determine whether a subcontext creation API is valid include operations to verify validity of arguments of subcontext creation API (e.g., arguments described herein at least in connection with FIG. 5 ) at step 604. In at least one embodiment, if it is determined that a subcontext creation API is valid ("YES" branch), then the process shown in block diagram 600 for performing a subcontext creation API continues at step 606. In at least one embodiment, if it is determined that a subcontext creation API is not valid ("NO" branch), then in step 604, process for performing a subcontext creation API shown in block diagram 600 continues in step 620, which is described below.In at least one embodiment, in step 606 of process shown in block diagram 600 for performing an API to generate a sub-context, a processor performing process performs one or more operations to obtain resources from a device. In at least one embodiment, in step 606, one or more operations to retrieve resources from a device include one or more operations to query a context (e.g., a primary context) of a device such as device 202 described herein at least in connection with FIG. 2 to determine a set of resources available to or otherwise associated with device as described herein at least in connection with FIG. 2. In at least one embodiment, an apparatus is an argument of an API to create a subcontext, as described herein at least in connection with FIG. 5. In at least one embodiment, after step 606, process shown in block diagram 600 for performing an API to create a sub-context continues at step 608.In at least one embodiment, in step 608 of process for performing a sub-context creation API, as shown in block diagram 600, a processor performing process performs one or more operations to determine whether a sub-context may be created based at least in part on a list of resources and / or one or more flags (e.g., received as arguments for a sub-context creation API, as described herein at least in connection with FIG. 5 ). In at least one embodiment, in one or more operations to determine whether a subcontext can be created, this determination is made based at least in part on available resources, a source context (e.g., from which a subcontext is to be created), a device on which context is to be created, etc. In at least one embodiment, after step 608, process for performing an API to create a subcontext shown in block diagram 600 continues at step 6104.In at least one embodiment, a processor performing process performs one or more operations to determine whether a subcontext can be created (e.g., based on a determination in step 608) in step 610 of process to perform an API to create a subcontext shown in block diagram 600. In at least one embodiment, if it is determined in step 604 that a subcontext can be created ("YES" branch), then process for performing an API to create a subcontext shown in block diagram 600 continues in step 612. In at least one embodiment, if it is determined in step 610 that a subcontext cannot be created ("NO" branch), then process for performing an API to create a subcontext shown in block diagram 600 continues in step 620, described below.In at least one embodiment, a processor performing this process performs one or more operations to generate a subcontext based on resources and flags received as arguments for a subcontext generation API in step 612 of subcontext generation API shown in block diagram 600. In at least one embodiment, a subcontext is created in step 612, as described herein in at least conjunction with FIGS. 1 and 2. In at least one embodiment, after step 612, process shown in block diagram 600 for performing an API to create a sub-context continues at step 614.In at least one embodiment, in step 614 of process shown in block diagram 600 for performing a sub-context creation API, a processor performing process performs one or more operations to assign a subset of a set of resources (e.g., a primary context and / or a sub-context) to a created sub-context. In at least one embodiment, in step 614, a subset of a set of resources is assigned to only one subcontext (e.g., a single context). In at least one embodiment, a subset of a set of resources is assigned to a subcontext by, for example, providing a list of subset to subcontext. In at least one embodiment, after step 614, process shown in block diagram 600 for performing an API to create a sub-context continues at step 616.In at least one embodiment, in step 616 of process for performing an API to create a sub-context depicted in block diagram 600, a processor performing process performs one or more operations to indicate a success indication (e.g., a success indication 522 described herein at least in connection with FIG. 5 ). In at least one embodiment, at step 616, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 616, process for performing an API to create a subcontext shown in block diagram 600 is terminated. In at least one embodiment, not shown in FIG. 6, after step 616, the process for performing an API to create a subcontext shown in block diagram 600 continues in step 602 described above.In at least one embodiment, a processor performing process performs one or more operations to indicate an error indication (e.g., an error indication 524, described herein at least in connection with FIG. 5 ) in step 618 of process to perform an API to create a subcontext shown in block diagram 600. In at least one embodiment, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ) in step 618. In at least one embodiment, after step 618, process for performing an API to create a subcontext shown in block diagram 600 is terminated. In at least one embodiment, not shown in FIG. 6, after step 618, process for performing an API to create a subcontext shown in block diagram 600 continues with step 602 described above.In at least one embodiment, operations of process for performing an API to create a subcontext shown in block diagram 600 are performed in a different order than shown in FIG. 600. In at least one embodiment, operations of process for performing an API to create a sub-context shown in block diagram 600 are performed simultaneously or in parallel. In at least one embodiment, operations of process for performing an API to create a subcontext shown in block diagram 600 that are not dependent on each other (e.g., are independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to create a subcontext shown in block diagram 600 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIGS. 5 and 6, and / or instructions, such as one or more circuits for performing an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIGS. 5 and 6, and / or instructions, such as one or more circuits for performing an application programming interface (API) to generate one or more masks to generate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs), which may be used to perform one or more kernels and / or otherwise to perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIGS. 5 and 6 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to perform an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIGS. 5 and 6 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to associate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or operations otherwise described herein. In at least one embodiment, not shown in FIGS. 5 and 6, one or more components described herein in connection with FIGS. 5 and 6 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to assign one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or otherwise perform operations described herein.FIG. 7 is a block diagram 700 illustrating an application programming interface (API) for destroying a subcontext in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a subcontext destroy API 702 to destroy a subcontext created by executing a subcontext create API, such as subcontext create API 502 described herein in connection with at least FIG. 5. In at least one embodiment, not shown in FIG. 7, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform sub-context destroy API 702 to execute an application programming interface (API) to release one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads. In at least one embodiment, not shown in FIG. 7, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a sub-context destroy API 702 to execute an application programming interface (API) to cause one or more masks to be disabled, wherein one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment, also not shown in FIG. 7, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform sub-context destroying API 702 to execute an application programming interface (API) to release one or more data structures to indicate which one or more streaming multiprocessors (SMs) of one or more processors to execute one or more software threads in response to receipt of a second API, such as those described herein.In at least one embodiment, subcontext destroy API 702, when invoked, receives one or more arguments indicative of information about operations to be performed using techniques such as those described herein. In at least one embodiment, sub-context destroy API 702, when invoked, receives one or more arguments that indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, sub-context destroy API 702 receives as input one or more arguments including sub-context 704. In at least one embodiment, a subcontext 704 is a data value that includes information that is used to identify, specify, or otherwise specify a subcontext to be destroyed with subcontext destroying API 702. In at least one embodiment, a subcontext identified, specified, or otherwise specified by subcontext 704 is one of a plurality of parameters that can be used by subcontext destroying API 702 to destroy a subcontext. In at least one embodiment, sub-context 704 is a data value that identifies, indicates, or otherwise specifies to an API, such as sub-context destroy API 702, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, subcontext destroy API 702 receives as input one or more arguments including one or more other arguments 718. In at least one embodiment, other arguments 718 are data that includes information to indicate other information that may be used to destroy a sub-context when executing sub-context destroying API 702.In at least one embodiment, not shown in FIG. 7, a processor executes one or more instructions to execute one or more APIs, such as Destroy SubKontext API 702, to execute an Application Programming Interface (API) to release one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads, using one or more arguments including, but not limited to, SubKontext 704 and / or other arguments 718. In at least one embodiment, not shown in FIG. 7, a processor executes one or more instructions to execute one or more APIs, such as subcontext destroy API 702, to execute an application programming interface (API) to cause one or more masks to be disabled, wherein one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more arguments including, but not limited to, subcontext 704 and / or other arguments 718.In at least one embodiment, sub-context destroy API 702, when invoked, causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, sub-context destroying API 702, when invoked, causes one or more APIs, such as one or more APIs 306, to add, insert, or otherwise include into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor, in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, in response to sub-context destroy API 702, one or more APIs 306 are to cause one or more processors to perform sub-context destroy API return 720, if performed. In at least one embodiment, sub-context destroy API return 720 is an instruction set that, when executed, generates and / or specifies one or more data values in response to sub-context destroy API 702. In at least one embodiment, sub-context destroy API return 720 indicates a success indication 722. In at least one embodiment, success indication 722 is data comprising any value to indicate success of sub-context destroy API 702. In at least one embodiment, success indicator 722 includes information indicating one or more specific types of incidents generated as a result of execution of sub-context destroy API 702. In at least one embodiment, success indicator 722 includes indicia indicating one or more other data values generated as a result of sub-context destroy API 702.In at least one embodiment, a fault indicator 724 is indicated at sub-context destroy API return 720. In at least one embodiment, fault indicator 724 is data comprising any value indicative of failure of sub-context destroy API 702. In at least one embodiment, fault indicator 724 includes information indicative of one or more specific types of faults generated as a result of performance of sub-context destroy API 702. In at least one embodiment, fault indicator 724 includes information indicative of one or more other data values generated as a result of sub-context destroy API 702.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, sub-context destroy API 702 that adds stream operations of various types that are executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a semaphore capture operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache memory of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more indications for communicating an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicative of communicating an operation to an external device use software code, such as example software code indicative of stream operations, as described herein at least in connection with FIG. 5.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, subcontext destroy API 702, one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be performed. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code, that indicates a function signature for a callback function, as described herein at least in connection with FIG. 5.In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by sub-context destroy API 702 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices to which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, one or more data structures of one or more APIs 306 usable to indicate one or more external devices for which one or more APIs 306 are to transmit one or more operations use software code, such as example software code, indicating a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code indicating a data structure to indicate type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions are to be executed in response to sub-context destroy API 702, as described above. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to sub-context destroy API 702 use software code, such as example software code, that indicates a stream operation API call in parallel computing environment 308, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are comparable to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to sub-context destroy API 702 as described herein. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code, that indicates addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308, as described herein at least in connection with FIG. 5.FIG. 8 is a block diagram 800 illustrating a process for performing an application programming interface (API) to destroy a sub-context, in accordance with at least one embodiment. In at least one embodiment, process for performing a sub-context destruction API shown in block diagram 800 is a process for performing sub-context destruction API 702 described herein at least in connection with FIG. 7. In at least one embodiment, a portion or all of process shown in block diagram 800 for performing an API to destroy a subcontext (or other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices as described in connection with FIGS. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor such as processor 110 described herein in at least connection with FIG. 1 performs one or more steps of process to execute an API to destroy a subcontext shown in block diagram 800. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to perform an API to destroy a subcontext shown in block diagram 800.In at least one embodiment, a processor performing process performs one or more operations to receive or otherwise obtain a sub-context destroying API, at step 802 of process performing an API shown in block diagram 800. In at least one embodiment, in step 802, an API for destroying a subcontext is received or otherwise provided to a processor such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, in step 802, an API to destroy a subcontext is received from a processor, such as processor 102, described herein at least in connection with FIG. 1. In at least one embodiment, an API for destroying a subcontext in step 802 includes one or more arguments as described herein at least in connection with FIG. 7. In at least one embodiment, after step 802, process for performing an API to destroy a sub-context shown in block diagram 800 continues at step 804.In at least one embodiment, a processor performing process performs one or more operations to determine whether a sub-context destruction API (e.g., received in step 802) is valid in step 804 of process performing a sub-context destruction API shown in block diagram 800. In at least one embodiment, operations to determine whether a subcontext destruction API is valid include operations to check validity of arguments of subcontext destruction API (e.g., arguments described herein at least in connection with FIG. 7 ) at step 804. In at least one embodiment, if a subcontext destruction API is determined to be valid ("YES" branch), then in step 804, process for performing a subcontext destruction API as shown in block diagram 800 continues in step 806. In at least one embodiment, if it is determined that a subcontext destruction API is not valid ("NO" branch), then in step 804, process for performing a subcontext destruction API shown in block diagram 800 continues in step 816, which is described below.In at least one embodiment, in step 806 of process for performing an API to destroy a subcon context shown in block diagram 800, a processor performing process performs one or more operations to identify a subcon context to be destroyed. In at least one embodiment, a subcontext is identified as an argument of a subcontext destroying API received in step 802. In at least one embodiment, after step 806, process for executing an API to destroy a subcontext shown in block diagram 800 continues at step 808.In at least one embodiment, in step 808 of process for performing an API to destroy a subcontext shown in block diagram 800, a processor performing process performs one or more operations to determine whether a subcontext to be destroyed has been identified (e.g., identified in step 806). In at least one embodiment, if it is determined in step 808 that a subcontext to be destroyed has been identified ("YES" branch), then process of performing an API to destroy a subcontext shown in block diagram 800 continues in step 810. In at least one embodiment, if it is determined at step 808 that a subcontext to be destroyed has not been identified ("NO" branch), then process of performing an API to destroy a subcontext shown in block diagram 800 continues at step 816, which is described below.In at least one embodiment, in step 810 of process for performing an API to destroy a subcontext shown in block diagram 800, a processor performing process performs one or more operations to stop one or more streams associated with a context to be destroyed (e.g., a subcontext identified in step 806), release all resources associated with context to be destroyed, and destroy context to be destroyed (e.g., indicate that context is not valid). In at least one embodiment, data streams are allowed to complete ongoing operations before being halted. In at least one embodiment, data streams are immediately stalled (e.g., they must not complete ongoing operations). In at least one embodiment, a plurality of data streams associated with a subcontext are stalled. In at least one embodiment, after step 810, process for performing an API to destroy a sub-context shown in block diagram 800 continues at step 812.In at least one embodiment, a processor performing process performs one or more operations to determine whether a subcontext to be destroyed has been destroyed (e.g., by performing step 710) at step 812 of process performing an API to destroy a subcontext shown in block diagram 800. In at least one embodiment, if it is determined in step 812 that a context to be destroyed has been destroyed ("YES" branch), then process of performing an API to destroy a sub-context shown in block diagram 800 continues in step 814. In at least one embodiment, if it is determined in step 812 that a context to be destroyed has not been destroyed ("NO" branch), then process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 816 described below.In at least one embodiment, a processor performing process performs one or more operations to indicate a success indication (e.g., a success indication 722, described herein at least in connection with FIG. 7 ) at step 814 of process performing an API to destroy a subcontext shown in block diagram 800. In at least one embodiment, at step 814, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 814, process for performing an API to destroy a subcontext shown in block diagram 800 is terminated. In at least one embodiment, not shown in FIG. 8, after step 814, process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 802 described above.In at least one embodiment, a processor performing process performs one or more operations to return an error indication (e.g., an error indication 724 described herein at least in connection with FIG. 7 ) in step 816 of process to perform an API to destroy a subcontext indicated in block diagram 800. In at least one embodiment, at step 816, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 816, process for performing an API to destroy a subcontext shown in block diagram 800 is terminated. In at least one embodiment, not shown in FIG. 8, after step 816, the process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 802 described above.In at least one embodiment, operations of process for performing an API to destroy a subcontext shown in block diagram 800 are performed in a different order than shown in FIG. 800. In at least one embodiment, operations of process for performing an API to destroy a sub-context shown in block diagram 800 are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to destroy a subcontext shown in block diagram 800 that are not dependent on each other (e.g., are independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process for performing an API to destroy a subcontext shown in block diagram 800 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform operations and / or instructions described herein in connection with FIGS. 7 and 8, such as one or more circuits to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads and / or operations otherwise described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to perform operations and / or instructions described herein in connection with FIGS. 7 and 8, such as one or more circuits to execute an application programming interface (API) to cause one or more masks to be deactivated, where one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for executing one or more kernels and / or for otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIGS. 7 and 8 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein. In at least one embodiment, components described herein in connection with FIGS. 7 and 8 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or operations otherwise described herein. In at least one embodiment, not shown in FIGS. 7 and 8, one or more components described herein in connection with FIGS. 7 and 8 include one or more components described herein in connection with FIGS. 26-58 to execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads and / or other operations described herein.FIG. 9 is a block diagram 900 illustrating an application programming interface (API) to obtain a sub-context of stream obtaining API, in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a subcontext of stream obtaining API 902 to obtain a subcontext associated with an identified stream executing using a channel as described herein at least in connection with FIGS. 1 and 2. In at least one embodiment, not shown in FIG. 9, one or more circuits of a processor, such as those described herein, executes one or more instructions to obtain a sub-context of stream API 902 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicative of a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored. In at least one embodiment, not shown in FIG. 9, one or more circuitry of a processor, such as those described herein, executes one or more instructions to execute sub-context-of-stream-obtaining API 902 to execute an application programming interface (API) to indicate one or more identifiers of one or more masks, where one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment, also not shown in FIG. 9, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute sub-context-of-stream-obtaining API 902 to cause an application programming interface (API) to store one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors in response to receipt of a second API, such as those described herein.In at least one embodiment, sub-context-of-stream-obtaining API 902, when invoked, receives one or more arguments that indicate information about operations to be performed using techniques such as those described herein. In at least one embodiment, subcontext from stream obtain API 902, when invoked, receives one or more arguments that indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, subcontext of stream obtaining API 902 receives as input one or more arguments including stream 904. In at least one embodiment, stream 904 is a data value that includes information that can be used to identify, specify, or otherwise specify a stream from which to retrieve a subcontext using subcontext of stream obtain API 902. In at least one embodiment, a stream identified, specified, or otherwise specified by stream 904 is one of a plurality of parameters that can be used by subcontext-of-stream-obtaining API 902 to obtain a subcontext of a stream. In at least one embodiment, stream 904 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, such as described herein, to an API, such as subcontext of stream obtaining API 902.In at least one embodiment, subcontext-of-stream-obtaining API 902 receives as input one or more arguments including subcontext return 906. In at least one embodiment, sub-context return 906 is a data value comprising information that can be used to identify, indicate, or otherwise specify a storage location to store a sub-context identified with sub-context of stream obtaining API 902. In at least one embodiment, a storage location for storing a subcontext identified, specified, or otherwise specified by subcontext return 906 is one of a plurality of parameters that can be used by subcontext of stream obtaining API 902 to obtain a subcontext from a stream. In at least one embodiment, sub-context return 906 is a data value for identifying, specifying, or otherwise specifying an API, such as sub-context from stream obtaining API 902, a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.In at least one embodiment, subcontext from stream obtain API 902 receives as input one or more arguments including one or more other arguments 918. In at least one embodiment, other arguments 918 are data that includes information to indicate other information that may be used in performing subcontext of stream obtaining API 902 to obtain a subcontext of a stream.In at least one embodiment, not shown in FIG. 9, a processor executes one or more instructions to execute one or more APIs, such as subcontext of stream obtaining API 902, to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored using one or more arguments including, but not limited to, stream 904, subcontext return 906, and / or other arguments 918. In at least one embodiment, not shown in FIG. 9, a processor executes one or more instructions to execute one or more APIs, such as subcontext of stream obtaining API 902, to execute an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more arguments including, but not limited to, stream 904, subcontext return 906, and / or other arguments 918.In at least one embodiment, sub-context-of-stream-obtaining API 902, when invoked, causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, sub-context-of-stream-obtaining API 902, when invoked, causes one or more APIs, such as one or more APIs 306, to add, insert, or otherwise include into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor, in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, one or more APIs 306, when executed, are to cause one or more processors to perform a sub-context-of-stream-obtained API return 920 in response to sub-context-of-stream-obtained API 902. In at least one embodiment, sub-context of stream obtaining API 920 is an instruction set that, when executed, generates and / or indicates one or more data values in response to sub-context of stream API 902. In at least one embodiment, sub-context-of-stream-obtaining API return 920 indicates success indication 922. In at least one embodiment, success indication 922 is data comprising any value indicative of success of sub-context-of-stream-obtaining API 902. In at least one embodiment, success indication 922 includes information indicating one or more specific types of success generated as a result of performance of sub-context-of-stream-obtaining API 902. In at least one embodiment, success indicator 922 includes information indicating one or more other data values generated as a result of sub-context-of-stream-obtaining API 902.In at least one embodiment, sub-context-of-stream-obtaining API return 920 indicates an error indicator 924. In at least one embodiment, fault indicator 924 is data comprising any value that indicates failure to query sub-context-of-stream-obtaining API 902. In at least one embodiment, error indicator 924 includes information indicative of one or more specific types of errors generated as a result of performance of sub-context-of-stream-obtaining API 902. In at least one embodiment, error indicator 924 includes information indicating one or more other data values generated as a result of sub-context-of-stream-obtaining API 902.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, sub-context-of-stream-obtaining API 902 that add to a stream various operations of various types that are performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a semaphore capture operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache memory of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more indications for communicating an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating transmission of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with FIG. 5.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, sub-context-of-stream-obtaining API 902, one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause execution of one or more callback functions. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code, that indicates a function signature for a callback function, as described herein at least in connection with FIG. 5.In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by subcontext of stream-getting API 902 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, one or more data structures of one or more APIs 306 usable to indicate one or more external devices for which one or more APIs 306 are to transmit one or more operations use software code, such as example software code, indicating a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code indicating a data structure to indicate type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions executed by one or more accelerators in heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions are to be executed in response to receiving sub-context-of-stream-receiving API 902, as described above. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions to be executed in response to sub-context stream obtaining API 902 use software code, such as example software code, that indicates a stream operation API call in a parallel computing environment 308, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are similar to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to sub-context of stream obtaining API 902 as described herein. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code, that indicates addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308, as described herein at least in connection with FIG. 5.FIG. 10 is a block diagram 1000 illustrating a process for performing an application programming interface (API) to obtain subcontext of stream API, in accordance with at least one embodiment. In at least one embodiment, process shown in block diagram 1000 for performing an API to obtain subcontext from a stream is a process to obtain subcontext from stream API 902, described herein at least in connection with FIG. 9. In at least one embodiment, a portion or all of a process for performing an API to obtain subcontext from a stream shown in block diagram 1000 (or any other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with FIGS. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor such as processor 110 described herein in at least connection with FIG. 1 performs one or more steps of process to perform an API for obtaining subcontext from a stream shown in block diagram 1000. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to execute an API for obtaining subcontext of stream API from a stream shown in block diagram 1000.In at least one embodiment, in step 1002 of process for performing an API to obtain subcontext from a stream shown in block diagram 1000, a processor performing process performs one or more operations to receive or otherwise obtain an API to obtain subcontext from a stream. In at least one embodiment, at step 1002, an API for obtaining subcontext from stream API is received or otherwise provided to a processor such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, in step 1002, an API for obtaining subcontext of stream API is received from a processor, such as processor 102, described herein at least in connection with FIG. 1. In at least one embodiment, an API for obtaining subcontext from stream API in step 1002 includes one or more arguments as described herein at least in connection with FIG. 9. In at least one embodiment, after step 1002, process for executing an API to obtain subcontext from stream API as shown in block diagram 1000 continues at step 1004.In at least one embodiment, in step 1004 of process shown in block diagram 1000 for executing a subcon context obtaining API of stream API, a processor performing process performs one or more operations to determine whether a subcon context obtaining API of stream API (e.g., received in step 1002) is valid. In at least one embodiment, operations to determine whether a stream API subcontext obtaining API is valid include, at step 1004, operations to check validity of arguments of stream API subcontext obtaining API (e.g., arguments described herein at least in connection with FIG. 9 ). In at least one embodiment, if it is determined in step 1004 that an API to obtain subcontext from a stream is valid ("YES" branch), then process for performing an API to obtain subcontext from a stream as shown in block diagram 1000 continues in step 1006. In at least one embodiment, if it is determined in step 1004 that an API to obtain subcontext from a stream is not valid ("NO" branch), then process for performing an API to obtain subcontext from a stream shown in block diagram 1000 continues in step 1014 described below.In at least one embodiment, in step 1006 of process for performing an API to obtain subcontext from a stream shown in block diagram 1000, a processor performing process performs one or more operations to identify a subcontext from a stream (e.g., a stream indicated by an argument for an API to obtain subcontext from a stream). In at least one embodiment, a subcontext maintains a list of streams associated with said context. In at least one embodiment, each stream maintains a list of subcontexts associated with that stream. In at least one embodiment, after step 1006, process for performing an API to obtain subcontext from stream APl as shown in block diagram 1000 continues at step 1008.In at least one embodiment, a processor performing process performs one or more operations to determine whether a subcontext has been identified (e.g., in step 1006) in step 1008 of process for performing an API to obtain subcontext from a stream shown in block diagram 1000. In at least one embodiment, if it is determined in step 1008 that a subcontext has been identified ("YES" branch), then process for performing an API to obtain subcontext from a stream shown in block diagram 1000 continues in step 1010. In at least one embodiment, if it is determined in step 1008 that a subcontext has not been identified ("NO" branch), then process for performing an API to obtain subcontext from a stream shown in block diagram 1000 continues in step 1014, which is described below.In at least one embodiment, a processor performing process performs one or more operations to store an identified subcontext (e.g., identified in step 1006) at a storage location indicated by an API for obtaining subcontext from a stream (e.g., received as an argument for API for obtaining subcontext from a stream, received in step 1002) in step 1010 of process for performing an API for obtaining subcontext from a stream. In at least one embodiment, after step 1010, process for performing an API to obtain subcontext from stream APl as shown in block diagram 1000 continues at step 1012.In at least one embodiment, in step 1012 of process for performing an API to obtain subcontext from a stream shown in block diagram 1000, a processor performing process performs one or more operations to return a success indication (e.g., a success indication 922, described herein at least in connection with FIG. 9 ). In at least one embodiment, at step 1012, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 1012, process for performing an API to obtain subcontext from a stream APl is ended as shown in block diagram 1000. In at least one embodiment, not shown in FIG. 6, after step 1012, process for performing an API to obtain subcontext from a stream shown in block diagram 1000 continues at step 1002 described above.In at least one embodiment, a processor performing process performs one or more operations to indicate an error indication (e.g., an error indication 924, described herein at least in connection with FIG. 9 ) at step 1014 of process for performing an API to obtain subcontext from a stream shown in block diagram 1000. In at least one embodiment, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ) in step 1014. In at least one embodiment, after step 1014, process for performing an API to obtain subcontext from a stream API is ended, as shown in block diagram 1000. In at least one embodiment, not shown in FIG. 10, after step 1014, process for performing an API to obtain subcontext from a stream shown in block diagram 1000 continues at step 1002 described above.In at least one embodiment, operations of process for executing an API for obtaining subcontext from a stream shown in block diagram 1000 are performed in a different order than shown in FIG. 1000. In at least one embodiment, operations of process for performing an API to obtain subcontext of stream API from a stream shown in block diagram 1000 are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to obtain subcontext from a stream shown in block diagram 1000 that are not dependent on each other (e.g., are independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to obtain subcontext from a stream shown in block diagram 1000 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations and / or instructions described herein in connection with FIGS. 9 and 10, such as one or more circuits for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise performing operations described herein. In at least one embodiment, one or more processors (e.g., one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to perform operations and / or instructions described herein in connection with FIGS. 9 and 10, such as one or more circuits to execute an application programming interface (API) to indicate one or more identifiers of one or more masks, where one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable to perform one or more kernels and / or otherwise perform operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIGS. 9 and 10 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise operations described herein to be performed. In at least one embodiment, components described herein in connection with FIGS. 9 and 10 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIGS. 9 and 10, one or more components described herein in connection with FIGS. 9 and 10 include one or more components described herein in connection with FIGS. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein.FIG. 11 is a block diagram 1100 illustrating an application programming interface (API) to obtain resources associated with a context, in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a context event obtain API 1102 to retrieve resources associated with a context. In at least one embodiment, not shown in FIG. 11, one or more circuits of a processor, such as those described herein, executes one or more instructions to perform an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for execution of one or more software threads. In at least one embodiment, not shown in FIG. 11, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform an application programming interface (API) 1102 to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for execution of one or more software kernels. In at least one embodiment, also not shown in FIG. 11, one or more circuits of a processor, such as those described herein, executes one or more instructions to execute an application programming interface (API) 1102 to indicate one or more streaming multiprocessors (SMs) of one or more processors available to be used to execute one or more software threads in response to receipt of a second API, such as those described herein.In at least one embodiment, upon its invocation, obtain context resource APl 1102 receives one or more arguments indicating information about operations to be performed using techniques such as those described herein. In at least one embodiment, get context resource API 1102, when invoked, receives one or more arguments that indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, obtain context resource API 1102 receives as input one or more arguments including resource return 1104. In at least one embodiment, resource return 1104 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location for storing resources (e.g., a resource descriptor) associated with a context that uses obtain context resource API 1102. In at least one embodiment, a storage location for storing resources identified, specified, or otherwise specified by resource return 1104 is one of a plurality of parameters that can be used by obtain context resource API 1102 to obtain context-associated resources. In at least one embodiment, resource return 1104 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to an API, such as obtain context resource API 1102, that are to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, get context resource APl 1102 receives as input one or more arguments including context 1106. In at least one embodiment, context 1106 is a data value that includes information that can be used to identify, specify, or otherwise specify a context or subcontext from which a set of resources can be identified and returned using obtain context resource API 1102. In at least one embodiment, a context or subcontext identified, specified, or otherwise specified by context 1106 is one of a plurality of parameters that can be used by obtain context resource API 1102 to obtain resources associated with a context. In at least one embodiment, context 1106 is a data value that identifies, indicates, or otherwise specifies to an API, such as obtain context resource API 1102, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, get context resource APl 1102 receives as input one or more arguments including one or more other arguments 1118. In at least one embodiment, other arguments 1118 are data including information to indicate other information that may be used in performing Obtain Context Resource API 1102 to obtain resources associated with a context.In at least one embodiment, not shown in FIG. 11, a processor executes one or more instructions to execute one or more APIs, such as obtain context resource API 1102, to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) from one or more processors available to be used to execute one or more software threads, using one or more arguments including, but not limited to, resource return 1104, context 1106, and / or other arguments 1118. In at least one embodiment, not shown in FIG. 11, a processor executes one or more instructions to execute one or more APIs, such as obtain context resource API 1102, to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs), which may be used to execute one or more software kernels using one or more arguments including, but not limited to, resource return 1104, context 1106, and / or other arguments 1118.In at least one embodiment, when invoked, Obtain Context Resource API 1102 causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, Obtain Context Resource API 1102, when invoked, causes one or more APIs, such as one or more APIs 306, to add, insert, or otherwise include into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor, in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, in response to obtain context resource API 1102, one or more APIs 306, if executed, are to cause one or more processors to perform obtain context resource API return 1120. In at least one embodiment, obtain context resource API return 1120 is an instruction set that, when executed, generates and / or indicates one or more data values in response to obtain context resource API 1102. In at least one embodiment, obtain context resource API return 1120 indicates success indication 1122. In at least one embodiment, success indication 1122 is data comprising any value to indicate success of Obtain Context Resource API 1102. In at least one embodiment, success indication 1122 includes information indicating one or more specific types of success generated as a result of performing Get Context Resource API 1102. In at least one embodiment, success indicator 1122 includes information indicating one or more other data values generated as a result of obtain context resource API 1102.In at least one embodiment, obtain context resource API return 1120 indicates an error indication 1124. In at least one embodiment, error indicator 1124 is data comprising any value to indicate an error of obtain context resource API 1102. In at least one embodiment, fault indicator 1124 includes information indicating one or more specific types of faults generated as a result of performing get context resource API 1102. In at least one embodiment, error indicator 1124 includes indicia indicating one or more other data values generated as a result of obtain context resource API 1102.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, obtain context resource API 1102 that adds various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include an Acquire Semaphore operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memories, such as L2 cache memories of a PPU, such as a GPU, and / or cache memories of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more operations to indicate communication of an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating communication of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with FIG. 5.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, obtain context resource API 1102, one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators in heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be performed. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code, that indicates a function signature for a callback function, as described herein at least in connection with FIG. 5.In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by obtain context resource API 1102 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, one or more data structures of one or more APIs 306 usable to indicate one or more external devices for which one or more APIs 306 are to transmit one or more operations use software code, such as example software code, indicating a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code indicating a data structure to indicate type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to be executed in response to obtain context resource API 1102, as described above. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to obtain context resource API 1102 use software code, such as example software code, that indicates a stream operation API call in a parallel computing environment 308, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are comparable to how one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors are added to one or more streams or instruction sets in response to obtain context resource API 1102 as described herein. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code, that indicates addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308, as described herein at least in connection with FIG. 5.FIG. 12 is a block diagram 1200 illustrating a process for performing an application programming interface (API) to obtain context resources in accordance with at least one embodiment. In at least one embodiment, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 is a process for performing obtain context resource API 1102, described herein at least in connection with FIG. 11. In at least one embodiment, a portion or all of process for performing an API for retrieving resources associated with a context depicted in block diagram 1200 (or other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices as described in connection with FIGS. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor, such as processor 110 described herein in at least connection with FIG. 1, executes one or more steps of process to execute an API to obtain resources associated with a context shown in block diagram 1200. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to execute an API to obtain resources associated with a context shown in block diagram 1200.In at least one embodiment, in step 1202 of process for performing an API to retrieve resources associated with a context depicted in block diagram 1200, a processor performing process performs one or more operations to receive or otherwise obtain an API to retrieve resources associated with a context. In at least one embodiment, in step 1202, an API for retrieving resources associated with a context is received or otherwise provided to a processor such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, at step 1202, an API for retrieving context-associated resources is received or otherwise provided by a processor, such as processor 102 described herein at least in connection with FIG. 1. In at least one embodiment, in step 1202, an API for retrieving context resources includes one or more arguments as described herein at least in connection with FIG. 11. In at least one embodiment, after step 1202, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues at step 1204.In at least one embodiment, a processor performing process performs one or more operations to determine whether an API for retrieving resources associated with a context shown in block diagram 1200 is valid, in step 1204 of process to perform an API for retrieving resources associated with a context (e.g., received in step 1202). In at least one embodiment, operations to determine whether a context-associated resource retrieval API is valid include, at step 1204, operations to check validity of arguments of context-associated resource retrieval API (e.g., arguments described herein at least in connection with FIG. 11 ). In at least one embodiment, if it is determined that a context-associated resource retrieval API is valid ("YES" branch), then in step 1204, process for performing a context-associated resource retrieval API as shown in block diagram 1200 continues in step 1206. In at least one embodiment, if it is determined that a context-associated resource retrieval API is not valid ("NO" branch), then in step 1204, process for performing a context-associated resource retrieval API is continued in step 1214, described below.In at least one embodiment, a processor performing process performs one or more operations to determine a request type of an API for retrieving resources associated with a context shown in block diagram 1200 in step 1206 of process to perform an API for retrieving resources associated with a context (e.g., received in step 1202). In at least one embodiment, a request type may be to obtain resources associated with a context. In at least one embodiment, a request type may be obtaining resources associated with a sub-context. In at least one embodiment, a request type may be to retrieve resources associated with a device. In at least one embodiment, a request type is received as an argument for an API to retrieve resources associated with a context (e.g., received in step 1202). In at least one embodiment, after step 1206, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues at step 1208.In at least one embodiment, a processor performing process performs, at step 1208 of process to perform an API to retrieve resources associated with a context shown in block diagram 1200, one or more operations to determine whether a request type (e.g., determined at step 1206) is to retrieve resources associated with a context. In at least one embodiment, if it is determined in step 1208 that a request type is to retrieve resources associated with a context ("YES" branch), then process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues in step 1216. In at least one embodiment, if it is determined in step 1208 that a request type is not intended to obtain resources associated with a context ("NO" branch), process of performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1210.In at least one embodiment, in step 1210 of process for performing an API to obtain resources associated with a context depicted in block diagram 1200, a processor performing process performs one or more operations to determine that a request type (e.g., determined in step 1206) is to obtain resources associated with a subcontext. In at least one embodiment, if it is determined in step 1210 that a request type is to obtain resources associated with a subcontext ("YES" branch), process of performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1222. In at least one embodiment, if it is determined in step 1210 that a request type is not destined to obtain resources associated with a subcontext ("NO" branch), process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1212.In at least one embodiment, a processor performing process performs one or more operations to determine whether a request type (e.g., determined in step 1206) is to retrieve resources associated with a device in step 1212 of process to perform an API to retrieve resources associated with a context shown in block diagram 1200. In at least one embodiment, if it is determined in step 1212 that a request type is to retrieve resources associated with a device ("YES" branch), then process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues in step 1228. In at least one embodiment, if it is determined in step 1212 that if it is determined that if it is determined that a request type is not to obtain resources associated with a device ("NO" branch), then process of performing an API to obtain resources associated with a context depicted in block diagram 1200 continues in step 1214.In at least one embodiment, in step 1214 of process for performing an API to retrieve resources associated with a context depicted in block diagram 1200, a processor performing process performs one or more operations to return an error indicator (e.g., error indicator 1124, described herein at least in connection with FIG. 11 ). In at least one embodiment, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ) in step 1214. In at least one embodiment, after step 1214, process for performing an API for retrieving resources associated with a context shown in block diagram 1200 is ended. In at least one embodiment, not shown in FIG. 12, after step 614, process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1202 described above.In at least one embodiment, a processor performing process performs one or more operations to identify a context (e.g., a context received as an argument for an API to retrieve resources associated with a context received in step 1202) using systems, methods, and operations described herein, in step 1216 of process to perform an API to retrieve resources associated with a context shown in block diagram 1200. In at least one embodiment, after step 1216, process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1218.In at least one embodiment, a processor performing process performs one or more operations to determine whether a context has been identified (e.g., in step 1216) in step 1218 of process to perform an API to retrieve resources associated with a context shown in block diagram 1200. In at least one embodiment, if it is determined in step 1118 that a context has been identified ("YES" branch), process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues in step 1220. In at least one embodiment, if it is determined at step 1204 that a context has not been identified ("NO" branch), process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1214, as described above.In at least one embodiment, a processor performing process performs one or more operations to store an identified context (e.g., identified in step 1216) in step 1220 of process to perform an API to retrieve resources associated with a context shown in block diagram 1200. In at least one embodiment, a processor performs one or more operations to store an identified context at a storage location received as an argument for an API to obtain resources associated with a context (e.g., received at step 1202) at step 1220. In at least one embodiment, after step 1220, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues at step 1230.In at least one embodiment, a processor performing process performs one or more operations to identify a subcontext (e.g., a subcontext received as an argument for an API for obtaining resources associated with a context received in step 1202) in step 1222 of process for performing an API for obtaining resources associated with a context shown in block diagram 1200, as described herein. In at least one embodiment, after step 1222, process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1224.In at least one embodiment, a processor performing process performs one or more operations to determine whether a subcontext has been identified (e.g., in step 1222) in step 1224 of process to perform an API to obtain resources associated with a context depicted in block diagram 1200. In at least one embodiment, if it is determined in step 1224 that a subcontext has been identified ("YES" branch), then process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1226. In at least one embodiment, if it is determined that a subcontext has not been identified ("NO" branch), then in step 1224, process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1214 described above.In at least one embodiment, a processor performing this process performs one or more operations to store an identified subcontext (e.g., identified in step 1222) in step 1226 of process to perform an API to obtain resources associated with a context depicted in block diagram 1200. In at least one embodiment, a processor performs one or more operations to store an identified subcontext at a storage location received as an argument for an API to obtain resources associated with a context (e.g., received in step 1202) in step 1226. In at least one embodiment, after step 1226, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues at step 1230.In at least one embodiment, in step 1228 of process for performing an API to retrieve resources associated with a context shown in block diagram 1200, a processor performing process performs one or more operations to store a device context (e.g., a primary or current context of a device such as device 202 described herein at least in connection with FIG. 2 ). In at least one embodiment, a processor performs one or more operations to store a device context at a storage location received as an argument for an API to obtain resources associated with a context (e.g., received at step 1202) at step 1228. In at least one embodiment, in step 1228, a device context is stored that is a primary, default, or current context of a device. In at least one embodiment, after step 1228, process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues at step 1230.In at least one embodiment, at step 1230 of process to perform an API to obtain resources associated with a context indicated in block diagram 1200, a processor performing process performs one or more operations to return a success indication (e.g., a success indication 1122, described herein at least in connection with FIG. 11). In at least one embodiment, at step 1230, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 1230, process for performing an API for retrieving resources associated with a context shown in block diagram 1200 is ended. In at least one embodiment, not shown in FIG. 12, after step 1230, process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1202 described above.In at least one embodiment, operations of process for performing an API for retrieving context resources shown in block diagram 1200 are performed in a different order than shown in FIG. 1200. In at least one embodiment, operations of process for performing an API for retrieving context resources shown in block diagram 1200 are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to retrieve context resources shown in block diagram 1200 that are not dependent on each other (e.g., are independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API for retrieving context resources shown in block diagram 1200 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIGS. 11 and 12 and / or instructions, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more available processors that may be used to perform one or more software threads and / or to otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIGS. 11 and 12, and / or instructions, such as one or more circuits for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in performing one or more software kernels and / or otherwise for performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIGS. 11 and 12 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 for performing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for execution of one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIGS. 11 and 12 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIGS. 11 and 12, one or more components described herein in connection with FIGS. 11 and 12 include one or more components described herein in connection with FIGS. 26-58. 26-58 to execute an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors that may be used for execution of one or more software threads and / or otherwise perform operations described herein.FIG. 13 is a block diagram 1300 illustrating an application programming interface (API) for partitioning context resources in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor perform a context resource partitioning APl 1302 to partition context resources according to one or more received partitioning criteria. In at least one embodiment, not shown in FIG. 13, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a context resource partitioning APl 1302 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use the corresponding plurality of SMs. In at least one embodiment, not shown in FIG. 13, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a context resource partitioning APl 1302 to execute an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor that may be used for execution of one or more software programs. In at least one embodiment, also not shown in FIG. 13, one or more circuits of a processor, such as those described herein, executes one or more instructions to perform a context resource partitioning APl 1302 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs in response to receipt of a second API, such as those described herein.In at least one embodiment, context resource partitioning API 1302, when invoked, receives one or more arguments that indicate information about operations to be performed using techniques such as those described herein. In at least one embodiment, upon invocation, context resource partitioning API 1302 receives one or more arguments to indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, context resource partitioning API 1302 receives as input one or more arguments including a partitioned resource return 1304. In at least one embodiment, resource return 1304 is a data value that includes information that can be used to identify, specify, or otherwise specify a storage location at which to store a partitioned context resource list as a result of using context resource partitioning API 1302. In at least one embodiment, a divided resource list return identified, indicated, or otherwise specified by resource return 1304 is one of a plurality of parameters that may be used by context resource division APl 1302. In at least one embodiment, API return 1304 of subdivided context resource list is a data value that identifies, indicates, or otherwise specifies to an API, such as context resource subdivision API 1302, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.In at least one embodiment, context resource partitioning API 1302 receives as input one or more arguments including a list of remaining resource return 1306. In at least one embodiment, return 1306 of remaining resource list is a data value that includes information usable for a storage location where a remaining resource list (e.g., a post-division remaining resource list) is to be stored as a result of using context resource division API 1302. In at least one embodiment, a resource return identified, indicated, or otherwise specified by resource return 1306 is one of a plurality of parameters that may be used by context resource partitioning APl 1302 to partition context resources. In at least one embodiment, API return 1306 is a data value that identifies, indicates, or otherwise specifies an API, such as context resource partitioning API 1302, a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.In at least one embodiment, context resource partitioning API 1302 receives as input one or more arguments including an input resource list 1308. In at least one embodiment, input resource list 1308 is a data value that includes information that can be used to identify, specify, or otherwise specify a list of resources to be divided using context resource partitioning API 1302. In at least one embodiment, an input resource list identified, indicated, or otherwise specified by input resource list 1308 is one of a plurality of parameters that may be used by context resource partitioning API 1302 to partition context resources. In at least one embodiment, input resource list 1308 is a data value for identifying, specifying, or otherwise specifying a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein to an API, such as context resource partition API 1302.In at least one embodiment, Unte Context Resource Partitioning API 1302 receives as input one or more arguments including flags 1310. In at least one embodiment, flags 1310 is a data value comprising information that can be used to identify, indicate, or otherwise specify one or more flags that indicate at least which resources (e.g., SMs) of input resource list 1308 are to be associated with a subdivided resource list (e.g., stored in subdivided resource list return 1304) and which remain (e.g., stored in remaining resource list return 1306) using context resource subdivision API 1302. In at least one embodiment, flags identified, indicated, or otherwise specified by flags 1310 are one of a plurality of parameters that may be used by context resource partition APl 1302 to partition context resources. In at least one embodiment, flags 1310 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to an API, such as context resource partitioning API 1302, that are to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, context resource partitioning API 1302 receives as input one or more arguments that include a minimum number 1312. In at least one embodiment, minimum number 1312 is a data value that includes information that can be used to identify, specify, or otherwise specify a minimum number (e.g., a smallest or minimum number of partitioned resources in a partitioned resource list) of partitioned resources obtained using context resource partitioning API 1302. In at least one embodiment, a minimum number identified, indicated, or otherwise specified by minimum number 1312 is one of a plurality of parameters that can be used by context resource subdivision API 1302 to subdivide context resources. In at least one embodiment, minimum number 1312 is a data value for identifying, specifying, or otherwise specifying a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein, to an API, such as context resource partitioning API 1302.In at least one embodiment, context resource subdivision API 1302 receives as input one or more arguments including one or more other arguments 1318. In at least one embodiment, other arguments 1318 are data including information to indicate other information that may be used in performing context resource partitioning API 1302 to partition context resources.In at least one embodiment, not shown in FIG. 13, a processor executes one or more instructions to execute one or more APIs, such as context resource partitioning API 1302, to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) to be stored by one or more processors according to a plurality of groups in which corresponding plurality of SMs are to be used using one or more arguments including, but not limited to, a partitioned resource resource list return 1304, a remaining resource resource list return 1306, an input resource list 1308, flags 1310, a minimum number 1312, and / or other arguments 1318. In at least one embodiment, not shown in FIG. 13, a processor executes one or more instructions to execute one or more APIs, such as context resource partition API 1302 to execute an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available to be used to execute one or more software cores, using one or more arguments including, but not limited to, partitioned resource return 1304, remaining resource return 1306, input resource list 1308, flags 1310, minimum number 1312, and / or other arguments 1318.In at least one embodiment, context resource partitioning API 1302, when invoked, causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, context resource partitioning APl 1302, when invoked, causes one or more APIs, such as one or more APIs 306, to add, insert, or otherwise include into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor, in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, in response to context resource partitioning API 1302, one or more APIs 306, if executed, are to cause one or more processors to perform a partitioning of context resource API return 1320. In at least one embodiment, context resource partitioning API return 1320 is an instruction set that, when executed, generates and / or indicates one or more data values in response to context resource partitioning API 1302. In at least one embodiment, context resource partition API return 1320 indicates success indication 1322. In at least one embodiment, success indication 1322 is data comprising any value to indicate success of context resource partitioning API 1302. In at least one embodiment, success indicator 1322 includes information indicating one or more specific types of success generated as a result of performing context resource partitioning API 1302. In at least one embodiment, success indicator 1322 includes information indicating one or more other data values generated as a result of context resource partitioning API 1302.In at least one embodiment, context resource partition API return 1320 indicates an error indication 1324. In at least one embodiment, fault indicator 1324 is data comprising any value indicative of failure of context resource partitioning API 1302. In at least one embodiment, fault indicator 1324 includes information indicating one or more specific types of faults generated as a result of execution of context resource partitioning API 1302. In at least one embodiment, fault indicator 1324 includes indicia indicating one or more other data values generated as a result of context resource partitioning API 1302.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, context resource partitioning API 1302 that adds various operations of different types to a stream executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a get semaphore operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memories, such as L2 cache memories of a PPU, such as a GPU, and / or cache memories of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more indications for communicating an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating communication of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with FIG. 5.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, context resource partitioning API 1302, one or more function signatures that can be used to indicate one or more callback functions for operations performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be performed. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code, that indicates a function signature for a callback function, as described herein at least in connection with FIG. 5.In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by context resource partitioning APl 1302 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, one or more data structures of one or more APIs 306 usable to indicate one or more external devices for which one or more APIs 306 are to transmit one or more operations use software code, such as example software code, indicating a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code indicating a data structure to indicate type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to execute in response to context resource partitioning API 1302 described above. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to context resource partitioning API 1302 use software code, such as example software code, that indicates a stream operation API call in parallel computing environment 308, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are comparable to how one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors are added to one or more streams or instruction sets in response to context resource partitioning API 1302 as described herein. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code, that indicates addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308, as described herein at least in connection with FIG. 5.FIG. 14 is a block diagram 1400 illustrating a process for performing an application programming interface (API) to partition context resources, in accordance with at least one embodiment. In at least one embodiment, process for performing a context resource partitioning API shown in block diagram 1400 is a context resource partitioning API 1302 process described herein at least in connection with FIG. 13. In at least one embodiment, a portion or all of process illustrated in block diagram 1400 for performing an API to partition context resources (or other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices as described in connection with FIGS. 26-58, configured with processor-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor such as processor 110 described herein in at least connection with FIG. 1 executes one or more steps of process to perform a context resource partitioning API shown in block diagram 1400. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to perform a context resource partitioning API shown in block diagram 1400.In at least one embodiment, in step 1402 of process for performing a context resource partitioning API shown in block diagram 1400, a processor performing process performs one or more operations to receive or otherwise obtain a context resource partitioning API. In at least one embodiment, at step 1402, an API for partitioning context resources is received by being otherwise provided to a processor such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, at step 1402, an API for partitioning context resources is received and otherwise provided by a processor, such as processor 102 described herein at least in connection with FIG. 1. In at least one embodiment, in step 1402, an API for partitioning context resources includes one or more arguments as described herein at least in connection with FIG. 13. In at least one embodiment, after step 1402, process shown in block diagram 1400 for performing an API to partition context resources continues at step 1404.In at least one embodiment, a processor performing this process performs one or more operations to determine whether a context resource partitioning API (e.g., received in step 1402) is valid in step 1404 of process shown in block diagram 1400 for performing a context resource partitioning API. In at least one embodiment, operations to determine whether a context resource partitioning API is valid include, at step 1404, operations to verify validity of arguments of context resource partitioning API (e.g., arguments described herein at least in connection with FIG. 13 ). In at least one embodiment, if it is determined that a context resource partitioning API is valid ("YES" branch), then in step 1404, process for performing a context resource partitioning API as shown in block diagram 1400 continues in step 1406. In at least one embodiment, if it is determined that a context resource partitioning API is not valid ("NElN" branch) in step 1404, then process shown in block diagram 1400 for performing a context resource partitioning API continues in step 1418, described below.In at least one embodiment, a processor performing process performs one or more operations to identify resources in an input resource list (e.g., received as an argument for a context resource partitioning API received in step 1402) in step 1406 of process shown in block diagram 1400 as described herein. In at least one embodiment, after step 1406, process for performing an API to partition context resources as shown in block diagram 1400 continues at step 1408.In at least one embodiment, a processor performing this process performs one or more input resource partitioning operations (e.g., identified in step 1406) using flags and / or minimum count (e.g., as arguments for a context resource partitioning API received in step 1402) in step 1408 of process shown in block diagram 1400 to perform a context resource partitioning API. In at least one embodiment, at step 1408, one or more operations are performed to partition input resources using an indication of available resources. In at least one embodiment, at step 1408, one or more input resource partitioning operations are performed to partition resources into two or more resource lists, as described herein. In at least one embodiment, no subdivision of input resources is made in step 1408 if no input resource list has been identified in step 1406. In at least one embodiment, after step 1408, process for performing an API to partition context resources, as shown in block diagram 1400, continues at step 1410.In at least one embodiment, a processor performing this process performs one or more operations to determine whether an input resource list has been subdivided (e.g., in step 1408) in step 1410 of process shown in block diagram 1400 for performing an API to subdivide context resources. In at least one embodiment, if it is determined in step 1410 that an input resource list has been subdivided ("YES" branch), process for performing a context resource subdivision API as shown in block diagram 1400 continues in step 1412. In at least one embodiment, if it is determined in step 1410 that an input resource list has not been partitioned ("NO" branch), process shown in block diagram 1400 for performing a context resource partitioning API continues in step 1418, which is described below.In at least one embodiment, in step 1412 of process shown in block diagram 1400 for performing a context resource partitioning API, a processor performing process performs one or more operations to store a partitioned input resource list (e.g., partitioned in step 1408) at a storage location received as an argument for a context resource partitioning API received in step 1402. In at least one embodiment, in step 1412, a subdivided input resource list is an empty list (e.g., does not contain resources). In at least one embodiment, after step 1412, process for performing an API to partition context resources, as shown in block diagram 1400, continues at step 1414.In at least one embodiment, a processor performing process performs one or more operations to store a remaining input resource list (e.g., resources remaining after an input resource list is divided in step 1408) at a storage location received as an argument for a context resource division API received in step 1402, in step 1414 of process shown in block diagram 1400. In at least one embodiment, in step 1412, a remaining input resource list is an empty list (e.g., it does not contain resources). In at least one embodiment, after step 1412, process for performing an API to partition context resources, as shown in block diagram 1400, continues at step 14164.In at least one embodiment, in step 1416 of process for performing an API to partition context resources indicated in block diagram 1400, a processor performing process performs one or more operations to return a success indication (e.g., a success indication 1322 described herein at least in connection with FIG. 13 ). In at least one embodiment, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ) in step 1416. In at least one embodiment, after step 1416, process for performing an API to partition context resources is ended, as shown in block diagram 1400. In at least one embodiment, not shown in FIG. 14, after step 1416, process for performing an API to partition context resources shown in block diagram 1400 continues at step 1402 described above.In at least one embodiment, in step 1418 of process for performing an API for partitioning context resources indicated in block diagram 1400, a processor performing process performs one or more operations to return an error indicator (e.g., an error indicator 1324 described herein at least in connection with FIG. 13 ). In at least one embodiment, in step 1418, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102 described herein at least in connection with FIG. 1 ). In at least one embodiment, after step 1418, process for performing an API to partition context resources is ended, as shown in block diagram 1400. In at least one embodiment, not shown in FIG. 14, after step 1418, process for performing an API to partition context resources shown in block diagram 1400 continues at step 1402 described above.In at least one embodiment, operations of process for performing an API to partition context resources shown in block diagram 1400 are performed in a different order than shown in FIG. 1400. In at least one embodiment, operations of process for performing an API to partition context resources shown in block diagram 1400 are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API to partition context resources shown in block diagram 1400 that are not dependent on each other (e.g., independent of order) are performed simultaneously or in parallel. In at least one embodiment, operations of process to perform an API for partitioning context resources shown in block diagram 1400 are performed by a plurality of threads executing on a processor such as those described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) will include one or more circuits to perform operations and / or instructions described herein in connection with FIGS. 13 and 14, such as one or more circuits to perform an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing operations described herein in connection with FIGS. 13 and 14, and / or instructions, such as one or more circuits for performing an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available for use in performing one or more software kernels and / or for otherwise performing operations described herein. In at least one embodiment, operations and / or instructions described herein in connection with FIGS. 13 and 14 include systems, methods, operations, and / or instructions described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which to use corresponding plurality of SMs and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with FIGS. 13 and 14 execute one or more processes described herein in connection with FIGS. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, not shown in FIGS. 13 and 14, one or more components described herein in connection with FIGS. 13 and 14 include one or more components described herein in connection with FIGS. 26-58. 26-58 to execute an application programming interface (API) to cause a corresponding plurality of identifiers of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.FIG. 15 is a block diagram 1500 illustrating an application programming interface (API) for generating a resource descriptor, according to at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a resource descriptor generation API 1502 to generate a descriptor for one or more resources. In at least one embodiment, not shown in FIG. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a resource descriptor generation APl 1502 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads. In at least one embodiment, not shown in FIG. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a resource descriptor generation API 1502, to execute an application programming interface (API), to generate a data structure containing information to indicate one or more streaming multiprocessors (SMs) of one or more GPUs that may be used to execute one or more software kernels. In at least one embodiment, also not shown in FIG. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a resource descriptor generation API 1502 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads in response to receipt of a second API, such as those described herein.In at least one embodiment, resource descriptor generation API 1502, when invoked, receives one or more arguments to indicate information about operations to be performed using techniques such as those described herein. In at least one embodiment, resource descriptor generation API 1502, when invoked, receives one or more arguments that indicate information about commands to execute using techniques such as those described herein.In at least one embodiment, resource descriptor generation API 1502 receives as input one or more arguments including resource descriptor return 1504. In at least one embodiment, resource descriptor return 1504 is a data value that includes information suitable for identifying, indicating, or otherwise specifying a storage location for storing a resource descriptor created with resource descriptor creation APl 1502. In at least one embodiment, a storage location identified, specified, or otherwise specified by resource descriptor return 1504 for a resource descriptor is one of a plurality of parameters that can be used by resource descriptor generation API 1502 to generate a resource descriptor. In at least one embodiment, resource descriptor return 1504 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein to an API, such as resource descriptor producer API 1502.In at least one embodiment, resource descriptor generate APl 1502 receives as input one or more arguments including resource list 1506. In at least one embodiment, resource list 1506 is a data value that includes information suitable for identifying, indicating, or otherwise specifying a resource list from which a resource descriptor is created using resource descriptor creation API 1502. In at least one embodiment, a resource list identified, indicated, or otherwise specified by resource list 1506 is one of a plurality of parameters that can be used by resource descriptor generation API 1502 to generate a resource descriptor. In at least one embodiment, resource list 1506 is a data value for identifying, indicating, or otherwise specifying a set of operations or instructions for an API, such as resource descriptor creation API 1502, to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.In at least one embodiment, resource descriptor generation API 1502 receives as input one or more arguments including a number of resources 1508. In at least one embodiment, number of resources 1508 is a data value comprising information suitable to identify, indicate, or otherwise specify a number of resources in resource list 1506 that may be used to generate a resource descriptor using resource descriptor generation API 1502. In at least one embodiment, number of resources identified, indicated, or otherwise specified by number of resources 1508 is one of a plurality of parameters that may be used by resource descriptor generation API 1502 to generate a resource descriptor. In at least one embodiment, number of resources 1508 is a data value that identifies, indicates, or otherwise specifies a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein, to an API, such as resource descriptor generation API 1502.In at least one embodiment, resource descriptor generation API 1502 receives as input one or more arguments including one or more other arguments 1518. In at least one embodiment, other arguments 1518 are data including information to indicate other information that may be used to generate a resource descriptor when executing resource descriptor generation API 1502.In at least one embodiment, not shown in FIG. 15, a processor executes one or more instructions to execute one or more APIs such as resource descriptor generation APl 1502 to execute an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads using one or more arguments including, but not limited to, resource descriptor return 1504, resource list 1506, number of resources 1508, and / or other arguments 1518. In at least one embodiment, not shown in FIG. 15, a processor executes one or more instructions to execute one or more APIs, such as resource descriptor generation API 1502, to execute an application programming interface (API) to generate a data structure containing information to indicate one or more streaming multiprocessors (SMs) of one or more GPUs that may be used to execute one or more software kernels using one or more arguments including, but not limited to, resource descriptor return 1504, resource list 1506, number of resources 1508, and / or other arguments 1518.In at least one embodiment, resource descriptor generation API 1502, when invoked, causes one or more APIs, such as one or more APIs 306 described herein at least in connection with FIG. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, resource descriptor generation API 1502, when invoked, causes one or more APls, such as one or more APIs 306, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor in a parallel computing environment, such as parallel computing environment 308 described herein at least in connection with FIG. 3.In at least one embodiment, one or more APIs 306, when executed, are to cause one or more processors to perform a resource descriptor creation return API 1520 in response to resource descriptor creation API 1502. In at least one embodiment, resource descriptor generate-generate API return 1520 is a set of instructions that, when executed, generate and / or indicate one or more data values in response to resource descriptor generate API 1502. In at least one embodiment, resource descriptor create API return 1520 generates a success indication 1522. In at least one embodiment, success indication 1522 is data comprising any value to indicate success of resource descriptor creation API 1502. In at least one embodiment, success indicator 1522 includes information indicating one or more specific types of success generated as a result of performance of resource descriptor generation API 1502. In at least one embodiment, success indicator 1522 includes information indicating one or more other data values generated as a result of resource descriptor generation API 1502.In at least one embodiment, an error indicator 1524 is specified at resource descriptor create API return 1520. In at least one embodiment, fault indicator 1524 is data comprising any value indicative of a fault of resource descriptor creation API 1502. In at least one embodiment, fault indicator 1524 includes information indicating one or more specific types of faults generated as a result of performance of resource descriptor generation API 1502. In at least one embodiment, error indicator 1524 includes information indicating one or more other data values generated as a result of resource descriptor generation API 1502.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, resource descriptor creation API 1502 that add various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a semaphore capture operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache memory of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include one or more indications for communicating an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating communication of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with FIG. 5.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, resource descriptor generation API 1502, one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be performed. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code, that indicates a function signature for a callback function, as described herein at least in connection with FIG. 5.In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations indicated by resource descriptor creation API 1502 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which one or more APIs 306 are to transmit one or more operations. In at least one embodiment, one or more data structures of one or more APIs 306 usable to indicate one or more external devices for which one or more APIs 306 are to transmit one or more operations use software code, such as example software code, indicating a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code indicating a data structure to indicate type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to execute in response to resource descriptor creation API 1502, as described above. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to generating resource descriptor generation API 1502 use software code, such as example software code, that indicates a stream operation API call in a parallel computing environment 308, as described herein at least in connection with FIG. 5.In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are similar to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to resource descriptor generation API 1502 as described herein. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code, that indicates addition of one or more operations or instructions to one or more executable graphs by one or more APIs 306 of parallel computing environment 308, as described herein at least in connection with FIG. 5.FIG. 16 is a block diagram 1600 illustrating a process for performing an application programming interface (API) to generate a resource descriptor, in accordance with at least one embodiment. In at least one embodiment, process for performing an API to generate a resource descriptor shown in block diagram 1600 is a process for performing resource descriptor generation API 1502 described herein at least in connection with FIG. 15. In at least one embodiment, a portion or all of a process for performing an API to generate a resource descriptor shown in block diagram 1600 (or any other processes described herein, or variations and / or combinations thereof) is performed under control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices as described in connection with FIGS. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) executing in common on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer readable storage medium in the form of a computer program comprising a plurality of computer readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer readable storage medium is a non-transitory computer readable medium. In at least one embodiment, a processor such as processor 110 described herein in at least connection with FIG. 1 performs one or more steps of process to execute an API to generate a resource descriptor shown in block diagram 1600. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of process to perform an API for creating a resource descriptor shown in block diagram 1600.In at least one embodiment, a processor performing process performs one or more operations to receive or otherwise obtain a resource descriptor generation API in step 1602 of process to perform an API shown in block diagram 1600. In at least one embodiment, at step 1602, an API for creating a resource descriptor is received or otherwise provided to a processor such as processor 110 described herein at least in connection with FIG. 1. In at least one embodiment, at step 1602, an API for generating a resource descriptor is received from a processor, such as processor 102 described herein at least in connection with FIG. 1. In at least one embodiment, in step 1602, an API for generating a resource descriptor includes one or more arguments as described herein at least in connection with FIG. 15. In at least one embodiment, after step 1602, process for performing an API to generate a resource descriptor as shown in block diagram 1600 continues at step 1604.In at least one embodiment, a processor performing process performs one or more operations to determine whether a resource descriptor generation API (e.g., received in step 1602) is valid in step 1604 of process for performing an API to generate a resource descriptor shown in block diagram 1600. In at least one embodiment, operations to determine whether a resource descriptor generation API is valid include operations to check validity of arguments of resource descriptor generation API (e.g., arguments described herein at least in connection with FIG. 1...
Claims
A processor comprising: one or more circuitry for performing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to perform one or more software threads.The processor of claim 1, wherein the API is to receive arguments comprising a subcontext that indicates the one or more data structures to be de-mapped.The processor of claim 1, wherein the API is to indicate that a subcontext comprising the one or more data structures has been released.The processor of claim 1, wherein the one or more data structures indicate a partitioning of a plurality of SMs of the one or more devices into at least the one or more SMs and a second one or more SMs that do not comprise the first one or more SMs.The processor of claim 1, wherein the API is to release one or more data structures that have been allocated by a second API to allocate the one or more data structures.The processor of claim 1, wherein the one or more data structures are comprised in a context of the one or more processors.The processor of claim 1, wherein the one or more data structures indicate a subcontext of a context of the one or more processors.A computer-implemented method, comprising: performing an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads.The computer-implemented method of claim 8, wherein performing the API comprises receiving arguments comprising a subcontext that indicates the one or more data structures to be released.The computer-implemented method of claim 8, wherein performing the API comprises indicating that a subcontext having the one or more data structures has been released.The computer-implemented method of claim 8, wherein the one or more data structures indicate partitioning of a plurality of SMs of the one or more devices into at least the one or more SMs and a second one or more SMs that do not comprise the first one or more SMs.The computer-implemented method of claim 8, wherein executing the API comprises de-mapping one or more data structures mapped by a second API to map the one or more data structures.The computer-implemented method of claim 8, wherein the one or more data structures are included in a context of the one or more processors.The computer-implemented method of claim 8, wherein the one or more data structures are comprised in a sub-context of a context of the one or more processors.A computer system comprising: one or more processors; and a memory storing executable instructions that, when executed by the one or more processors, execute an application programming interface (API) to enable one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors to use to execute one or more software threads.The computer system of claim 15, wherein the API is to receive arguments comprising a subcontext that indicates the one or more data structures to be de-mapped.The computer system of claim 15, wherein the API is to indicate that a subcontext comprising the one or more data structures has been released.The computer system of claim 15, wherein the one or more data structures indicate a partitioning of a plurality of SMs of the one or more devices into at least the one or more SMs and a second one or more SMs that do not comprise the first one or more SMs.The computer system of claim 15, wherein the API is to release one or more data structures that have been allocated by a second API to allocate the one or more data structures.The computer system of claim 15, wherein the one or more data structures are comprised in a subcontext of a context of the one or more processors.
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.18/593,593
US-PATENTANMELDUNGNR.18/593,710
US-PATENTANMELDUNGNR.18/593,578
US-PATENTANMELDUNGNR.18/593,731
US-PATENTANMELDUNGNR.18/593,720