Application programming interface for specifying the availability of multiprocessors

APIs for managing GPU resources through subcontext creation, destruction, and synchronization enhance the execution of software programs on GPUs, addressing inefficiencies in memory and resource utilization.

DE102025101960A1Pending Publication Date: 2025-07-31NVIDIA CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
DE102025101960
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-01-21
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing technologies face challenges in optimizing the execution of computer programs by improving memory and resource utilization on graphics processing units (GPUs).

Method used

The implementation of application programming interfaces (APIs) to manage resources by creating, destroying, and partitioning subcontexts, obtaining resources, and synchronizing events on GPUs, allowing for efficient allocation and utilization of GPU resources.

Benefits of technology

Enhances the execution of software programs on GPUs by optimizing memory and resource usage, leading to improved performance and efficiency in executing computational operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Devices, systems, and techniques for performing computational operations. In at least one embodiment, a processor executes an application programming interface to specify one or more streaming multiprocessors of one or more processors that can be used to execute one or more software threads.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to Provisional U.S. Application No. 63 / 625,278 (Attorney Docket No. 0112912-A27PR0), entitled "APPLICATION PROGRAMMING INTERFACE TO MANAGE RESOURCES," filed January 25, 2024, the entire contents of which are incorporated by reference into this application. This application also incorporates for all purposes the entire disclosure of concurrently filed U.S. patent application No. 18 / 593,578, entitled "APPLICATION PROGRAMMING INTERFACE TO ALLOCATE A DATA STRUCTURE" (Attorney Docket No. 0112912-A27US0), concurrently filed U.S. patent application No. 18 / 593,588, entitled "APPLICATION PROGRAMMING INTERFACE TO DEALLOCATE A DATA STRUCTURE" (Attorney Docket No. 0112912-C03US0), concurrently filed U.S. patent application No. 18 / 593,593, entitled "APPLICATION PROGRAMMING INTERFACE TO STORE AN IDENTIFIER OF A DATA STRUCTURE" (Attorney Docket No. 0112912-C04US0), concurrently filed U.S. patent application No.18 / 593,716 entitled “APPLICATION PROGRAMMING INTERFACE TO STORE IDENTIFIERS OF MULTIPROCESSOR GROUPS” (Attorney Docket No. 0112912-C06US0), concurrently filed U.S. patent application No. 18 / 593,720 entitled “APPLICATION PROGRAMMING INTERFACE TO INDICATE MULTIPROCESSOR GROUPS” (Attorney Docket No. 0112912-C07US0), concurrently filed U.S. patent application No. 18 / 593,722 entitled “APPLICATION PROGRAMMING INTERFACE TO READ FROM A DATA STRUCTURE” (Attorney Docket No. 0112912-C08US0), concurrently filed U.S. patent application No. 18 / 593,726 entitled “APPLICATION PROGRAMMING INTERFACE TO COMMUNICATE CONTEXT” (Attorney Docket No. 0112912-C09US0), and concurrently filed U.S. patent application No. 18 / 593,731, entitled “APPLICATION PROGRAMMING INTERFACE TO WAIT FOR CONTEXT” (Attorney Docket No. 0112912-C10US0). AREA

[0002] At least one embodiment relates to processing resources used to execute one or more software programs on a graphics processing unit ("GPU"). For example, at least one embodiment relates to implementing application programming interfaces for managing software program resources. BACKGROUND

[0003] Executing computer programs can require significant amounts of memory, time, or resources. The amount of memory, time, and / or resources used to execute computer programs can be improved. Despite advances that speed up or otherwise assist the execution of various components of a computer program, there are still challenges in executing computer programs with improved use of memory, time, and / or resources. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram illustrating a computer system for performing applications of computing resources via application programming interfaces (APIs) according to at least one embodiment; Fig. 2 is a block diagram illustrating contexts and subcontexts according to at least one embodiment; Fig. 3 is a block diagram illustrating a software program executed by one or more processors according to at least one embodiment; Fig. 4 is a block diagram illustrating a process for executing one or more application programming interfaces (APIs), in accordance with at least one embodiment; Fig. 5 is a block diagram illustrating an application programming interface (API) for creating a subcontext according to at least one embodiment; Fig. 6 is a block diagram illustrating a process for performing an application programming interface (API) to create a subcontext, in accordance with at least one embodiment; Fig. 7 is a block diagram illustrating an application programming interface (API) for destroying a subcontext, in accordance with at least one embodiment; Fig. 8 is a block diagram illustrating a process for implementing an application programming interface (API) to destroy a subcontext in accordance with at least one embodiment; Fig. 9 is a block diagram illustrating an application programming interface (API) for obtaining a subcontext from a stream, in accordance with at least one embodiment; Fig. 10 is a block diagram illustrating a process for performing an application programming interface (API) to obtain a subcontext from a data stream, in accordance with at least one embodiment; Fig. 11 is a block diagram illustrating an application programming interface (API) for obtaining resources associated with a context, in accordance with at least one embodiment; Fig. 12 is a block diagram illustrating a process for performing an application programming interface (API) to obtain resources associated with a context, in accordance with at least one embodiment; Fig. 13 is a block diagram illustrating an application programming interface (API) for partitioning context resources, in accordance with at least one embodiment; Fig. 14 is a block diagram illustrating a process for implementing an application programming interface (API) for partitioning context resources, in accordance with at least one embodiment; Fig. 15 is a block diagram illustrating an application programming interface (API) for generating a resource descriptor in accordance with at least one embodiment; Fig. 16 is a block diagram illustrating a process for implementing an application programming interface (API) to generate a resource descriptor in accordance with at least one embodiment; Fig. 17 is a block diagram illustrating an application programming interface (API) for retrieving device resources of a context, in accordance with at least one embodiment; Fig. 18 is a block diagram illustrating a process for performing an application programming interface (API) to obtain device resources of a context in accordance with at least one embodiment; Fig. 19 is a block diagram illustrating an application programming interface (API) for recording a context event according to at least one embodiment; Fig. 20 is a block diagram illustrating a process for implementing an application programming interface (API) for recording a context event, in accordance with at least one embodiment; Fig. 21 is a block diagram illustrating an application programming interface (API) for waiting for a context event, in accordance with at least one embodiment; Fig. 22 is a block diagram illustrating a process for performing an application programming interface (API) to wait for a context event, in accordance with at least one embodiment; Fig. 23 is a block diagram illustrating an exemplary software stack in which application programming interfaces (APIs) are processed, in accordance with at least one embodiment; Fig. 24 is a block diagram showing a processor and modules according to at least one embodiment; Fig. 25 is a block diagram illustrating a driver and / or runtime including one or more libraries to provide one or more application programming interfaces (APIs), according to at least one embodiment; Fig. 26 shows an exemplary data center according to at least one embodiment; Fig. 27 shows a processing system according to at least one embodiment; Fig. 28 shows a computer system according to at least one embodiment; Fig. 29 shows a system according to at least one embodiment; Fig. 30 shows an exemplary integrated circuit according to at least one embodiment; Fig. 31 shows a computer system according to at least one embodiment; Fig. 32 shows an APU according to at least one embodiment; Fig. 33 shows a CPU according to at least one embodiment; Fig. 34 shows an exemplary accelerator integration slice, according to at least one embodiment; Fig. 35A and Fig. 35B illustrate exemplary graphics processors according to at least one embodiment; Fig. 36A shows a graphics core according to at least one embodiment; Fig. 36B shows a GPGPU according to at least one embodiment; Fig. 37A shows a parallel processor according to at least one embodiment; Fig. 37B shows a processing cluster in accordance with at least one embodiment; Fig. 37C illustrates a graphics multiprocessor in accordance with at least one embodiment; Fig. 38 shows a graphics processor according to at least one embodiment; Fig. 39 shows a processor according to at least one embodiment; Fig. 40 shows a processor according to at least one embodiment; Fig. 41 shows a graphics processor core in accordance with at least one embodiment; Fig. 42 shows a PPU according to at least one embodiment; Fig. 43 shows a GPC according to at least one embodiment; Fig. 44 illustrates a streaming multiprocessor in accordance with at least one embodiment; Fig. 45 shows a software stack of a programming platform according to at least one embodiment; Fig. 46 shows a CUDA implementation of a software stack of Fig. 45 according to at least one embodiment; Fig. 47 shows an ROCm implementation of a software stack of Fig. 45, in accordance with at least one embodiment; Fig. 48 shows an OpenCL implementation of a software stack of Fig. 45 in accordance with at least one embodiment; Fig. 49 illustrates software supported by a programming platform in accordance with at least one embodiment; Fig. 50 shows the compilation of code for execution on the programming platforms of the Fig. 45 - 48, according to at least one embodiment; Fig. 51 shows in more detail how to compile code for execution on the programming platforms of the Fig. 45-48, in accordance with at least one embodiment; Fig. 52 illustrates translating source code prior to compiling source code according to at least one embodiment; Fig. 53A shows a system configured to compile and execute CUDA source code using different types of processing units, according to at least one embodiment; Fig. Figure 53B shows a system configured to run CUDA source code from Fig. 53A compiles and executes using a CPU and a CUDA-enabled graphics processor according to at least one embodiment; Fig. Figure 53C shows a system configured to run the CUDA source code from Fig. 53A compiles and executes using a CPU and a non-CUDA-capable GPU in accordance with at least one embodiment; Fig. 54 shows an example kernel translated by the CUDA-to-HIP translation tool of Fig. 53C in accordance with at least one embodiment; Fig. 55 shows a non-CUDA capable GPU from Fig. 53C in greater detail, in accordance with at least one embodiment; Fig. 56 shows how threads of an example CUDA grid are allocated to different processing units of Fig. 55, in accordance with at least one embodiment; Fig. 57 shows how existing CUDA code may be migrated to Data Parallel C++ code in accordance with at least one embodiment; and Fig. 58 shows components of a system for accessing a large language model according to at least one embodiment. DETAILED DESCRIPTION

[0004] In at least one embodiment, software (e.g., a computer program) executing on one or more processors causes software executing on other processors (e.g., in a data center) to perform operations to manage resources associated with the software executing on other processors. In at least one embodiment, software executing on one or more processors executes one or more application programming interfaces (APIs) to create a subcontext that manages a subset of resources that can be used to execute software programs.In at least one embodiment, software executing on one or more processors executes one or more APIs to create a subcontext that manages a subset of resources that can be used to execute software programs, thereby causing software executing on other processors to perform operations to manage resources associated with the software executing on other processors.

[0005] In at least one embodiment, software executing one or more processors executes one or more APIs to destroy a subcontext. In at least one embodiment, software executing one or more processors executes one or more APIs to destroy a subcontext, thereby causing software executing other processors to perform operations to manage resources associated with software executing other processors. In at least one embodiment, a subcontext destroyed by an application programming interface (API) to destroy a subcontext is a subcontext created by an API for creating a subcontext, as described above.

[0006] In at least one embodiment, software executing on one or more processors executes one or more APIs to create a subcontext from a data stream (e.g., a series of operations performed in a particular order by one or more other processors). In at least one embodiment, software executing on one or more processors executes one or more APIs to create a subcontext from a data stream, causing software executing on other processors to perform operations to manage resources associated with the software executing on other processors.In at least one embodiment, a subcontext created by an API for creating a subcontext to a data stream functions like a subcontext created by an API for creating a subcontext as described above, and is a subcontext that can be destroyed with an API for destroying a subcontext, also as described above.

[0007] In at least one embodiment, software executing on one or more processors executes one or more APIs to obtain resources associated with a context or subcontext. In at least one embodiment, software executing on one or more processors executes one or more APIs to obtain resources associated with a context or subcontext, thereby causing software executing on other processors to perform operations to manage resources associated with the software executing on other processors.In at least one embodiment, the resources obtained when executing an API for obtaining resources associated with a context or subcontext are resources associated with a context or subcontext created as described above (e.g., using an API for creating a subcontext or an API for creating a subcontext from a stream).

[0008] In at least one embodiment, software executing one or more processors executes one or more APIs to partition context resources according to one or more criteria. In at least one embodiment, software executing one or more processors executes one or more APIs to partition context resources, causing software executing other processors to perform operations to manage resources associated with software executing other processors. In at least one embodiment, the context resources partitioned by a context resource partitioning API are resources associated with a context or subcontext created as described above (e.g., using a subcontext creation API or a subcontext creation from stream API).In at least one embodiment, resource partitions (e.g., obtained through a context resource partitioning API) are used to create subcontexts as described herein.

[0009] In at least one embodiment, software executing one or more processors executes one or more APIs to generate a resource descriptor from a list of resources. In at least one embodiment, software executing one or more processors executes one or more APIs to generate a resource descriptor, causing software executing other processors to perform operations to manage resources associated with the software executing other processors. In at least one embodiment, the list of resource descriptors generated by an API is used to generate a resource descriptor to create and / or manage subcontext, as described above.

[0010] In at least one embodiment, software executing one or more processors executes one or more APIs to obtain device resources associated with a context or subcontext. In at least one embodiment, software executing one or more processors executes one or more APIs to obtain device resources, causing software executing other processors to perform operations to manage resources associated with the software executing other processors. In at least one embodiment, device resources obtained by executing a device resource retrieval API are usable to create a subcontext and / or to further subdivide resources, as described above.

[0011] In at least one embodiment, software executing one or more processors executes one or more APIs to record a context event (e.g., to emit an event associated with another context so that synchronization between contexts can be performed). In at least one embodiment, software executing one or more processors executes one or more APIs to record a context event, causing software executing other processors to perform operations to manage resources associated with software executing other processors. In at least one embodiment, a context event recorded by an API to record a context event is used by a context associated with an API to listen for a context event (see below).

[0012] In at least one embodiment, software executing one or more processors executes one or more APIs to listen for a context event (e.g., to listen for an event associated with another context so that synchronization between contexts can be performed). In at least one embodiment, software executing one or more processors executes one or more APIs to listen for a context event, causing software executing other processors to perform operations to manage resources associated with software executing other processors. In at least one embodiment, a context event of an API listening for a context event is used by a context associated with an API to record a context event, as described above.

[0013] Fig. 1 is a block diagram 100 illustrating a computer system for performing applications of computing resources via application programming interfaces (APIs), according to at least one embodiment. In at least one embodiment, a processor 102 executes one or more software programs 104. In at least one embodiment, the processor 102 is a processor as described below. In at least one embodiment, the processor 102 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general purpose graphics processing unit (GPGPU), a compute cluster, and / or a combination of these and / or other such processors. In at least one embodiment, the processor 102 is part of a computer system as described herein. In at least one embodiment described in Fig. 1, processor 102 is a processor of a client computing system. In at least one embodiment, a client computing system includes one or more client devices as described herein. In at least one embodiment, a client computing system includes one or more devices that are clients of a cloud computing environment as described herein.

[0014] In at least one embodiment, the software programs 104 include one or more computational operations, such as computational operations for training a neural network, executing a neural network, executing a Compute Uniform Device Architecture (CUDA) program, executing a large language model, performing a rendering operation, performing data analysis, and / or performing other operations, including, but not limited to, those described herein. In at least one embodiment, the software programs 104 include software as described herein.

[0015] In at least one embodiment, software programs 104 perform resource operations 106 using resource operation APIs 108. In at least one embodiment, resource operations 106 are operations for managing resources associated with processor 110. In at least one embodiment, processor 110 is a processor as described below. In at least one embodiment, processor 110 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general purpose graphics processing unit (GPGPU), a compute cluster, and / or a combination of these and / or other such processors. In at least one embodiment, processor 110 is part of a computer system as described herein.

[0016] In at least one embodiment described in Fig. 1, processor 110 is one of a plurality of processors of a high-performance computing system. In at least one embodiment, a high-performance computing system is a computing system that includes a plurality of processors to perform computational operations such as those described herein. In at least one embodiment, a high-performance computing system is a distributed computing system. In at least one embodiment, a high-performance computing system is a deep learning computing system. In at least one embodiment, operations to perform computational operations are performed by a high-performance computing system using the systems, methods, operations, and / or techniques described herein. In at least one embodiment, a high-performance computing system includes a cloud computing environment as described herein.In at least one embodiment, a high-performance computing system includes one or more processors, such as processor 110. In at least one embodiment, a high-performance computing system includes one or more graphics processors, as described herein. In at least one embodiment, the processors of a high-performance computing system include one or more central processing units (CPUs), graphics processing units (GPUs), parallel processing units (PPUs), general purpose graphics processing units (GPGPUs), compute clusters, and / or a combination of these and / or other such processors, as described herein.

[0017] In at least one embodiment, a high-performance computing system includes a plurality of processors, such as processor 110, having an associated set of resources that can be used by a processor, such as processor 102, to perform computational operations, including one or more resource operations 106 using one or more resource operation APIs 108. In at least one embodiment, resource operations 106 are computational operations for managing resources of processes performed by one or more processors using systems, methods, and / or operations as described herein.

[0018] In at least one embodiment, the resource operations APIs 108 include one or more APIs for managing resources of processors such as processor 110 that can be used to execute software programs 104. In at least one embodiment, resource operations APIs 108 include one or more APIs such as Get Subcontext from Stream API 502 (herein at least in connection with Fig. 5 and Fig. 6), Subcontext Destroy API 702 (herein at least in connection with Fig. 7 and Fig. 8), Get Subcontext from Stream API 902 (herein at least in connection with Fig. 9 and Fig. 10), Get Context Resource API 1102 (herein at least in connection with Fig. 11 and Fig. 12), context resource subdivision-APl 1302 (here at least in connection with Fig. 13 and Fig. 14), Get Device Resources API 1502 (here at least in conjunction with Fig. 15 and Fig. 16), Get Device Resources API 1702 (here at least in conjunction with Fig. 17 and Fig. 18), Context Event Capture API 1902 (herein at least in connection with Fig. 19 and Fig. 20), Wait for Context Event API 2102 (herein at least in connection with Fig. 21 and Fig. 22) and / or other such resource operations APIs.

[0019] In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more contexts 112 that include context resources 114. In at least one embodiment, a context 112 includes a description of resources available to a processor, such as processor 110, as described below at least in connection with Fig. 2. In at least one embodiment, processor resources 114 include resources available to processor 110 for executing one or more workloads (see below). In at least one embodiment, a context 112 is a primary context of processor 110, which is a previously created and / or default context of processor 110.

[0020] In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more subcontexts (e.g., subcontext 118A-118N) that can be used to execute one or more workloads (e.g., workload 120A-120N). In at least one embodiment, processor 110 receives resource operations APIs 108 and performs one or more operations to create one or more subcontexts based on one or more subsets of processor resources 114 (e.g., subset 116A-116N). In at least one embodiment, subcontexts (e.g., subcontext 118A-118N) are used to execute workloads (e.g., workload 120A-120N).In at least one embodiment, a subcontext is derived from context 112 that uses a subset of the resources available to processor 110 (e.g., as described at least in connection with . Fig. 2). In at least one embodiment, a subcontext is referred to as a green context. In at least one embodiment, a subset of processor resources 114 (e.g., subset 116A-116N) includes some or all of processor resources 114 (e.g., a first subset of processor resources 114 may include a first portion of said processor resources, a second subset of processor resources 114 may include a second portion of said processor resources, etc.). In at least one embodiment, a subset of processor resources is not empty (e.g., it includes at least a portion of processor resources 114).In at least one embodiment, a workload (e.g., workload 120A-120N) comprises at least a portion of computational operations of software programs 104 to be executed by one or more processors, such as processor 110, using systems, methods, and operations as described herein. In at least one embodiment, a workload (e.g., workload 120A-120N) is also referred to as a software workload. In at least one embodiment, a workload (e.g., workload 120A-120N) is also referred to as a kernel. In at least one embodiment, a workload (e.g., workload 120A-120N) is also referred to as a software kernel.

[0021] In at least one embodiment described in Fig. 1, the processor 110 receives resource operations APIs 108 and performs one or more operations to create a subcontext using subcontext creation API 502, as described herein at least in connection with Fig. 5 and Fig. 6. In at least one of Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations to destroy a subcontext using the subcontext destroy API 702, which is used here at least in conjunction with Fig. 7 and Fig. 8. In at least one of Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations to obtain a subcontext from a stream by using the processor 110 at least in conjunction with Fig. 9 and Fig. 10 described subcontext-from-stream API 902. In at least one embodiment described in Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations to obtain resources associated with a context using the Get Context Resource API 1102, which is used here at least in conjunction with Fig. 11 and Fig. 12. In at least one embodiment described in Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations for partitioning context resources using the context resource partitioning API 1302, which is illustrated here at least in conjunction with Fig. 13 and Fig. 14. In at least one embodiment described in Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations to generate a resource descriptor using the generate resource descriptor API 1502, which is used here at least in conjunction with Fig. 15 and Fig. 16. In at least one of Fig. 1, the processor 110 receives resource operations APIs 108 and performs one or more operations to retrieve resources of a device using Get Device Resources API 1702 (herein at least in connection with Fig. 17 and Fig. 18). In at least one of Fig. 1, the processor 110 receives the resource operations APIs 108 and performs one or more operations to record a context event using the context event recording API 1902, which is used here at least in conjunction with the Fig. 19 and Fig. 20. In at least one of Fig. 1, the processor 110 receives resource operations APIs 108 and performs one or more operations to wait for a context event using wait-for-context-event API 2102, which is used here at least in conjunction with Fig. 21 and Fig. 22 is described.

[0022] In at least one embodiment, a plurality of subsets of processor resources 114 (e.g., a plurality of subsets of subset 116A-116N) are used by and / or associated with a subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A-116N) is used by and / or associated with a subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A-116N) is used by and / or associated with a plurality of subcontexts (e.g., a plurality of subcontexts of subcontext 118A-118N).In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A-116N) is different than a number of subcontexts (e.g., subcontext 118A-118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A-116N) is identical to a number of subcontexts (e.g., subcontext 118A-118N).

[0023] In at least one embodiment, a plurality of subcontexts (e.g., a plurality of subcontexts of subcontext 118A-118N) are used by and / or associated with a subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or associated with a subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subset 116A-116N) is used by and / or associated with a plurality of subcontexts (e.g., a plurality of subcontexts of subcontext 118A-118N).In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A-116N) is different than a number of subcontexts (e.g., subcontext 118A-118N). In at least one embodiment, a number of subsets of processor resources 114 (e.g., subset 116A-116N) is identical to a number of subcontexts (e.g., subcontext 118A-118N).

[0024] In at least one embodiment, as used in each implementation described herein, terms such as "module" and nominalized verbs (e.g., subcontext, workload, context, and / or other terms) each refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein, unless the context indicates otherwise or explicitly states otherwise.In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any implementation described herein may include, for example, individually or in any combination, hard-wired circuits, programmable circuits, state machine circuits, fixed-function circuits, execution unit circuits, and / or firmware that stores instructions executed by programmable circuits. In at least one embodiment, modules may be embodied collectively or individually as circuits that are part of a larger system, such as an integrated circuit (IC), system-on-chip (SoC), etc.

[0025] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to implement the methods described herein in connection with Fig. 1, such as one or more circuits to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) should be used by one or more processors. 1, such as one or more circuits to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) should be used by one or more processors to execute one or more software threads and / or otherwise to perform operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to implement the methods described herein in connection with. Fig. 1, such as one or more circuits for executing an application programming interface (API) to generate one or more masks indicating one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more software threads and / or otherwise perform the operations described herein. 1, such as one or more circuits for executing an application programming interface (API) to generate one or more masks indicating one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the circuits described herein in connection with Fig. 1 described systems, methods, operations and / or instructions to execute an application programming interface (API) for allocating one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise to perform the operations described herein. In at least one embodiment, the systems described in connection with Fig. 1 described components one or more of the components associated with the Fig. 1-25 to execute an application programming interface (API) for allocating one or more data structures that indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or to otherwise perform the operations described herein. In at least one embodiment not described in Fig. 1 include one or more components described herein in connection with Fig. 1 described components one or more of the components associated with the Fig. 26-58 to implement an application programming interface (API) for allocating one or more data structures containing the indication of which of one or more streaming multiprocessors (SMs) is to be used by one or more processors to execute one or more software threads and / or to otherwise perform the operations described herein.

[0026] In at least one embodiment not described in Fig. 1, a non-transitory machine-readable medium comprises an instruction set that, when executed by one or more processors, performs the operations described herein at least in conjunction with the Fig. 1-25, such as operations for executing an application programming interface (API) to allocate one or more data structures that indicate which of one or more streaming multiprocessors (SMs) of one or more processors is to be used to execute one or more software threads and / or to otherwise perform operations described herein. 1-25, such as operations for executing an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors is to be used to execute one or more software threads and / or to otherwise perform operations described herein. In at least one embodiment not described in Fig. 1, a non-transitory machine-readable medium comprises an instruction set that, when executed by one or more processors, performs the operations described herein at least in conjunction with the Fig. 1-25, such as operations for executing an application programming interface (API) to generate one or more masks indicating one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein.

[0027] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to implement the methods described herein in connection with Fig. 1, such as one or more circuits to execute an application programming interface (API), to release one or more data structures, to indicate which of one or more streaming multiprocessors (SMs) of one or more processors is to be used. 1, such as one or more circuits to execute an application programming interface (API), to release one or more data structures, to indicate which of one or more streaming multiprocessors (SMs) of one or more processors is to be used to execute one or more software threads and / or to otherwise perform the operations described herein.In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to implement the methods described herein in connection with. Fig. 1, such as one or more circuits for executing an application programming interface (API) to cause the deactivation of one or more masks, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processors. 1, such as one or more circuits for executing an application programming interface (API) to cause the deactivation of one or more masks, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the circuits described herein in connection with Fig. 1 described systems, methods, acts, and / or instructions to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) should be used by one or more processors to execute one or more software threads and / or otherwise perform acts described herein. In at least one embodiment, the systems described in connection with Fig. 1 described components one or more of the components associated with the Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) should be used by one or more processors to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment not described in Fig. 1 include one or more components described herein in connection with Fig. 1 contain one or more components which are used here in conjunction with the Fig. 26-58 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) should be used by one or more processors to execute one or more software threads and / or otherwise perform operations described herein.

[0028] In at least one embodiment described in Fig. 1 is not shown, stored on a non-transitory machine-readable medium is a set of instructions which, when executed by one or more processors, are intended to perform operations described herein at least in connection with Fig. 1-25, such as operations for performing an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, a set of instructions is stored on a non-transitory machine-readable medium which, when executed by one or more processors, are intended to perform operations described here at least in connection with Fig. 1-25, such as operations for performing an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for performing one or more kernels and / or otherwise performing operations described herein.

[0029] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more interfaces to implement the methods described herein in connection with Fig. 1, such as one or more circuits to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to execute the operations and / or instructions described herein in connection with Fig. 1, such as one or more circuits to execute an application programming interface (API), to specify one or more identifiers of one or more masks, wherein the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for executing one or more kernels and / or otherwise performing operations described herein. In at least one embodiment, the methods described herein in connection with Fig. 1 described processes and / or commands systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, one or more of the methods described herein in connection with Fig. 1 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein.

[0030] In at least one embodiment described in Fig. 1 is not shown, a non-transitory, machine-readable medium comprises an instruction set which, when executed by one or more processors, described here at least in conjunction with Fig. 1-25, such as operations for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, Fig. 1 is not shown, stored on a non-transitory machine-readable medium is a set of instructions which, when executed by one or more processors, are intended to perform operations described herein at least in connection with Fig. 1-25, such as operations for performing an application programming interface (API) to specify one or more identifiers of one or more masks, wherein the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for performing one or more kernels and / or otherwise performing operations described herein.

[0031] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for implementing the operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) for specifying one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in executing one or more software kernels and / or for performing the operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 1 described processes and / or commands systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, comprise one or more components which are used here in conjunction with Fig. 1, one or more components described herein in connection with Fig. 26-58 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available to be used to perform one or more software threads and / or otherwise operations described herein.

[0032] In at least one embodiment described in Fig. 1, a non-transitory machine-readable medium comprises a set of instructions that, when executed by one or more processors, are operable to perform operations described herein at least in connection with Fig. 1-25, such as operations to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, Fig. 1 is not shown, a non-transitory machine-readable medium stores an instruction set which, when executed by one or more processors, which are used here at least in conjunction with Fig. 1-25, such as operations to perform an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in performing one or more software kernels and / or otherwise performing operations described herein.

[0033] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) may include one or more circuits for use in conjunction with Fig. 1 described operations and / or instructions, such as one or more circuits to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more interfaces for performing operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available for use in implementing one or more software kernels and / or otherwise implementing operations described herein. In at least one embodiment, the methods described herein in connection with Fig. 1 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, one or more of the methods described herein in connection with Fig. 1 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0034] In at least one embodiment described in Fig. 1 is not shown, stored on a non-transitory machine-readable medium is a set of instructions which, when executed by one or more processors, are intended to perform operations described herein at least in connection with Fig. 1-25, such as operations for performing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 1, a non-transitory machine-readable medium comprises an instruction set which, when executed by one or more processors, is operable to perform operations described herein at least in conjunction with Fig. 1-25, such as operations to execute an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available to be used to execute one or more software kernels and / or otherwise perform operations described herein.

[0035] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 1, such as one or more circuits for performing an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing the operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) to generate a data structure containing information to specify one or more streaming multiprocessors (SMs) of one or more GPUs as usable for implementing one or more software kernels and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 1 described processes and / or commands systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations otherwise described herein are performed. In at least one embodiment described in Fig. 1 is not shown, comprise one or more components which are used here in conjunction with Fig. 1, one or more components described herein in connection with Fig. 26-58 to implement an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations otherwise described herein are performed.

[0036] In at least one embodiment described in Fig. 1 is not shown, a non-transitory, machine-readable medium comprises an instruction set which, when executed by one or more processors, described here at least in conjunction with Fig. 1-25, such as performing operations to perform an application programming interface (API) to indicate one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or otherwise perform operations described herein. In at least one embodiment, Fig. 1 is not shown, a non-transitory machine-readable medium stores an instruction set which, when executed by one or more processors, which are used here at least in conjunction with Fig. 1-25, such as operations to execute an application programming interface (API) to create a data structure including information indicating that one or more streaming multiprocessors (SMs) of one or more GPUs may be used to execute one or more software cores and / or otherwise perform operations described herein.

[0037] In at least one embodiment, one or more processors (e.g., one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to, in conjunction with Fig. 1 described operations and / or instructions, such as one or more circuits to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits to perform the operations described herein in connection with Fig. 1 described operations and / or instructions, such as one or more circuits for implementing an application programming interface (API) to cause a number of streaming multiprocessors specified by one or more masks to be used to implement one or more specified software kernels and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 1 described processes and / or commands systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, one or more of the methods described herein in connection with Fig. 1 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or to otherwise perform operations described herein.

[0038] In at least one embodiment described in Fig. 1 is not shown, stored on a non-transitory machine-readable medium is a set of instructions which, when executed by one or more processors, are intended to perform operations described herein at least in connection with Fig. 1-25, such as operations for performing an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 1, a non-transitory machine-readable medium comprises an instruction set which, when executed by one or more processors, is operable to perform operations described herein at least in conjunction with Fig. 1-25, such as operations for executing an application programming interface (API) to cause a number of streaming multiprocessors specified by one or more masks to be usable to execute one or more specified software cores and / or to otherwise perform operations described herein.

[0039] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 1 described operations and / or instructions, such as one or more circuits for performing an application programming interface (API) to cause the context of one or more first software instructions to be passed to one or more second software instructions and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of the processor to be provided to one or more resources of a second context of the processor and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 1 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API), to cause the context of one or more first software programs to be communicated to one or more second software programs, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, one or more of the methods described herein in connection with Fig. 1 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API), to cause the context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein.

[0040] In at least one embodiment described in Fig. 1 is not shown, a non-transitory machine-readable medium stores an instruction set which, when executed by one or more processors, which are used here at least in conjunction with Fig. 1-25, such as operations for performing an application programming interface (API) to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, Fig. 1, a non-transitory machine-readable medium comprises an instruction set which, when executed by one or more processors, is operable to perform operations described herein at least in conjunction with Fig. 1-25, such as operations for executing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of the processor to be provided to one or more resources of a second context of the processor, and / or to otherwise perform operations described herein.

[0041] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for, in conjunction with, Fig. 1 described operations and / or instructions, such as one or more circuits to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to perform operations described herein in connection with Fig. 1, such as one or more circuits for implementing an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of the processor until one or more indications of one or more events of a second context are generated and / or otherwise perform operations described herein. In at least one embodiment, the applications and / or instructions described herein in connection with Fig. 1 described processes and / or commands systems, procedures, processes and / or commands described herein in connection with the Fig. 1-25 to execute an application programming interface (API), to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 1, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 1 is not shown, one or more of the methods described herein in connection with Fig. 1 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0042] In at least one embodiment described in Fig. 1 is not shown, a set of instructions is stored on a non-transitory machine-readable medium which, when executed by one or more processors, which are used here at least in conjunction with Fig. 1-25, such as operations for executing an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions, and / or otherwise perform operations described herein. In at least one embodiment, Fig. 1, a non-transitory machine-readable medium comprises a set of instructions which, when executed by one or more processors, described here at least in conjunction with Fig. 1-25, such as operations for performing an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of the processor until one or more indications of one or more events of a second context are generated and / or otherwise perform operations described herein.

[0043] Fig. 2 is a block diagram 200 illustrating contexts and subcontexts in accordance with at least one embodiment. In at least one embodiment, an apparatus 202 is used to execute one or more workloads (e.g., one or more of workloads 120A-120N, which are described herein at least in connection with Fig. 1). In at least one embodiment, the device 202 comprises one or more processors such as the processor 110, which is described here at least in conjunction with Fig. 1. In at least one embodiment, the device 202 includes one or more channels used to execute streams of workloads (e.g., a first workload followed by a second workload, etc.). In at least one embodiment described in Fig. 2, device 202 includes a hardware context (e.g., a set of hardware resources). In at least one embodiment, device 202 includes a plurality of hardware contexts. In at least one embodiment, performance of device 202 is improved when device 202 includes only one hardware context. In at least one embodiment, a hardware context of device 202 includes one or more primary contexts associated with that hardware context. In at least one embodiment, a primary context includes a description of and / or references to resources of device 202 (e.g., the hardware context). In at least one embodiment, a primary context is a one-to-one mapping with a hardware context of device 202 (e.g., it includes only references to hardware resources of device 202).

[0044] In at least one embodiment, one or more channels of device 202 are primary context channels 212 (e.g., C1 204, C2 206, C3 208, and C4 210) used by a primary context to execute software workloads using device 202. In at least one embodiment, primary context channels 212 are used exclusively by a primary context to execute software workloads. In at least one embodiment, one or more channels of device 202 are preferred selection channels 232 (e.g., C5 214, C6 220, C7 226, and C8 230). In at least one embodiment, preferred selection channels 232 are used by subcontexts to execute software workloads. In at least one embodiment, the subcontexts are dynamically assigned to one or more preferred selection channels 232.In at least one embodiment, a subcontext (also referred to as a green context) is a subset of a context (e.g., a subset of resources of device 202) that does not require a context switch (e.g., a hardware context switch). In at least one embodiment, a subcontext is a one-to-one mapping of a subset of resources of a hardware context. In at least one embodiment, software programs, such as software programs 104, described here at least in connection with. Fig. 1, perform computational operations to switch between subcontexts when workloads, such as workload 120A-120N, are executed using device 202. In at least one embodiment, one or more subcontexts are managed by a single thread executing on device 202. In at least one embodiment, only one subcontext managed by a thread may be current at a time. In at least one embodiment, a thread performs operations to switch between the subcontexts managed by that thread to keep one subcontext current at a time.

[0045] In at least one embodiment, a subcontext is associated with a plurality of channels of the preferred selection channels 232 (e.g., subcontext 216 associated with both C5 214 and C6 220, and subcontext 218 also associated with both C5 214 and C6 220). In at least one embodiment, a subcontext is associated with a single channel (e.g., subcontext 222 associated with C7 226, and subcontext 224 also associated with C7 226). In at least one embodiment, a subcontext is associated exclusively with one channel (e.g., subcontext 228 associated with C8 230).

[0046] In at least one embodiment, a primary context is referred to as a full context. In at least one embodiment, a full context includes access to all execution resources of device 202. In at least one embodiment, a primary context is a default context that a CUDA runtime uses to manage the resources of device 202. In at least one embodiment, a primary context is thread-safe. In at least one embodiment, a subcontext is mapped to a primary context and may use a subset of resources of device 202. In at least one embodiment, a subcontext may also use other resources (e.g., resources from other devices and / or contexts).

[0047] In at least one embodiment, a subcontext comprises a resource group for managing a subset of resources used by the subcontext. In at least one embodiment, a subcontext may implicitly specify a set of resources that a software library or framework (e.g., as described herein) may use. In at least one embodiment, a resource group represents one or more resources that a particular context, subcontext, or API may use. In at least one embodiment, a resource group is a descriptor (e.g., a description of resources) that may be partitioned by one or more APIs, such as those described herein, to manage resources (e.g., in a hierarchical manner).In at least one embodiment, a resource group includes descriptions or listings of a variety of different resources of device 202, including, but not limited to, streaming multiprocessors (SMs), device interconnects (e.g., copy and compute hardware channels), processor bindings, software schedulers, work distribution, etc. In at least one embodiment, other APIs (e.g., CUDA APIs) may be assigned to a subcontext (e.g., a current subcontext) and utilize resources of that subcontext. In at least one embodiment, other APIs (e.g., CUDA APIs) may be assigned directly to a resource group (e.g., without using a subcontext).

[0048] In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2, such as one or more circuits for performing an application programming interface (API) for allocating one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig.2 is not shown, comprise one or more components which are used here in conjunction with Fig. 2, one or more components described herein in connection with Fig. 26-58 to implement an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0049] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2, such as one or more circuits for performing an application programming interface (API) for enabling one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 2 is not shown, comprise one or more components which are used here in conjunction with Fig. 2, one or more components described herein in connection with Fig. 26-58 to perform an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0050] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2 described operations and / or instructions, such as one or more circuits for implementing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, comprise one or more components which are described here in connection with Fig. 2, one or more components described here in connection with Fig. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein.

[0051] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing functions described herein in connection with Fig. 2, such as one or more circuits for implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, the applications and / or instructions described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that may be used to execute one or more software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, comprise one or more components which are described here in connection with Fig. 2, one or more components described here in connection with Fig. 26-58 to perform an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to perform one or more software threads, and / or to otherwise perform operations described herein.

[0052] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2 described operations and / or instructions, such as one or more circuits for implementing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of identifiers of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, one or more of the methods described herein in connection with Fig. 2 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0053] In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2, such as one or more circuits for performing an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or for otherwise performing the operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, comprise one or more components which are described here in connection with Fig. 2, one or more components described here in connection with Fig. 26-58 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein.

[0054] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits to implement the methods described herein in connection with Fig. 2, such as one or more circuits to execute an application programming interface (API), to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, one or more of the methods described herein in connection with Fig. 2 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or to otherwise perform operations described herein.

[0055] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 2, such as one or more circuits for implementing an application programming interface (API) to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, one or more of the methods described herein in connection with Fig. 2 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API), to cause the context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein.

[0056] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing operations described herein in connection with Fig. 2, such as one or more circuits for implementing an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 2 described processes and / or commands systems, methods, processes and / or commands described herein in connection with the Fig. 1-25 to execute an application programming interface (API), to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 2, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 2 is not shown, one or more of the methods described herein in connection with Fig. 2 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0057] Fig. 3 is a block diagram 300 illustrating a software program executed by one or more processors according to at least one embodiment. In at least one embodiment, the block diagram 300 illustrates a software program 304 executed by a processor, such as a central processing unit (CPU) 302, as well as a graphics processing unit (GPU) 310 and an accelerator 314 within a heterogeneous processor. In at least one embodiment, the CPU 302 is a processor such as the processor 102, which is illustrated here at least in conjunction with Fig. 1. In at least one embodiment, CPU 302 is a processor such as processor 110, which is described here at least in connection with Fig. 1. In at least one embodiment, a CPU 302 is any processor having an architecture further described herein. In at least one embodiment, a CPU 302 is any general-purpose processor having any architecture further described herein. In at least one embodiment, a processor, such as a CPU 302, includes circuitry for performing one or more computational operations. In at least one embodiment, a processor, such as a CPU 302, includes any configuration of circuitry for performing one or more computational operations further described herein.

[0058] In at least one embodiment, a processor, such as a central processing unit (CPU) 302, executes a parallel computing environment 308. In at least one embodiment, a processor, such as a CPU 302, executes a parallel computing environment 308, such as the Compute Uniform Device Architecture (CUDA). In at least one embodiment, the parallel computing environment 308 includes instructions that, when executed by one or more processors, such as CPUs 302, facilitate the execution of one or more software programs by one or more CPUs 302, one or more parallel processing units (PPUs), such as GPUs 310, and / or one or more accelerators 314 within a heterogeneous processor.

[0059] In at least one embodiment, one or more PPUs are processors that include one or more circuits for performing parallel computational operations, such as GPUs 310 and any other parallel processors further described herein. In at least one embodiment, a GPU 310 is hardware that includes circuits for performing one or more computational operations, as described below in connection with various embodiments. In at least one embodiment, a GPU 310 includes one or more processing cores, each of which performs one or more computational operations. In at least one embodiment, a GPU 310 includes one or more processing cores to perform one or more parallel computational operations. In at least one embodiment, a GPU 310 is packaged with a CPU 302 or other processors as a system-on-chip (SoC).In at least one embodiment, a GPU 310 is packaged on a common chip or other substrate with a CPU 302 or other processors as a system-on-chip (SoC). In at least one embodiment, one or more accelerators 314 within heterogeneous processors are hardware that includes one or more circuits for performing specific computational operations, such as a deep learning accelerator (DLA), a programmable image processing accelerator (PVA), a field-programmable gate array (FPGA), or another accelerator described further herein. In at least one embodiment, an accelerator 314 is packaged within a heterogeneous processor together with a CPU 302 or other processors as a system-on-chip (SoC).In at least one embodiment, an accelerator 314 is packaged within a heterogeneous processor on a common die or other substrate with a CPU 302 or other processors as a system-on-chip (SoC). In at least one embodiment, one or more CPUs 302, one or more GPUs 310, or other PPUs, and / or accelerators 314 are packaged within heterogeneous processors as a system-on-chip (SoC). In at least one embodiment, one or more CPUs 302, one or more GPUs 310, or other PPUs, and / or accelerators 314 are packaged within heterogeneous processors on a common die or other substrate as a system-on-chip (SoC).

[0060] In at least one embodiment, parallel computing environment 308, such as CUDA, includes libraries and other software programs for performing one or more computational operations using one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within a heterogeneous processor. In at least one embodiment, parallel computing environment 308 includes libraries and other software programs that, when executed by one or more processors, such as one or more CPUs 302, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within a heterogeneous processor to perform one or more computational operations.In at least one embodiment, parallel computing environment 308 includes libraries that, when executed, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors to perform mathematical operations. In at least one embodiment, parallel computing environment 308 includes libraries that, when executed, cause one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors to perform any further operations described herein.

[0061] In at least one embodiment, one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors perform one or more computational operations in response to one or more application programming interfaces (APIs). In at least one embodiment, an API is a set of instructions that, when executed by one or more processors, such as CPUs 302, causes one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 within heterogeneous processors to perform one or more computational operations.In at least one embodiment, parallel computing environment 308 includes one or more APIs 306 that, when executed by one or more processors, such as CPUs 302, cause one or more PPUs, such as GPUs 310 and / or one or more accelerators 314 within heterogeneous processors, to perform one or more computational operations. In at least one embodiment, one or more APIs 306 include one or more functions that, when executed, cause one or more processors, such as CPUs 302, to perform one or more operations, such as computational operations, error reporting, scheduling of other operations to be performed by GPUs 310 and / or accelerators 314 within heterogeneous processors, or any other operation further described herein.In at least one embodiment, one or more APIs 306 include one or more functions that, when executed, cause one or more PPUs, such as GPUs 310, to perform one or more operations, such as computation, error reporting, or any other operation further described herein. In at least one embodiment, one or more APIs 306 include one or more functions such as those described below in connection with [ ]. Fig. 5-22, which, when executed, cause one or more accelerators 314 in heterogeneous processors to perform one or more operations, such as computational operations, error reporting, or any other operation described herein. In at least one embodiment, one or more APIs 306 include one or more functions to cause a CPU 302 to perform one or more computational operations in response to information or events generated by one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 in heterogeneous processors.In at least one embodiment, one or more APIs 306 include one or more functions that, when invoked, cause a CPU 302 to perform one or more computational operations in response to information or events generated by one or more PPUs, such as GPUs 310, and / or one or more accelerators 314 in heterogeneous processors.

[0062] In at least one embodiment, a processor, such as CPU 302, executes one or more software programs 304. In at least one embodiment, one or more software programs are sets of instructions that, when executed, cause one or more processors, such as CPUs 302, PPUs such as GPUs 310, and / or accelerators 314 in heterogeneous processors, to perform computational operations. In at least one embodiment, software programs 304 include instructions and / or operations executed by one or more PPUs, such as GPUs 310. In at least one embodiment, one or more software programs 304 include GPU-specific code 312 and / or accelerator-specific code 316. In at least one embodiment, the instructions and / or operations to be executed by one or more PPUs, such as GPUs 310, are PPU-specific or GPU-specific code 312.In at least one embodiment, GPU-specific code 312 is a set of software instructions and / or other operations, as further described herein, to be executed by one or more GPUs 310. In at least one embodiment, software programs 304 include instructions and / or operations to be executed by one or more accelerators 314 in heterogeneous processors. In at least one embodiment, the instructions and / or operations to be executed by one or more accelerators 314 in heterogeneous processors are accelerator-specific code 316. In at least one embodiment, accelerator-specific code 316 is a set of software instructions and / or other operations, as further described herein, to be executed by one or more accelerators 314.In at least one embodiment, PPU-specific or GPU-specific code 312 and / or accelerator-specific code 316 is executed in response to one or more APIs 306, as described below in connection with. Fig. 5-22 described.

[0063] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 3 described operations and / or instructions, such as one or more circuits for implementing an application programming interface (API) for allocating one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or other operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 3 is not shown, comprise one or more components which are used here in conjunction with Fig. 3, one or more components described herein in connection with Fig. 26-58 to implement an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0064] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 3, such as one or more circuits for performing an application programming interface (API) for enabling one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 3 is not shown, comprise one or more components which are used here in conjunction with Fig. 3, one or more components described herein in connection with Fig. 26-58 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0065] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 3, such as one or more circuits for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, one or more of the methods described herein in connection with Fig. 3 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein.

[0066] In at least one embodiment, one or more processors (e.g., those described herein) include one or more interfaces for performing the functions described herein in connection with Fig. 3, such as one or more circuits for implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, the applications and / or instructions described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that may be used to execute one or more software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, comprise one or more components which are used here in conjunction with Fig. 3, one or more components described herein in connection with Fig. 26-58 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available to be used to perform one or more software threads and / or otherwise operations described herein.

[0067] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing operations described herein in connection with Fig. 3, such as one or more circuits for implementing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of identifiers of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, one or more of the methods described herein in connection with Fig. 3 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0068] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 3, such as one or more circuits for implementing an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, comprise one or more components which are used here in conjunction with Fig. 3, one or more components described herein in connection with Fig. 26-58 to implement an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which one or more corresponding groups of software threads are scheduled and / or operations otherwise described herein are performed.

[0069] In at least one embodiment, one or more processors (e.g., as described herein) include one or more interfaces for communicating with Fig. 3 described operations and / or instructions, such as one or more circuits to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment, the circuits described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, one or more of the methods described herein in connection with Fig. 3 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators and / or to otherwise perform operations described herein.

[0070] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 3, such as one or more circuits for implementing an application programming interface (API) to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 3 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, one or more of the methods described herein in connection with Fig. 3 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API), to cause the context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein.

[0071] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing operations described herein in connection with Fig. 3, such as one or more circuits for implementing an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 3 described processes and / or commands systems, methods, processes and / or commands described herein in connection with the Fig. 1-25 to execute an application programming interface (API), to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 3, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 3 is not shown, one or more of the methods described herein in connection with Fig. 3 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0072] Fig. 4 is a block diagram 400 illustrating a process for executing one or more application programming interfaces (APIs) in accordance with at least one embodiment. In at least one embodiment, the process illustrated in block diagram 4 for executing one or more APIs utilizes one or more accelerators within a heterogeneous processor by a parallel computing environment, such as parallel computing environment 308, as described herein at least in connection with Fig. 3. In at least one embodiment, the process shown in block diagram 400 for executing one or more APIs 402 begins at step 404, wherein one or more processors execute a software program comprising one or more instructions that, when executed, cause the one or more processors and / or one or more other processors, such as graphics processing units (GPUs) and / or one or more accelerators within a heterogeneous processor or processors, to perform one or more computational operations. In at least one embodiment, a software program to be executed by one or more processors at step 404 comprises one or more instructions that, when executed, cause execution of one or more APIs 306 of a parallel computing environment 308, as described above.In at least one embodiment, after step 404, the process shown in block diagram 400 continues to execute one or more APIs in step 406.

[0073] In at least one embodiment, a processor performing the process shown in block diagram 400 for performing one or more APIs determines in step 406 of the process for performing one or more APIs shown in block diagram 400 whether an API such as that described herein at least in connection with Fig. 5-22 (e.g., Create Subcontext API 502, Destroy Subcontext API 702, Get Context Resource From Stream API 902, Get Context From Stream API 1102, Partition Context Resource API 1302, Create Resource Descriptor API 1502, Get Device Resource API 1702, Get Context Resource Capture API 1902, and / or Wait for Context Event API 2102). In at least one embodiment, if it is determined in step 406 that an API is not executing ("NO" branch), the process shown in block diagram 400 continues to execute one or more APIs in step 416. In at least one embodiment, if it is determined in step 406 that an API is being executed ("YES" branch), the process shown in block diagram 400 for executing one or more APIs continues in step 408.

[0074] In at least one embodiment, in step 408 of the process shown in block diagram 400 for performing one or more APIs, a processor performing the process shown in block diagram 400 for performing one or more APIs executes an API as described herein at least in connection with Fig. 5-22. In at least one embodiment, in step 408, one or more processors execute one or more instructions to make one or more API calls, as described herein at least in connection with Fig. 5-22 (e.g., Create Subcontext API 502, Destroy Subcontext API 702, Get Subcontext From Stream API 902, Get Context Resource API 1102, Partition Context Resource API 1302, Create Resource Descriptor API 1502, Get Resource Descriptor API 1702, Record Context Event API 1902, and / or Wait for Context Event API 2102) by the one or more processors and / or by one or more other processors, such as GPUs and / or accelerators within a heterogeneous processor, as described above. In at least one embodiment, after step 408, the process shown in block diagram 400 continues to execute one or more APIs in step 410.

[0075] In at least one embodiment, a processor performing the process illustrated in block diagram 400 for performing one or more APIs determines in step 410 of the process for performing one or more APls illustrated in block diagram 400 whether a return value as a result of executing one or more instructions for performing one or more API calls, as described herein at least in connection with Fig. 5-22 (e.g., Create Subcontext API 502, Destroy Subcontext API 702, Get Subcontext from Stream API 902, Get Context Event API 1102, Partition Context Resource API 1302, Create Resource Descriptor API 1502, Get Resource Descriptor API 1702, Record Context Event API 1902, and / or Wait for Context Event API 2102) by the one or more processors and / or by one or more other processors, such as GPUs and / or accelerators within a heterogeneous processor, as described above. In at least one embodiment, a processor performing the process shown in block diagram 400 for executing one or more APIs determines in step 410 whether to return a value using an API return, as described herein at least in connection with Fig. 5-22 (e.g., Get Context Resource API Return 520, Destroy Subcontext API Return 720, Get Context From Stream API Return 920, Get Context Resource API Return 1120, Get Context Resource API Return 1320, Get Context Event Capture API Return 1520, Get Resource Descriptor API Return 1720, Resource Descriptor Capture API Return 1920, and / or Wait for Context Event API Return 2120). In at least one embodiment, if it is determined in step 410 that a return value is returned ("YES" branch), the process continues to perform one or more of the APIs shown in block diagram 400 in step 412.In at least one embodiment, if it is determined in step 410 that no return value is returned ("NO" branch), the process continues to perform one or more of the APIs illustrated in block diagram 400 in step 414.

[0076] In at least one embodiment, a processor performing the process shown in block diagram 400 to execute one or more APIs sets a return value in step 412. In at least one embodiment, a return value is set in step 412 by storing the return value in a memory location defined by an API, as described herein at least in connection with Fig. 5-22 (e.g., Get Subcontext API 502, Destroy Subcontext API 702, Get Subcontext from Stream API 902, Get Context Event API 1102, Subdivide Context Resource API 1302, Create Resource Descriptor API 1502, Device Resource API 1702, Capture Context Event API 1902, and / or Wait for Context Event API 2102). In at least one embodiment, in step 412, a return value is determined by storing the return value in a memory location that includes an API return, as described at least in connection with Fig. 5-22 (e.g., Get Context Resource API Return 520, Subcontext Destroy API Return 720, Get Context From Stream API Return 920, Get Context Resource API Return 1120, Get Context Resource API Return 1320, Get Resource Descriptor Capture API Return 1520, Get Resource Descriptor Capture API Return 1720, Get Resource Descriptor Capture API Return 1920, and / or Wait for Context Event API Return 2120). In at least one embodiment, after step 412, the process shown in block diagram 400 continues to execute one or more APIs in step 414.

[0077] In at least one embodiment, a processor performing the process illustrated in block diagram 400 to execute one or more APIs returns the success or failure (e.g., an error) in step 414 using an API return, as described at least in connection with Fig. 5-22 (e.g., obtaining subcontext API return 520, destroying subcontext API return 720, obtaining subcontext from stream API return 920, obtaining context resources API return 1120, partitioning context resources API return 1320, creating resource descriptors API return 1520, obtaining device resources API return 1720, recording context events API return 1920, and / or waiting for context events API return 2120). In at least one embodiment, after step 414, the process shown in block diagram 400 continues to execute one or more APIs in step 416.

[0078] In at least one embodiment, a processor performing the process shown in block diagram 400 for executing one or more APIs determines, in step 416 of the process shown in block diagram 400 for executing one or more APIs, whether execution of the software program (e.g., started in step 404) is complete. In at least one embodiment, a processor performing the process shown in block diagram 400 for executing one or more APIs determines, in step 416, that execution of the software program (e.g., started in step 404) is complete, based at least in part on whether one or more processors are executing instructions of the software program (e.g., started in step 404).In at least one embodiment, in step 416, if it is determined that the execution of the software program (e.g., started in step 404) is complete, the process shown in block diagram 400 for executing one or more APIs is terminated (418). In at least one embodiment, in step 416, if it is determined that the execution of the software program (e.g., started in step 404) is not complete, the process shown in block diagram 400 for executing one or more APIs is continued in step 404 to continue the execution of one or more instructions of a software program.

[0079] In at least one embodiment, the operations of the process shown in block diagram 400 for executing one or more APIs are performed in a different order than in Fig. 4. In at least one embodiment, the operations of the process shown in block diagram 400 for executing one or more APIs are performed concurrently or in parallel. In at least one embodiment, operations of the process shown in block diagram 400 for executing one or more APIs that are not dependent on one another (e.g., are independent of order) are performed concurrently or in parallel. In at least one embodiment, the operations of the process for executing one or more of the APIs shown in block diagram 400 are performed by a plurality of threads executing on a processor such as those described herein.

[0080] In at least one embodiment, one or more processors (e.g., those described herein) include one or more circuits for performing the functions described herein in connection with Fig. 4, such as one or more circuits for performing an application programming interface (API) for allocating one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are used here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to implement an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0081] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 4, such as one or more circuits for performing an application programming interface (API) for enabling one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) for exposing one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are used here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to perform an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0082] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits to implement the methods described herein in connection with Fig. 4, such as one or more circuits to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein.

[0083] In at least one embodiment, components used herein in conjunction with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are described here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein.

[0084] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing functions described herein in connection with Fig. 4, such as one or more circuits for implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for use in executing one or more software threads and / or otherwise performing operations described herein. In at least one embodiment, the applications and / or instructions described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that may be used to execute one or more software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are used here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to implement an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein.

[0085] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 4 described operations and / or instructions, such as one or more circuits for implementing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause a corresponding plurality of identifiers of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, one or more of the methods described herein in connection with Fig. 4 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0086] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing the functions described herein in connection with Fig. 4, such as one or more circuits for implementing an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are described here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads and / or otherwise perform operations described herein.

[0087] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits to implement the methods described herein in connection with Fig. 4, such as one or more circuits to execute an application programming interface (API), to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, comprise one or more components which are used here in connection with Fig. 4, one or more components described here in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more indicators of one or more numbers of one or more streaming multiprocessors (SMs) of one or more processors to be read from one or more data structures storing the one or more indicators, and / or to otherwise perform operations described herein.

[0088] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing functions described herein in connection with Fig. 4, such as one or more circuits for performing an application programming interface (API) to cause the context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, the operations and / or instructions described herein in connection with Fig. 4 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API), to cause the context of one or more first software instructions to be passed to one or more second software instructions, and / or to otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, one or more of the methods described herein in connection with Fig. 4 described components one or more herein in connection with Fig. 26-58 to perform an application programming interface (API), to cause the context of one or more first software instructions to be communicated to one or more second software instructions, and / or to otherwise perform operations described herein.

[0089] In at least one embodiment, one or more processors (e.g., as described herein) include one or more circuits for performing operations described herein in connection with Fig. 4, such as one or more circuits for performing an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs, and / or to otherwise perform operations described herein. In at least one embodiment, the operations described herein in connection with Fig. 4 described processes and / or commands Systems, procedures, processes and / or commands described herein in connection with the Fig. 1-25 to execute an application programming interface (API), to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software instructions, and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 4, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more second instructions to wait until the one or more second instructions receive a context corresponding to one or more first software programs and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 4 is not shown, one or more of the methods described herein in connection with Fig. 4 described components one or more herein in connection with Fig. 26-58 to execute an application programming interface (API) to cause one or more second instructions to wait to execute until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0090] Fig. 5 is a block diagram 500 illustrating an application programming interface (API) for creating a subcontext in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a subcontext creation API 502 to create a subcontext of a primary context using resources of the primary context. In at least one embodiment, Fig. 5, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a subcontext creation API 502 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads. In at least one embodiment, as described in Fig. 5, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a subcontext creation API 502 to create an application programming interface (API) to generate one or more masks that specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels. In at least one embodiment, also not shown in Fig. 5, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a create subcontext API 502 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads in response to receiving a second API, such as those described herein.

[0091] In at least one embodiment, the subcontext creation API 502, when invoked, receives one or more arguments that specify information about operations to be performed using techniques such as those described herein. In at least one embodiment, the subcontext creation API 502, when invoked, receives one or more arguments that specify information about commands to be executed using techniques such as those described herein.

[0092] In at least one embodiment, the subcontext creation API 502 receives as input one or more arguments comprising a subcontext return 504. In at least one embodiment, the subcontext return 504 is a data value comprising information that can be used to identify, indicate, or otherwise specify a location for storing a created subcontext (e.g., created with the subcontext creation API 502). In at least one embodiment, a subcontext return location identified, indicated, or otherwise specified by the subcontext return 504 is one of a plurality of parameters that can be used by the subcontext creation API 502 to create a subcontext.In at least one embodiment, subcontext return 504 is a data value for identifying, indicating, or otherwise specifying an API, such as create subcontext API 502, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0093] In at least one embodiment, the Create Subcontext API 502 receives as input one or more arguments comprising a resource descriptor 506. In at least one embodiment, the resource descriptor 506 is a data value comprising information that can be used to identify, indicate, or otherwise specify a location of a resource descriptor to create a subcontext (e.g., created with the Create Subcontext API 502). In at least one embodiment, a resource descriptor identified, indicated, or otherwise specified by the Resource Descriptor API 506 is one of a plurality of parameters that can be used by the Create Subcontext API 502 to create a subcontext.In at least one embodiment, resource descriptor 506 is a data value for identifying, indicating, or otherwise specifying a set of operations or instructions for an API, such as create subcontext API 502, to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.

[0094] In at least one embodiment, the subcontext creation API 502 receives as input one or more arguments comprising a device 508. In at least one embodiment, the device 508 is a data value that includes information that can be used to identify, indicate, or otherwise specify a device associated with a subcontext (e.g., created using the subcontext creation API 502). In at least one embodiment, a device identified, indicated, or otherwise specified by the device 508 is one of a plurality of parameters that can be used by the subcontext creation API 502 to create a subcontext.In at least one embodiment, device 508 is a data value that identifies, indicates, or otherwise specifies to an API such as create subcontext API 502 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0095] In at least one embodiment, the Create Subcontext API 502 receives as input one or more arguments comprising flags 510. In at least one embodiment, flags 510 is a data value comprising information used to identify, indicate, or otherwise specify flags when creating a subcontext (e.g., using Create Subcontext API 502). In at least one embodiment, flags 510 is a data value indicating which SMs of a device (e.g., indicated by device 508) are to be used for a subcontext created by Create Subcontext API 502. In at least one embodiment, the flags identified, indicated, or otherwise specified by flags 510 are one of a plurality of parameters that can be used by Create Subcontext API 502 to create a subcontext.In at least one embodiment, flags 510 is a data value that identifies, indicates, or otherwise specifies to an API, such as subcontext creation API 502, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.

[0096] In at least one embodiment, the subcontext creation API 502 receives as input one or more arguments that include one or more other arguments 518. In at least one embodiment, the other arguments 518 are data that include information to indicate other information that can be used in executing the subcontext creation API 502 to create a subcontext.

[0097] In at least one embodiment not described in Fig. 5, a processor executes one or more instructions to execute one or more APIs, such as create subcontext API 502, to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads, using one or more arguments, including, but not limited to, subcontext return 504, resource descriptor 506, device 508, flags 510, and / or other arguments 518. In at least one embodiment not shown in Fig. 5, a processor executes one or more instructions to execute one or more APIs, such as create subcontext API 502 to execute an application programming interface (API) to generate one or more masks to indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels, using one or more arguments, including, but not limited to, subcontext return 504, resource descriptor 506, device 508, flags 510, and / or other arguments 518.

[0098] In at least one embodiment, the subcontext creation API 502, when invoked, causes one or more APIs such as one or more APIs 306 described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the Create Subcontext API 502, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or set of instructions to be executed by one or more accelerators within a heterogeneous processor.

[0099] In at least one embodiment, one or more APIs 306, in response to the create subcontext API 502, are to cause one or more processors to perform a create subcontext API return 520, if performed. In at least one embodiment, the create subcontext API return 520 is a set of instructions that, when executed, generates and / or specifies one or more data values in response to the create subcontext API 502. In at least one embodiment, the create subcontext API return 520 specifies a success indicator 522. In at least one embodiment, the success indicator 522 is data comprising any value to indicate the success of the create subcontext API 502.In at least one embodiment, success indicator 522 includes information indicating one or more specific types of successes generated as a result of performing subcontext creation API 502. In at least one embodiment, success indicator 522 includes information indicating one or more other data values generated as a result of subcontext creation API 502.

[0100] In at least one embodiment, the create subcontext API return 520 indicates an error indicator 524. In at least one embodiment, the error indicator 524 is data comprising any value indicating the failure of the create subcontext API 502. In at least one embodiment, the error indicator 524 includes information indicating one or more specific types of errors generated as a result of performing the create subcontext API 502. In at least one embodiment, the error indicator 524 includes indications indicating one or more other data values generated as a result of the create subcontext API 502.

[0101] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, create subcontext API 502, that add various operations of different types to a data stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include a semaphore operation. In at least one embodiment, the stream operations include a release semaphore operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, such as L2 cache of a PPU, such as a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more operations to indicate the transmission of an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, the example software code specifying the types of stream operations is as follows: / ** * Types of stream operations * / typedef enum { / **< Acquire Semaphore * / CUSOCKET_STREAM_OP_SEMA_ACQ, / * *< Release Semaphore * / CUSOCKET_ STREAM OP _ SEMA REL, / **< Flush GPU L2 Cache * / CUSOCKET_STREAM_OP_GPU_L2_FLUSH, / **< Invalidate GPU L2 Cache * / CUSOCKET_STREAM_OP_GPU_L2_INVALIDATE, / **< Submit Operation to External Device * / CUSOCKET_STREAM_OP_EXTERNAL_DEVICE_SUBMIT} cuSocketStreamOpType;.

[0102] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, create subcontext API 502, one or more function signatures that can be used to specify one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be executed. In at least one embodiment, example software code specifying a function signature for a callback function is as follows: / ** * Function signature of the callback function to submit to an external device. * / typedef unsigned int (*cuSocketExternalDeviceSubmitCallback)(void *submitArgs);

[0103] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by the Create Subcontext API 502 to one or more APIs 306, one or more data structures from one or more APIs 306 are usable to specify one or more external devices for which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, example software code specifying a data structure representing a device node for one or more accelerators within heterogeneous processors is as follows: / ** * Structure representing the external device node that submits information about a particular task for an external device. * / typedef struct { void *submitArgs; cuSocketExternalDeviceSubmitCallback callback;} cuSocketExternalDeviceNodeParams;。

[0104] In at least one embodiment, one or more data structures of one or more APIs 306 are to be used to specify the type and data of one or more operations to be performed by one or more accelerators in heterogeneous processors. In at least one embodiment, example software code specifying a data structure to specify the type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors is as follows: / ** * Structure that tracks the type and data for stream operations. The data is filled with * semaphore address and payload for the types * ::CUSOCKET_STREAM_OP_SEMA_ACQ and * ::CUSOCKET_STREAM_OP_SEMA_REL * / typedef struct { / ** * Type of stream operation * / cuSocketStreamOpType type; union { / ** * Parameters for semaphore * / struct { / ** * Address of the semaphore to be acquired or released.* / void *semaAddr; / ** * Semaphore payload value. * / unsigned int payload;} sema; / ** * The specific task that needs to be passed to the external device. * / cuSocketExternalDeviceNodeParams task;} data;} cuSocketStreamOp;.

[0105] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions executed by one or more accelerators in heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to be executed in response to the create subcontext API 502, as described above. In at least one embodiment, example software code indicating a stream operation API call in the parallel computing environment 308, such as CUDA, is as follows: / ** * Submits a list of operations to a CUDA stream. * * - param[in] usrStream - The stream to which the operations are submitted.* * - param[in] streamOp - The list of operations to submit. * * - param[in] count - The number of operations to submit. * * - Returns CUDA_SUCCESS on success; otherwise, an appropriate error is raised. * / CUresult cuSocketStreamOps( CUstream usrStream, cuSocketStreamOp *streamOp, unsigned int count, unsigned int flags );.

[0106] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs, similar to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to the create subcontext API 502. In at least one embodiment, example software code indicating the addition of one or more operations or instructions to one or more executable graphs through one or more APIs 306 of the parallel computing environment 308 is as follows: / ** * Submits a task for an external device to a CUDA stream. * * - param[in] graphNode - The newly created node.* * - param[in] graph - The graph in which this node should be added. * * - param[in] dependencies - The dependencies that must be satisfied before this node can * be executed. * - param[in] numDependencies - The number of dependencies. * - param[in] nodeParams - The node's execution parameters. * * - Returns CUDA_SUCCESS on success, otherwise raises an appropriate error. * / CUresult cuSocketAddExternalDeviceNode ( CUgraphNode* graphNode, CUgraph graph, CUgraphNode* dependencies, unsigned int numDependencies, cuSocketExternalDeviceNodeParams* nodeParams ); .

[0107] Fig. 6 is a block diagram 600 illustrating a process for executing an application programming interface (API) to create a subcontext, in accordance with at least one embodiment. In at least one embodiment, the process illustrated in block diagram 600 for executing an API to create a subcontext is a process for executing the subcontext creation API 502, which is illustrated here at least in conjunction with Fig. 5. In at least one embodiment, part or all of the process shown in block diagram 600 for performing an API for creating a subcontext (or other processes described herein or variations and / or combinations thereof) is performed under the control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with Fig. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions that can be executed by one or more processors such as those described herein. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, a processor such as the one described herein, at least in conjunction with Fig. 1, processor 110 performs one or more steps of the process to execute an API to create a subcontext shown in block diagram 600. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of the process to execute an API to create a subcontext shown in block diagram 600.

[0108] In at least one embodiment, a processor performing the process performs one or more operations to receive or otherwise obtain an API for creating a subcontext, as illustrated in block diagram 600, at step 602 of the process. In at least one embodiment, in step 602, an API for creating a subcontext is received or otherwise provided to a processor, such as processor 110, which is illustrated here at least in connection with Fig. 1. In at least one embodiment, in step 602, an API for creating subcontext is called by a processor, such as processor 102, which is used here at least in conjunction with Fig. 1 described in step 602. In at least one embodiment, an API for creating a subcontext includes one or more arguments, described herein at least in connection with Fig. 5. In at least one embodiment, after step 602, the process continues to perform an API for creating a subcontext, as shown in block diagram 600, in step 604.

[0109] In at least one embodiment, in step 604 of the process shown in block diagram 600 for performing a subcontext creation API, a processor performing the process performs one or more operations to determine whether a subcontext creation API (e.g., received in step 602) is valid. In at least one embodiment, the operations to determine whether a subcontext creation API is valid include, in step 604, operations to check the validity of arguments of the subcontext creation API (e.g., arguments described here at least in connection with Fig. 5). In at least one embodiment, if a subcontext creation API is determined to be valid (a "YES" branch) at step 604, the process shown in block diagram 600 for performing a subcontext creation API continues in step 606. In at least one embodiment, if a subcontext creation API is determined to be invalid (a "NO" branch) at step 604, the process shown in block diagram 600 for performing a subcontext creation API continues in step 620, described below.

[0110] In at least one embodiment, in step 606 of the process shown in block diagram 600 for performing an API for creating a subcontext, a processor performing the process performs one or more operations to obtain resources from a device. In at least one embodiment, in step 606, one or more operations to retrieve resources from a device include one or more operations to query a context (e.g., a primary context) of a device such as that described herein at least in connection with Fig. 2 to determine a set of resources available to or otherwise allocated to the device, as described here at least in connection with Fig. 2. In at least one embodiment, a device is an argument of an API to create a subcontext, as described herein at least in connection with Fig. 5. In at least one embodiment, after step 606, the process shown in block diagram 600 for executing an API to create a subcontext continues in step 608.

[0111] In at least one embodiment, in step 608 of the process of performing a subcontext creation API, as shown in block diagram 600, a processor performing the process performs one or more operations to determine whether a subcontext is to be created based at least in part on a list of resources and / or one or more flags (e.g., received as arguments to a subcontext creation API, as described herein at least in connection with Fig. 5). In at least one embodiment, in one or more operations to determine whether a subcontext can be created, this determination is made based at least in part on available resources, a source context (e.g., from which a subcontext is to be created), a device on which the context is to be created, etc. In at least one embodiment, after step 608, the process continues to execute an API to create a subcontext shown in block diagram 600 in step 6104.

[0112] In at least one embodiment, in step 610 of the process of performing an API to create a subcontext shown in block diagram 600, a processor performing the process performs one or more operations to determine whether a subcontext can be created (e.g., based on a determination in step 608). In at least one embodiment, if it is determined in step 604 that a subcontext can be created ("YES" branch), the process of performing an API to create a subcontext shown in block diagram 600 continues in step 612. In at least one embodiment, if it is determined in step 610 that a subcontext cannot be created ("NO" branch), the process of performing an API to create a subcontext shown in block diagram 600 continues in step 620, which is described below.

[0113] In at least one embodiment, in step 612 of the process depicted in block diagram 600 for performing a subcontext creation API, a processor performing this process performs one or more operations to create a subcontext based on resources and flags received as arguments to a subcontext creation API. In at least one embodiment, in step 612, a subcontext is created, as described herein at least in connection with Fig. 1 and Fig. 2. In at least one embodiment, after step 612, the process shown in block diagram 600 for executing an API to create a subcontext continues in step 614.

[0114] In at least one embodiment, in step 614 of the process shown in block diagram 600 for performing a subcontext creation API, a processor performing the process performs one or more operations to assign a subset of a set of resources (e.g., a primary context and / or a subcontext) to a created subcontext. In at least one embodiment, in step 614, a subset of a set of resources is assigned exclusively to a subcontext (e.g., a single context). In at least one embodiment, a subset of a set of resources is assigned to a subcontext, for example, by making a list of the subset available to the subcontext. In at least one embodiment, after step 614, the process shown in block diagram 600 for performing a subcontext creation API continues in step 616.

[0115] In at least one embodiment, in step 616 of the process for performing an API for creating a subcontext depicted in block diagram 600, a processor performing the process performs one or more operations to indicate a success indicator (e.g., a success indicator 522, herein at least in connection with Fig. 5). In at least one embodiment, in step 616, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 616, the process for executing an API to create a subcontext shown in block diagram 600 is terminated. In at least one embodiment, which is described in Fig. 6, the process for executing an API for creating a subcontext shown in block diagram 600 continues after step 616 in step 602 described above.

[0116] In at least one embodiment, in step 618 of the process of performing an API to create a subcontext depicted in block diagram 600, a processor performing the process performs one or more operations to indicate an error indication (e.g., an error indication 524, herein at least in connection with Fig. 5). In at least one embodiment, in step 618, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 618, the process for executing an API to create a subcontext shown in block diagram 600 is terminated. In at least one embodiment, which is described in Fig. 6 is not shown, the process for executing an API for creating a subcontext shown in block diagram 600 continues after step 618 with step 602 described above.

[0117] In at least one embodiment, the operations of the process for performing an API to create a subcontext shown in block diagram 600 are performed in a different order than shown in FIG. 600. In at least one embodiment, the operations of the process for performing an API to create a subcontext shown in block diagram 600 are performed concurrently or in parallel. In at least one embodiment, operations of the process for performing an API to create a subcontext shown in block diagram 600 that are not dependent on one another (e.g., are order independent) are performed concurrently or in parallel.In at least one embodiment, the operations of the process for performing an API for creating a subcontext shown in block diagram 600 are performed by a plurality of threads executing on a processor such as those described herein.

[0118] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 5 and Fig. 6, such as one or more circuits for performing an application programming interface (API) for allocating one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to perform one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 5 and Fig. 6, such as one or more circuits for implementing an application programming interface (API) for generating one or more masks to generate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to implement one or more kernels and / or otherwise perform operations described herein. In at least one embodiment, the methods described herein in connection with Fig. 5 and Fig. 6 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to perform an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 5 and Fig. 6, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 5 and Fig. 6 is not shown, comprise one or more components which are used here in connection with Fig. 5 and Fig. 6, one or more components described herein in connection with Fig. 26-58 to perform an application programming interface (API) to allocate one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein.

[0119] Fig. 7 is a block diagram 700 illustrating an application programming interface (API) for destroying a subcontext in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a subcontext destroy API 702 to destroy a subcontext created by executing an API for creating a subcontext, such as a subcontext create API 502, described herein at least in conjunction with Fig. 5. In at least one embodiment described in Fig. 7, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform subcontext destroy API 702 to execute an application programming interface (API) to release one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads. In at least one embodiment, as described in Fig. 7, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a subcontext destroy API 702 to execute an application programming interface (API) to cause one or more masks to be disabled, where the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment, as described in Fig. 7, one or more circuits of a processor, such as those described herein, execute one or more subcontext destroy API 702 instructions to execute an application programming interface (API) to release one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads in response to receiving a second API, such as those described herein.

[0120] In at least one embodiment, when invoked, the subcontext destroy API 702 receives one or more arguments specifying information about operations to be performed using techniques such as those described herein. In at least one embodiment, when invoked, the subcontext destroy API 702 receives one or more arguments specifying information about commands to be executed using techniques such as those described herein.

[0121] In at least one embodiment, the subcontext destroy API 702 receives as input one or more arguments comprising a subcontext 704. In at least one embodiment, a subcontext 704 is a data value that includes information used to identify, indicate, or otherwise specify a subcontext to be destroyed with the subcontext destroy API 702. In at least one embodiment, a subcontext identified, indicated, or otherwise specified by subcontext 704 is one of a plurality of parameters that can be used by the subcontext destroy API 702 to destroy a subcontext.In at least one embodiment, subcontext 704 is a data value that identifies, indicates, or otherwise specifies to an API such as subcontext destroy API 702 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0122] In at least one embodiment, the subcontext destroy API 702 receives as input one or more arguments that include one or more other arguments 718. In at least one embodiment, the other arguments 718 are data that include information to indicate other information that can be used in executing the subcontext destroy API 702 to destroy a subcontext.

[0123] In at least one embodiment described in Fig. 7, a processor executes one or more instructions to execute one or more APIs, such as subcontext destroy API 702, to execute an application programming interface (API) to release one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads, using one or more arguments, including, but not limited to, subcontext 704 and / or other arguments 718. In at least one embodiment, shown in Fig. 7, a processor executes one or more instructions to execute one or more APIs, such as subcontext destroy API 702, to execute an application programming interface (API) to cause one or more masks to be disabled, where the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more arguments, including, but not limited to, subcontext 704 and / or other arguments 718.

[0124] In at least one embodiment, the subcontext destroy API 702, when invoked, causes one or more APIs such as one or more APIs 306, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the subcontext destroy API 702, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0125] In at least one embodiment, one or more APIs 306, in response to the subcontext destroy API 702, are to cause one or more processors to perform a subcontext destroy API return 720, if performed. In at least one embodiment, the subcontext destroy API return 720 is a set of instructions that, when executed, creates and / or indicates one or more data values in response to the subcontext destroy API 702. In at least one embodiment, the subcontext destroy API return 720 indicates a success indication 722. In at least one embodiment, the success indication 722 is data comprising any value to indicate the success of the subcontext destroy API 702.In at least one embodiment, success indicator 722 includes information indicating one or more specific types of successes generated as a result of performing subcontext destroy API 702. In at least one embodiment, success indicator 722 includes information indicating one or more other data values generated as a result of subcontext destroy API 702.

[0126] In at least one embodiment, an error indicator 724 is provided in the subcontext destroy API return 720. In at least one embodiment, the error indicator 724 is data comprising any value indicating the failure of the subcontext destroy API 702. In at least one embodiment, the error indicator 724 includes information indicating one or more specific types of errors generated as a result of performing the subcontext destroy API 702. In at least one embodiment, the error indicator 724 includes information indicating one or more other data values generated as a result of the subcontext destroy API 702.

[0127] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, the subcontext destroy API 702, which adds various types of operations to a stream performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include a semaphore acquisition operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more indications for submitting an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating submission of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with. Fig. 5 described.

[0128] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, subcontext destroy API 702, one or more function signatures that can be used to specify one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be executed. In at least one embodiment, one or more operations that cause one or more callback functions to be executed use software code, such as example software code that specifies a function signature for a callback function, as described herein at least in connection with Fig. 5 described.

[0129] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by subcontext destroy API 702 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which the one or more APIs 306 are to communicate the one or more operations.In at least one embodiment, one or more data structures of one or more APIs 306 usable for specifying one or more external devices for which the one or more APIs 306 are to communicate the one or more operations utilize software code, such as example software code that specifies a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with. Fig. 5 described.

[0130] In at least one embodiment, one or more data structures of one or more APIs 306 are used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code that indicates a data structure to indicate the type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with Fig. 5 described.

[0131] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions are to be executed in response to the subcontext destroy API 702, as described above.In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to the subcontext destroy API 702 use software code, such as example software code specifying a stream operation API call in the parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0132] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are similar to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to the subcontext destroy API 702, as described herein.In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code that indicates adding one or more operations or instructions to one or more executable graphs through one or more APIs 306 of the parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0133] Fig. 8 is a block diagram 800 illustrating a process for implementing an application programming interface (API) for destroying a subcontext in accordance with at least one embodiment. In at least one embodiment, the process for implementing an API for destroying a subcontext shown in block diagram 800 is a process for implementing the subcontext destroy API 702, which is illustrated here at least in connection with Fig. 7. In at least one embodiment, part or all of the process shown in block diagram 800 for performing an API for destroying a subcontext (or other processes described herein or variations and / or combinations thereof) is performed under the control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with Fig. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions that can be executed by one or more processors such as those described herein. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, a processor such as the one described herein, at least in conjunction with Fig. 1 performs one or more steps of the process to execute an API for destroying a subcontext shown in block diagram 800. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of the process to execute an API for destroying a subcontext shown in block diagram 800.

[0134] In at least one embodiment, a processor performing the process performs one or more operations to receive or otherwise obtain a subcontext destruction API shown in block diagram 800 at step 802 of the process. In at least one embodiment, in step 802, a subcontext destruction API is received or otherwise provided to a processor such as processor 110, which is illustrated here at least in connection with Fig. 1. In at least one embodiment, in step 802, an API for destroying a subcontext is received from a processor, such as processor 102, which is used here at least in conjunction with Fig. 1. In at least one embodiment, an API for destroying a subcontext in step 802 includes one or more arguments as described herein at least in connection with Fig. 7. In at least one embodiment, after step 802, the process continues to perform an API for destroying a subcontext shown in block diagram 800 in step 804.

[0135] In at least one embodiment, in step 804 of the process for performing a subcontext destruction API shown in block diagram 800, a processor performing the process performs one or more operations to determine whether a subcontext destruction API (e.g., received in step 802) is valid. In at least one embodiment, the operations for determining whether a subcontext destruction API is valid include, in step 804, operations for checking the validity of arguments of the subcontext destruction API (e.g., arguments described herein at least in connection with Fig. 7). In at least one embodiment, if a subcontext destruction API is determined to be valid (a "YES" branch) at step 804, the process of performing a subcontext destruction API as shown in block diagram 800 continues to step 806. In at least one embodiment, if a subcontext destruction API is determined to be invalid (a "NO" branch) at step 804, the process of performing a subcontext destruction API shown in block diagram 800 continues to step 816, described below.

[0136] In at least one embodiment, in step 806 of the process for executing an API for destroying a subcontext shown in block diagram 800, a processor performing the process performs one or more operations to identify a subcontext to be destroyed. In at least one embodiment, a subcontext is identified as an argument of a subcontext destroy API received in step 802. In at least one embodiment, after step 806, the process for executing an API for destroying a subcontext shown in block diagram 800 continues in step 808.

[0137] In at least one embodiment, in step 808 of the process for performing an API to destroy a subcontext shown in block diagram 800, a processor performing the process performs one or more operations to determine whether a subcontext to be destroyed has been identified (e.g., identified in step 806). In at least one embodiment, if it is determined in step 808 that a subcontext to be destroyed has been identified ("YES" branch), the process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 810. In at least one embodiment, if it is determined in step 808 that a subcontext to be destroyed has not been identified ("NO" branch), the process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 816, which is described below.

[0138] In at least one embodiment, in step 810 of the process for performing an API to destroy a subcontext shown in block diagram 800, a processor performing the process performs one or more operations to stop one or more streams associated with a context to be destroyed (e.g., a subcontext identified in step 806), release all resources associated with the context to be destroyed, and destroy the context to be destroyed (e.g., indicate that the context is not valid). In at least one embodiment, the streams are allowed to complete any ongoing operations before being stopped. In at least one embodiment, streams are stopped immediately (e.g., are not allowed to complete any ongoing operations). In at least one embodiment, a plurality of streams associated with a subcontext are stopped.In at least one embodiment, after step 810, the process continues to perform an API to destroy a subcontext shown in block diagram 800 in step 812.

[0139] In at least one embodiment, at step 812 of the process for performing an API to destroy a subcontext shown in block diagram 800, a processor performing the process performs one or more operations to determine whether a subcontext to be destroyed has been destroyed (e.g., by performing step 710). In at least one embodiment, if it is determined in step 812 that a context to be destroyed has been destroyed ("YES" branch), the process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 814. In at least one embodiment, if it is determined in step 812 that a context to be destroyed has not been destroyed ("NO" branch), the process for performing an API to destroy a subcontext shown in block diagram 800 continues in step 816, described below.

[0140] In at least one embodiment, at step 814 of the process for performing an API to destroy a subcontext depicted in block diagram 800, a processor performing the process performs one or more operations to provide a success indicator (e.g., a success indicator 722, described herein at least in connection with Fig. 7). In at least one embodiment, in step 814, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 814, the process for executing an API for destroying a subcontext shown in block diagram 800 is terminated. In at least one embodiment shown in Fig. 8, after step 814, the process for executing an API for destroying a subcontext shown in block diagram 800 continues in step 802 described above.

[0141] In at least one embodiment, in step 816 of the process for performing an API to destroy a subcontext indicated in block diagram 800, a processor performing the process performs one or more operations to return an error indication (e.g., an error indication 724, herein at least in connection with Fig. 7). In at least one embodiment, in step 816, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 816, the process for executing an API for destroying a subcontext shown in block diagram 800 is terminated. In at least one embodiment shown in Fig. 8, the process for executing an API for destroying a subcontext shown in block diagram 800 continues after step 816 in step 802 described above.

[0142] In at least one embodiment, the operations of the process for performing an API to destroy a subcontext shown in block diagram 800 are performed in a different order than shown in FIG. 800. In at least one embodiment, the operations of the process for performing an API to destroy a subcontext shown in block diagram 800 are performed concurrently or in parallel. In at least one embodiment, operations of the process for performing an API to destroy a subcontext shown in block diagram 800 that are not dependent on one another (e.g., are order independent) are performed concurrently or in parallel.In at least one embodiment, the operations of the process for performing an API for destroying a subcontext shown in block diagram 800 are performed by a plurality of threads executing on a processor such as those described herein.

[0143] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits to perform operations and / or instructions described herein in connection with Fig. 7 and Fig. 8, such as one or more circuits to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or other operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) may include one or more circuits to execute operations and / or instructions described herein in connection with Fig. 7 and Fig. 8, such as one or more circuits to execute an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for executing one or more kernels and / or otherwise performing operations described herein. In at least one embodiment, the methods described herein in connection with Fig. 7 and Fig. 8 described processes and / or commands Systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 7 and Fig. 8, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or otherwise operations described herein. In at least one embodiment described in Fig. 7 and Fig. 8 is not shown, comprise one or more components which are used here in conjunction with Fig. 7 and Fig. 8, one or more components described here in connection with Fig. 26-58 to execute an application programming interface (API) to expose one or more data structures to indicate which of one or more streaming multiprocessors (SMs) of one or more processors should be used to execute one or more software threads and / or other operations described herein.

[0144] Fig. 9 is a block diagram 900 illustrating an application programming interface (API) for obtaining a subcontext-from-stream API in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a subcontext-from-stream API 902 to obtain a subcontext associated with an identified stream executing using a channel, as described herein at least in connection with Fig. 1 and Fig. 2. In at least one embodiment described in Fig. 9, one or more circuits of a processor, such as those described herein, execute one or more instructions to obtain a subcontext of stream API 902 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored. In at least one embodiment, Fig. 9, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute Get Subcontext of Stream API 902 to execute an application programming interface (API) to specify one or more identifiers of one or more masks, where the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels. In at least one embodiment, also described in Fig. 9, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute Get Subcontext of Stream API 902 to cause an application programming interface (API) to store one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors in response to receiving a second API, such as those described herein.

[0145] In at least one embodiment, the Get Subcontext from Stream API 902, when invoked, receives one or more arguments that specify information about operations to be performed using techniques such as those described herein. In at least one embodiment, the Get Subcontext from Stream API 902, when invoked, receives one or more arguments that specify information about commands to be executed using techniques such as those described herein.

[0146] In at least one embodiment, Get Subcontext from Stream API 902 receives as input one or more arguments comprising a stream 904. In at least one embodiment, stream 904 is a data value that comprises information that can be used to identify, indicate, or otherwise specify a stream from which a subcontext is to be retrieved using Get Subcontext from Stream API 902. In at least one embodiment, a stream identified, indicated, or otherwise specified by stream 904 is one of a plurality of parameters that can be used by Get Subcontext from Stream API 902 to obtain a subcontext from a stream.In at least one embodiment, stream 904 is a data value that identifies, indicates, or otherwise specifies to an API such as Get Subcontext of Stream API 902 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0147] In at least one embodiment, Get Subcontext from Stream API 902 receives as input one or more arguments comprising a subcontext return 906. In at least one embodiment, subcontext return 906 is a data value comprising information that can be used to identify, indicate, or otherwise specify a storage location for storing a subcontext identified by Get Subcontext from Stream API 902. In at least one embodiment, a storage location for storing a subcontext identified, indicated, or otherwise specified by Subcontext return 906 is one of a plurality of parameters that can be used by Get Subcontext from Stream API 902 to obtain a subcontext from a stream.In at least one embodiment, subcontext return 906 is a data value for identifying, indicating, or otherwise specifying an API, such as Get Subcontext from Stream API 902, of a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0148] In at least one embodiment, Get Subcontext from Stream API 902 receives as input one or more arguments that include one or more other arguments 918. In at least one embodiment, other arguments 918 are data that include information to indicate other information that can be used in performing Get Subcontext from Stream API 902 to obtain a subcontext from a stream.

[0149] In at least one embodiment described in Fig. 9, a processor executes one or more instructions to execute one or more APIs, such as Get Subcontext from Stream API 902 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored using one or more arguments, including, but not limited to, stream 904, subcontext return 906, and / or other arguments 918. In at least one embodiment not shown in Fig. 9, a processor executes one or more instructions to execute one or more APIs, such as Get Subcontext of Stream API 902, to execute an application programming interface (API) to specify one or more identifiers of one or more masks, where the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) that can be used to execute one or more kernels, using one or more arguments, including, but not limited to, stream 904, subcontext return 906, and / or other arguments 918.

[0150] In at least one embodiment, the Get Subcontext of Stream API 902, when invoked, causes one or more APIs such as one or more APIs 306 described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the Get Subcontext of Stream API 902, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0151] In at least one embodiment, one or more APIs 306, when executed, are to cause one or more processors to perform a Get Subcontext of Stream API return 920 in response to the Get Subcontext of Stream API 902. In at least one embodiment, the Get Subcontext of Stream API 920 is a set of instructions that, when executed, generates and / or specifies one or more data values in response to the Get Subcontext of Stream API 902. In at least one embodiment, the Get Subcontext of Stream API return 920 specifies a success indication 922. In at least one embodiment, the Success Indicator 922 is data comprising any value indicating the success of the Get Subcontext of Stream API 902.In at least one embodiment, success indicator 922 includes information indicating one or more specific types of successes generated as a result of performing Get Subcontext from Stream API 902. In at least one embodiment, success indicator 922 includes information indicating one or more other data values generated as a result of Get Subcontext from Stream API 902.

[0152] In at least one embodiment, the Get Subcontext from Stream API return 920 indicates an error indicator 924. In at least one embodiment, the error indicator 924 is data comprising any value indicating the failure of the Get Subcontext from Stream API 902 query. In at least one embodiment, the error indicator 924 includes information indicating one or more specific types of errors generated as a result of performing the Get Subcontext from Stream API 902. In at least one embodiment, the error indicator 924 includes information indicating one or more other data values generated as a result of the Get Subcontext from Stream API 902.

[0153] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, Get Subcontext of Stream API 902, that add various operations of different types to a stream performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include a semaphore acquire operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more indications for submitting an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating submission of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with. Fig. 5 described.

[0154] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, Get Subcontext of Stream API 902, one or more function signatures that can be used to specify one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause execution of one or more callback functions. In at least one embodiment, one or more operations that cause execution of one or more callback functions use software code, such as example software code that specifies a function signature for a callback function, as described herein at least in connection with Fig. 5 described.

[0155] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by Get Subcontext of Stream API 902 to one or more APIs 306, one or more data structures from one or more APIs 306 are usable to specify one or more external devices for which the one or more APIs 306 are to communicate the one or more operations.In at least one embodiment, one or more data structures of one or more APIs 306 usable for specifying one or more external devices for which the one or more APIs 306 are to communicate the one or more operations utilize software code, such as example software code that specifies a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with. Fig. 5 described.

[0156] In at least one embodiment, one or more data structures of one or more APIs 306 are used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code that indicates a data structure to indicate the type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with Fig. 5 described.

[0157] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions executed by one or more accelerators in heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions are to be executed in response to obtaining the subcontext from the Get Stream API 902, as described above.In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other subset of instructions to be executed in response to Get Subcontext of Stream API 902 use software code, such as example software code specifying a Stream Operations API call in a parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0158] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are similar to the manner in which one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are to be added to one or more streams or instruction sets in response to Get Subcontext from Stream API 902, as described herein.In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code that indicates adding one or more operations or instructions to one or more executable graphs through one or more APIs 306 of the parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0159] Fig. 10 is a block diagram 1000 illustrating a process for implementing an application programming interface (API) for obtaining subcontext from a stream API in accordance with at least one embodiment. In at least one embodiment, the process illustrated in block diagram 1000 for implementing an API for obtaining subcontext from a stream is a process for obtaining subcontext from a stream API 902, which is illustrated here at least in connection with Fig. 9. In at least one embodiment, part or all of the process for performing an API for obtaining subcontext from a stream depicted in block diagram 1000 (or any other processes described herein or variations and / or combinations thereof) is performed under the control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with Fig. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions that can be executed by one or more processors such as those described herein. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, a processor such as the one described herein, at least in conjunction with Fig. 1, processor 110 performs one or more steps of the process to execute an API for obtaining subcontext from a stream shown in block diagram 1000. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of the process to execute an API for obtaining subcontext from a stream shown in block diagram 1000.

[0160] In at least one embodiment, in step 1002 of the process for performing an API for obtaining subcontext from a stream illustrated in block diagram 1000, a processor performing the process performs one or more operations to receive or otherwise obtain an API for obtaining subcontext from a stream. In at least one embodiment, in step 1002, an API for obtaining subcontext from stream API is received or otherwise provided to a processor such as processor 110, which is illustrated here at least in connection with Fig. 1. In at least one embodiment, in step 1002, an API for obtaining subcontext from Stream API is received from a processor such as processor 102, which is described here at least in conjunction with Fig. 1. In at least one embodiment, an API for obtaining subcontext from Stream API in step 1002 comprises one or more arguments as described herein at least in connection with Fig. 9. In at least one embodiment, after step 1002, the process of executing an API to obtain subcontext from Stream API, as shown in block diagram 1000, continues in step 1004.

[0161] In at least one embodiment, in step 1004 of the process shown in block diagram 1000 for executing an API for obtaining subcontext from stream API, a processor performing the process performs one or more operations to determine whether an API for obtaining subcontext from stream API (e.g., received in step 1002) is valid. In at least one embodiment, the operations for determining whether an API for obtaining subcontext from stream API is valid include, in step 1004, operations for checking the validity of arguments of the API for obtaining subcontext from stream API (e.g., arguments described herein at least in connection with Fig. 9). In at least one embodiment, if it is determined in step 1004 that an API for obtaining subcontext from a stream is valid ("YES" branch), the process of performing an API for obtaining subcontext from a stream as shown in block diagram 1000 continues in step 1006. In at least one embodiment, if it is determined in step 1004 that an API for obtaining subcontext from a stream is not valid ("NO" branch), the process of performing an API for obtaining subcontext from a stream shown in block diagram 1000 continues in step 1014 described below.

[0162] In at least one embodiment, in step 1006 of the process of performing an API to obtain subcontext from a stream shown in block diagram 1000, a processor performing the process performs one or more operations to identify a subcontext of a stream (e.g., a stream specified by an argument to an API to obtain subcontext from a stream). In at least one embodiment, a subcontext maintains a list of streams associated with the named context. In at least one embodiment, each stream maintains a list of subcontexts associated with that stream. In at least one embodiment, after step 1006, the process of performing an API to obtain subcontext from stream API, as shown in block diagram 1000, continues in step 1008.

[0163] In at least one embodiment, in step 1008 of the process of performing an API to obtain subcontext from a stream depicted in block diagram 1000, a processor performing the process performs one or more operations to determine whether a subcontext has been identified (e.g., in step 1006). In at least one embodiment, in step 1008, if it is determined that a subcontext has been identified ("YES" branch), the process of performing an API to obtain subcontext from a stream depicted in block diagram 1000 continues in step 1010. In at least one embodiment, in step 1008, if it is determined that a subcontext has not been identified ("NO" branch), the process of performing an API to obtain subcontext from a stream depicted in block diagram 1000 continues in step 1014, which is described below.

[0164] In at least one embodiment, in step 1010 of the process of performing an API for obtaining a subcontext from a stream illustrated in block diagram 1000, a processor performing the process performs one or more operations to store an identified subcontext (e.g., identified in step 1006) at a memory location specified by an API for obtaining subcontext from a stream (e.g., received as an argument to the API for obtaining subcontext from a stream received in step 1002). In at least one embodiment, after step 1010, the process of performing an API for obtaining subcontext from stream API, as shown in block diagram 1000, continues in step 1012.

[0165] In at least one embodiment, in step 1012 of the process of performing an API for obtaining subcontext from a stream depicted in block diagram 1000, a processor performing the process performs one or more operations to return a success indication (e.g., a success indication 922, herein at least in connection with Fig. 9). In at least one embodiment, in step 1012, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 1012, the process for executing an API for obtaining subcontext from a stream API, as shown in block diagram 1000, is terminated. In at least one embodiment described in Fig. 6, after step 1012, the process of performing an API for obtaining subcontext from a stream, illustrated in block diagram 1000, continues in step 1002 described above.

[0166] In at least one embodiment, at step 1014 of the process of performing an API for obtaining subcontext from a stream depicted in block diagram 1000, a processor performing the process performs one or more operations to provide an error indication (e.g., an error indication 924, herein at least in connection with Fig. 9). In at least one embodiment, in step 1014, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in conjunction with Fig. 1). In at least one embodiment, after step 1014, the process for performing an API for obtaining subcontext from a stream API, as shown in block diagram 1000, is terminated. In at least one embodiment described in Fig. 10, after step 1014, the process of performing an API for obtaining subcontext from a stream, illustrated in block diagram 1000, continues in step 1002 described above.

[0167] In at least one embodiment, the operations of the process for performing an API for obtaining subcontext from a stream shown in block diagram 1000 are performed in a different order than shown in FIG. 1000. In at least one embodiment, the operations of the process for performing an API for obtaining subcontext from a stream shown in block diagram 1000 are performed concurrently or in parallel. In at least one embodiment, operations of the process for performing an API for obtaining subcontext from a stream shown in block diagram 1000 that are not dependent on one another (e.g., are order independent) are performed concurrently or in parallel.In at least one embodiment, operations of the process for performing an API for obtaining subcontext from a stream shown in block diagram 1000 are performed by a plurality of threads executing on a processor such as those described herein.

[0168] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 9 and Fig. 10, such as one or more circuits for performing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) may include one or more circuits to perform the operations described herein in connection with Fig. 9 and Fig. 10, such as one or more circuits to execute an application programming interface (API), to specify one or more identifiers of one or more masks, wherein the one or more masks specify one or more subsets of streaming multiprocessors of one or more graphics processing units (GPUs) usable for executing one or more kernels and / or otherwise performing operations described herein. In at least one embodiment, the methods described herein in connection with Fig. 9 and Fig. 10 described processes and / or commands systems, methods, processes and / or commands described herein in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or to otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 9 and Fig. 10, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 9 and Fig. 10 is not shown, comprise one or more components which are used here in connection with Fig. 9 and Fig. 10, one or more components described herein in connection with Fig. 26-58 to perform an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) of one or more processors to be stored and / or otherwise perform operations described herein.

[0169] Fig. 11 is a block diagram 1100 illustrating an application programming interface (API) for obtaining resources associated with a context, in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a context event obtain API 1102 to retrieve resources associated with a context. In at least one embodiment, Fig. 11, one or more circuits of a processor, such as those described herein, execute one or more instructions to implement an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available for executing one or more software threads. In at least one embodiment, as described in Fig. 11, one or more circuits of a processor, such as those described herein, execute one or more instructions to implement an application programming interface (API) 1102 to specify one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for executing one or more software kernels. In at least one embodiment, also not shown in Fig. 11, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute an application programming interface (API) 1102 to specify one or more streaming multiprocessors (SMs) of one or more processors available to be used to execute one or more software threads in response to receiving a second API, such as those described herein.

[0170] In at least one embodiment, the Get Context Resource API 1102, when invoked, receives one or more arguments specifying information about the operations to be performed using techniques such as those described herein. In at least one embodiment, the Get Context Resource API 1102, when invoked, receives one or more arguments specifying information about commands to be executed using techniques such as those described herein.

[0171] In at least one embodiment, the Get Context Resource API 1102 receives as input one or more arguments comprising a resource return 1104. In at least one embodiment, the resource return 1104 is a data value comprising information that can be used to identify, indicate, or otherwise specify a location for storing resources (e.g., a resource descriptor) associated with a context using the Get Context Resource API 1102. In at least one embodiment, a location for storing resources identified, indicated, or otherwise specified by the resource return 1104 is one of a plurality of parameters that can be used by the Get Context Resource API 1102 to obtain resources associated with a context.In at least one embodiment, the resource return 1104 is a data value that identifies, indicates, or otherwise specifies to an API such as the Get Context Resource API 1102 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0172] In at least one embodiment, the Get Context Resource API 1102 receives as input one or more arguments comprising a context 1106. In at least one embodiment, the context 1106 is a data value comprising information that can be used to identify, indicate, or otherwise specify a context or subcontext from which a set of resources can be identified and returned using the Get Context Resource API 1102. In at least one embodiment, a context or subcontext identified, indicated, or otherwise specified by context 1106 is one of a plurality of parameters that can be used by the Get Context Resource API 1102 to obtain resources associated with a context.In at least one embodiment, context 1106 is a data value that identifies, indicates, or otherwise specifies to an API such as Get Context Resource API 1102 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0173] In at least one embodiment, the Get Context Resource API 1102 receives as input one or more arguments that include one or more other arguments 1118. In at least one embodiment, the other arguments 1118 are data that include information to indicate other information that can be used in executing the Get Context Resource API 1102 to obtain resources associated with a context.

[0174] In at least one embodiment described in Fig. 11, a processor executes one or more instructions to execute one or more APIs such as Get Context Resource API 1102 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) from one or more processors available to be used to execute one or more software threads, using one or more arguments, including, but not limited to, resource return 1104, context 1106, and / or other arguments 1118. In at least one embodiment, shown in Fig. 11, a processor executes one or more instructions to execute one or more APIs, such as Get Context Resource API 1102 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) that may be used to execute one or more software kernels using one or more arguments, including, but not limited to, resource return 1104, context 1106, and / or other arguments 1118.

[0175] In at least one embodiment, Get Context Resource API 1102, when invoked, causes one or more APIs such as one or more APIs 306 described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the Get Context Resource API 1102, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0176] In at least one embodiment, in response to Get Context Resource API 1102, one or more APIs 306, if executed, are to cause one or more processors to perform a Get Context Resource API return 1120. In at least one embodiment, Get Context Resource API return 1120 is a set of instructions that, when executed, generates and / or indicates one or more data values in response to Get Context Resource API 1102. In at least one embodiment, Get Context Resource API return 1120 indicates a success indication 1122. In at least one embodiment, Success Indicator 1122 is data comprising any value to indicate the success of Get Context Resource API 1102.In at least one embodiment, success indicator 1122 includes information indicating one or more specific types of successes generated as a result of performing Get Context Resource API 1102. In at least one embodiment, success indicator 1122 includes information indicating one or more other data values generated as a result of Get Context Resource API 1102.

[0177] In at least one embodiment, the Get Context Resource API return 1120 indicates an error indicator 1124. In at least one embodiment, the error indicator 1124 is data comprising any value to indicate an error of the Get Context Resource API 1102. In at least one embodiment, the error indicator 1124 includes information indicating one or more specific types of errors generated as a result of performing the Get Context Resource API 1102. In at least one embodiment, the error indicator 1124 includes indications indicating one or more other data values generated as a result of the Get Context Resource API 1102.

[0178] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, Get Context Resource API 1102, that add various operations of different types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include an acquire semaphore operation. In at least one embodiment, the stream operations include a release semaphore operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, such as L2 cache of a PPU, such as a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more operations to indicate the delivery of an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating the delivery of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with. Fig. 5 described.

[0179] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, the Get Context Resource API 1102, one or more function signatures that can be used to specify one or more callback functions for operations to be performed by one or more accelerators in heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be executed. In at least one embodiment, one or more operations that cause one or more callback functions to be executed use software code, such as example software code that specifies a function signature for a callback function, as described herein at least in connection with Fig. 5 described.

[0180] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by Get Context Resource API 1102 to one or more APIs 306, one or more data structures from one or more APIs 306 are usable to specify one or more external devices for which the one or more APIs 306 are to communicate the one or more operations.In at least one embodiment, one or more data structures of one or more APIs 306 usable for specifying one or more external devices for which the one or more APIs 306 are to communicate the one or more operations utilize software code, such as example software code that specifies a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with. Fig. 5 described.

[0181] In at least one embodiment, one or more data structures of one or more APIs 306 are used to specify the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more data structures of one or more APIs 306 used to specify the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code specifying a data structure to specify the type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with Fig. 5 described.

[0182] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to be executed in response to the Get Context Resource API 1102, as described above.In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to Get Context Resource API 1102 use software code, such as example software code specifying a stream operations API call in a parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0183] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are comparable to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are added to one or more streams or instruction sets in response to the Get Context Resource API 1102, as described herein.In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be performed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code that indicates adding one or more operations or instructions to one or more executable graphs through one or more APIs 306 of the parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0184] Fig. 12 is a block diagram 1200 illustrating a process for performing an application programming interface (API) to obtain context resources in accordance with at least one embodiment. In at least one embodiment, the process illustrated in block diagram 1200 for performing an API for retrieving resources associated with a context is a process for performing Get Context Resource API 1102, which is illustrated here at least in connection with Fig. 11. In at least one embodiment, part or all of the process for performing an API for retrieving resources associated with a context depicted in block diagram 1200 (or other processes described herein or variations and / or combinations thereof) is performed under the control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with Fig. 26-58, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions that can be executed by one or more processors such as those described herein. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, a processor such as the one described herein, at least in connection with Fig. 1, performs one or more steps of the process to execute an API to obtain resources associated with a context shown in block diagram 1200. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of the process to execute an API to obtain resources associated with a context shown in block diagram 1200.

[0185] In at least one embodiment, in step 1202 of the process for performing an API for retrieving resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to receive or otherwise obtain an API for retrieving resources associated with a context. In at least one embodiment, in step 1202, an API for retrieving resources associated with a context is received or otherwise provided to a processor, such as processor 110, depicted here at least in connection with Fig. 1. In at least one embodiment, an API for retrieving resources associated with a context is provided in step 1202 from a processor such as processor 102, which is used here at least in conjunction with Fig. 1 described in step 1202. In at least one embodiment, an API for retrieving context resources includes one or more arguments as described herein at least in connection with Fig. 11. In at least one embodiment, after step 1202, the process continues to step 1204 to perform an API for retrieving resources associated with a context shown in block diagram 1200.

[0186] In at least one embodiment, in step 1204 of the process for performing an API for retrieving resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine whether an API for retrieving resources associated with a context (e.g., received in step 1202) is valid. In at least one embodiment, the operations for determining whether an API for retrieving resources associated with a context is valid include, in step 1204, operations for checking the validity of arguments of the API for retrieving resources associated with a context (e.g., arguments depicted here at least in connection with Fig. 11). In at least one embodiment, if it is determined in step 1204 that an API for retrieving resources associated with a context is valid ("YES" branch), the process of performing an API for retrieving resources associated with a context, as shown in block diagram 1200, continues in step 1206. In at least one embodiment, if it is determined in step 1204 that an API for retrieving resources associated with a context is not valid ("NO" branch), the process of performing an API for retrieving resources associated with a context shown in block diagram 1200 continues in step 1214, described below.

[0187] In at least one embodiment, in step 1206 of the process of performing an API to retrieve resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine a request type of an API to retrieve resources associated with a context (e.g., received in step 1202). In at least one embodiment, a request type may be to obtain resources associated with a context. In at least one embodiment, a request type may be to obtain resources associated with a subcontext. In at least one embodiment, a request type may be to retrieve resources associated with a device.In at least one embodiment, a request type is received as an argument to an API to retrieve resources associated with a context (e.g., received in step 1202). In at least one embodiment, after step 1206, the process continues to perform an API to retrieve resources associated with a context shown in block diagram 1200, in step 1208.

[0188] In at least one embodiment, in step 1208 of the process for performing an API to retrieve resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine whether a request type (e.g., determined in step 1206) is to retrieve resources associated with a context. In at least one embodiment, if it is determined in step 1208 that a request type is to retrieve resources associated with a context ("YES" branch), the process for performing an API to retrieve resources associated with a context depicted in block diagram 1200 continues in step 1216.In at least one embodiment, if it is determined in step 1208 that a request type is not intended to obtain resources associated with a context ("NO" branch), the process of performing an API to obtain resources associated with a context depicted in block diagram 1200 continues in step 1210.

[0189] In at least one embodiment, in step 1210 of the process for performing an API to obtain resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine that a request type (e.g., determined in step 1206) is to obtain resources associated with a subcontext. In at least one embodiment, if it is determined in step 1210 that a request type is to obtain resources associated with a subcontext ("YES" branch), the process for performing an API to obtain resources associated with a context depicted in block diagram 1200 continues in step 1222.In at least one embodiment, if it is determined in step 1210 that a request type is not intended to obtain resources associated with a subcontext ("NO" branch), the process of performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1212.

[0190] In at least one embodiment, in step 1212 of the process for performing an API to retrieve resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine whether a request type (e.g., determined in step 1206) is to retrieve resources associated with a device. In at least one embodiment, if it is determined in step 1212 that a request type is to retrieve resources associated with a device ("YES" branch), the process for performing an API to retrieve resources associated with a context depicted in block diagram 1200 continues in step 1228.In at least one embodiment, if it is determined in step 1212 that a request type is not to obtain resources associated with a device ("NO" branch), the process of performing an API to obtain resources associated with a context depicted in block diagram 1200 continues in step 1214.

[0191] In at least one embodiment, in step 1214 of the process of performing an API for retrieving resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to return an error indicator (e.g., an error indicator 1124, herein at least in connection with Fig. 11). In at least one embodiment, in step 1214, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 1214, the process for performing an API for retrieving resources associated with a context shown in block diagram 1200 is terminated. In at least one embodiment shown in Fig. 12, after step 614, the process of executing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1202 described above.

[0192] In at least one embodiment, in step 1216 of the process for performing an API to retrieve resources associated with a context shown in block diagram 1200, a processor performing the process performs one or more operations to identify a context (e.g., a context received as an argument to an API to retrieve resources associated with a context received in step 1202) using the systems, methods, and operations described herein. In at least one embodiment, after step 1216, the process for performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1218.

[0193] In at least one embodiment, in step 1218 of the process of performing an API to retrieve resources associated with a context shown in block diagram 1200, a processor performing the process performs one or more operations to determine whether a context has been identified (e.g., in step 1216). In at least one embodiment, if it is determined in step 1218 that a context has been identified ("YES" branch), the process of performing an API to retrieve resources associated with a context shown in block diagram 1200 continues to step 1220. In at least one embodiment, if it is determined in step 1204 that a context has not been identified ("NO" branch), the process of performing an API to obtain resources associated with a context shown in block diagram 1200 continues to step 1214, as described above.

[0194] In at least one embodiment, in step 1220 of the process of performing an API to retrieve resources associated with a context shown in block diagram 1200, a processor performing the process performs one or more operations to store an identified context (e.g., identified in step 1216). In at least one embodiment, in step 1220, a processor performs one or more operations to store an identified context in a memory location received as an argument to an API to obtain resources associated with a context (e.g., received in step 1202). In at least one embodiment, after step 1220, the process of performing an API to retrieve resources associated with a context shown in block diagram 1200 continues in step 1230.

[0195] In at least one embodiment, in step 1222 of the process of performing an API to obtain resources associated with a context shown in block diagram 1200, a processor performing the process performs one or more operations to identify a subcontext (e.g., a subcontext received as an argument to an API to obtain resources associated with a context received in step 1202), as described herein. In at least one embodiment, after step 1222, the process of performing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1224.

[0196] In at least one embodiment, in step 1224 of the process of performing an API to obtain resources associated with a context depicted in block diagram 1200, a processor performing the process performs one or more operations to determine whether a subcontext has been identified (e.g., in step 1222). In at least one embodiment, if it is determined in step 1224 that a subcontext has been identified ("YES" branch), the process of performing an API to obtain resources associated with a context depicted in block diagram 1200 continues to step 1226.In at least one embodiment, if it is determined in step 1224 that a subcontext has not been identified (“NO” branch), the process of performing an API to obtain resources associated with a context depicted in block diagram 1200 continues in step 1214 described above.

[0197] In at least one embodiment, in step 1226, a processor performing this process performs one or more operations to store an identified subcontext (e.g., identified in step 1222) of the process for performing an API to obtain resources associated with a context shown in block diagram 1200. In at least one embodiment, in step 1226, a processor performs one or more operations to store an identified subcontext in a memory location received as an argument to an API to obtain resources associated with a context (e.g., received in step 1202). In at least one embodiment, after step 1226, the process for performing an API to retrieve resources associated with a context shown in block diagram 1200 continues in step 1230.

[0198] In at least one embodiment, in step 1228 of the process for performing an API for retrieving resources associated with a context shown in block diagram 1200, a processor performing the process performs one or more operations to store a device context (e.g., a primary or current context of a device such as device 202, described here at least in connection with Fig. 2). In at least one embodiment, in step 1228, a processor performs one or more operations to store a device context received as an argument to an API to obtain resources associated with a context (e.g., received in step 1202) in a memory location. In at least one embodiment, in step 1228, a device context that is a primary, default, or current context of a device is stored. In at least one embodiment, after step 1228, the process continues to perform an API to retrieve resources associated with a context shown in block diagram 1200, in step 1230.

[0199] In at least one embodiment, at step 1230 of the process of performing an API to obtain resources associated with a context indicated in block diagram 1200, a processor performing the process performs one or more operations to return a success indication (e.g., a success indication 1122, herein at least in connection with Fig. 11). In at least one embodiment, in step 1230, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 1230, the process for performing an API for retrieving resources associated with a context shown in block diagram 1200 is terminated. In at least one embodiment shown in Fig. 12, after step 1230, the process of executing an API to obtain resources associated with a context shown in block diagram 1200 continues in step 1202 described above.

[0200] In at least one embodiment, the operations of the process for performing an API for retrieving context resources shown in block diagram 1200 are performed in a different order than that shown in FIG. 1200. In at least one embodiment, the operations of the process for performing an API for retrieving context resources shown in block diagram 1200 are performed concurrently or in parallel. In at least one embodiment, operations of the process for performing an API for retrieving context resources shown in block diagram 1200 that are not dependent on one another (e.g., are order independent) are performed concurrently or in parallel.In at least one embodiment, the operations of the process for performing an API for retrieving context resources shown in block diagram 1200 are performed by a plurality of threads executing on a processor such as those described herein.

[0201] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 11 and Fig. 12, such as one or more circuits for performing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more available processors that may be used to perform one or more software threads and / or to otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) include one or more circuits for performing operations described herein in connection with Fig. 11 and Fig. 12, such as one or more circuits for implementing an application programming interface (API) for specifying one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for use in executing one or more software kernels and / or otherwise performing operations described herein. In at least one embodiment, the circuits described herein in connection with the Fig. 11 and Fig. 12 described processes and / or commands that are used here in connection with the Fig. 1-25, systems, methods, acts, and / or instructions for performing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available to execute one or more software threads and / or otherwise perform the acts described herein. In at least one embodiment, components described herein in connection with Fig. 11 and Fig. 12, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 11 and Fig. 12 is not shown, comprise one or more components which are described here in connection with Fig. 11 and Fig. 12, one or more components described here in connection with Fig. 26-58 to execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that may be used to execute one or more software threads and / or otherwise perform operations described herein.

[0202] Fig. 13 is a block diagram 1300 illustrating an application programming interface (API) for partitioning context resources in accordance with at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a context resource partitioning API 1302 to partition context resources according to one or more received partitioning criteria. In at least one embodiment, Fig. 13, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a context resource partitioning API 1302 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used. In at least one embodiment, as described in Fig. 13, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a context resource partitioning API 1302 to execute an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor that can be used to execute one or more software programs. In at least one embodiment, as described in Fig. 13, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a context resource partitioning API 1302 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used in response to receiving a second API, such as those described herein.

[0203] In at least one embodiment, when invoked, the context resource partitioning API 1302 receives one or more arguments specifying information about operations to be performed using techniques such as those described herein. In at least one embodiment, when invoked, the context resource partitioning API 1302 receives one or more arguments specifying information about commands to be executed using techniques such as those described herein.

[0204] In at least one embodiment, the context resource subdivision API 1302 receives as input one or more arguments comprising a subdivision resource return 1304. In at least one embodiment, the resource return 1304 is a data value comprising information that can be used to identify, indicate, or otherwise specify a storage location where a subdivision context resource list is to be stored as a result of using the context resource subdivision API 1302. In at least one embodiment, a subdivision resource list return identified, indicated, or otherwise specified by the resource return 1304 is one of a plurality of parameters that can be used by the context resource subdivision API 1302.In at least one embodiment, the partitioned context resource list API return 1304 is a data value that identifies, indicates, or otherwise specifies to an API, such as the partitioned context resource API 1302, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.

[0205] In at least one embodiment, the context resource subdivide API 1302 receives as input one or more arguments comprising a remaining resource list return 1306. In at least one embodiment, the remaining resource list return 1306 is a data value comprising information usable for a memory location in which to store a remaining resource list (e.g., a list of resources remaining after subdividing) as a result of using the context resource subdivide API 1302. In at least one embodiment, a resource return identified, indicated, or otherwise specified by the resource return 1306 is one of a plurality of parameters that can be used by the context resource subdivide API 1302 to subdivide context resources.In at least one embodiment, the API return 1306 is a data value that identifies, indicates, or otherwise specifies to an API, such as the context resource partitioning API 1302, a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.

[0206] In at least one embodiment, context resource subdivision API 1302 receives as input one or more arguments comprising an input resource list 1308. In at least one embodiment, input resource list 1308 is a data value comprising information that can be used to identify, indicate, or otherwise specify a list of resources to be subdivision using context resource subdivision API 1302. In at least one embodiment, an input resource list identified, indicated, or otherwise specified by input resource list 1308 is one of a plurality of parameters that can be used by context resource subdivision API 1302 to subdivision context resources.In at least one embodiment, input resource list 1308 is a data value for identifying, indicating, or otherwise specifying a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein, to an API, such as context resource partitioning API 1302.

[0207] In at least one embodiment, the context resource partitioning API 1302 receives as input one or more arguments comprising flags 1310. In at least one embodiment, flags 1310 is a data value comprising information that can be used to identify, indicate, or otherwise specify one or more flags that indicate at least which resources (e.g., SMs) of the input resource list 1308 are to be allocated to a partitioned resource list (e.g., stored in the partitioned resource list return 1304) and which remain (e.g., stored in the remaining resource list return 1306) when using the context resource partitioning API 1302.In at least one embodiment, the flags identified, indicated, or otherwise specified by flags 1310 are one of a variety of parameters that may be used by context resource partitioning API 1302 to partition context resources. In at least one embodiment, flags 1310 is a data value that identifies, indicates, or otherwise specifies to an API such as context resource partitioning API 1302 a set of operations or instructions to be executed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein.

[0208] In at least one embodiment, the context resource partitioning API 1302 receives as input one or more arguments comprising a minimum count 1312. In at least one embodiment, the minimum count 1312 is a data value comprising information that can be used to identify, indicate, or otherwise specify a minimum number (e.g., a smallest or minimum number of partitioned resources in a partitioned resource list) of partitioned resources obtained using the context resource partitioning API 1302. In at least one embodiment, a minimum count identified, indicated, or otherwise specified by the minimum count 1312 is one of a plurality of parameters that can be used by the context resource partitioning API 1302 to partition context resources.In at least one embodiment, the minimum number 1312 is a data value for identifying, indicating, or otherwise specifying a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor, as described herein, to an API, such as context resource partitioning API 1302.

[0209] In at least one embodiment, the context resource partitioning API 1302 receives as input one or more arguments that include one or more other arguments 1318. In at least one embodiment, the other arguments 1318 are data that include information to indicate other information that can be used in performing the context resource partitioning API 1302.

[0210] In at least one embodiment not described in Fig. 13, a processor executes one or more instructions to execute one or more APIs, such as context resource partitioning API 1302, to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) to be stored by one or more processors according to a plurality of groups in which to use the corresponding plurality of SMs using one or more arguments, including, but not limited to, a partitioned resource resource list return 1304, a remaining resource resource list return 1306, an input resource list 1308, flags 1310, a minimum number 1312, and / or other arguments 1318. In at least one embodiment not described in Fig. 13, a processor executes one or more instructions to execute one or more APIs, such as the context resource subdivision API 1302, to execute an application programming interface (API) to generate an indication of one or more subsets of a processor's streaming multiprocessors (SMs) available to be used to execute one or more software cores, using one or more arguments, including, but not limited to, subdivided resource return 1304, remaining resource return 1306, input resource list 1308, flags 1310, minimum count 1312, and / or other arguments 1318.

[0211] In at least one embodiment, the context resource partitioning API 1302, when invoked, causes one or more APIs such as one or more APIs 306, described herein at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the context resource partitioning API 1302, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described here at least in connection with Fig. 3, to add, insert, or otherwise include one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0212] In at least one embodiment, in response to the context resource partitioning API 1302, one or more APIs 306, when executed, are to cause one or more processors to perform a context resource partitioning API return 1320. In at least one embodiment, the context resource partitioning API return 1320 is a set of instructions that, when executed, generates and / or indicates one or more data values in response to the context resource partitioning API 1302. In at least one embodiment, the context resource partitioning API return 1320 indicates a success indicator 1322. In at least one embodiment, the success indicator 1322 is data comprising any value to indicate the success of the context resource partitioning API 1302.In at least one embodiment, success indicator 1322 includes information indicating one or more specific types of successes generated as a result of performing context resource subdivision API 1302. In at least one embodiment, success indicator 1322 includes information indicating one or more other data values generated as a result of context resource subdivision API 1302.

[0213] In at least one embodiment, the context resource subdivision API return 1320 indicates an error indicator 1324. In at least one embodiment, the error indicator 1324 is data comprising any value indicating the failure of the context resource subdivision API 1302. In at least one embodiment, the error indicator 1324 includes information indicating one or more specific types of errors generated as a result of performing the context resource subdivision API 1302. In at least one embodiment, the error indicator 1324 includes indications indicating one or more other data values generated as a result of the context resource subdivision API 1302.

[0214] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, context resource partitioning API 1302, that add various operations of different types to a stream executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, stream operations include a get semaphore operation. In at least one embodiment, stream operations include a release semaphore operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, such as L2 cache of a PPU, such as a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more indications for submitting an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating submission of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with. Fig. 5 described.

[0215] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, the context resource partitioning API 1302, one or more function signatures that can be used to specify one or more callback functions for operations performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be executed. In at least one embodiment, one or more operations that cause one or more callback functions to be executed use software code, such as example software code that specifies a function signature for a callback function, as described herein at least in connection with Fig. 5 described.

[0216] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by the context resource partitioning API 1302 to one or more APIs 306, one or more data structures of one or more APIs 306 are usable to specify one or more external devices for which the one or more APIs 306 are to communicate the one or more operations.In at least one embodiment, one or more data structures of one or more APIs 306 usable for specifying one or more external devices for which the one or more APIs 306 are to communicate the one or more operations utilize software code, such as example software code that specifies a data structure representing a device node for one or more accelerators within heterogeneous processors, as described herein at least in connection with. Fig. 5 described.

[0217] In at least one embodiment, one or more data structures of one or more APIs 306 are used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more data structures of one or more APIs 306 used to indicate the type and data of one or more operations specified by one or more operations to be performed by one or more accelerators within heterogeneous processors use software code, such as example software code that indicates a data structure to indicate the type and data of one or more operations to be performed by one or more accelerators within heterogeneous processors, as described herein at least in connection with Fig. 5 described.

[0218] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be added to a stream or other set of instructions to be executed by one or more accelerators within heterogeneous processors. In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions are to be executed in response to the context resource partitioning API 1302 described above.In at least one embodiment, instructions that cause one or more operations or instructions to be added to a stream or other set of instructions to be executed in response to the context resource partitioning API 1302 use software code, such as example software code specifying a stream operation API call in the parallel computing environment 308, as described at least in connection with. Fig. 5 described.

[0219] In at least one embodiment, one or more APIs 306 include instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs. In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs are comparable to how one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors are added to one or more streams or instruction sets in response to the context resource partitioning API 1302, as described herein.In at least one embodiment, instructions that, when executed, cause one or more operations or instructions to be executed by one or more accelerators within heterogeneous processors to be added to one or more executable graphs use software code, such as example software code that indicates adding one or more operations or instructions to one or more executable graphs through one or more APIs 306 of the parallel computing environment 308, as described herein at least in connection with. Fig. 5 described.

[0220] Fig. 14 is a block diagram 1400 illustrating a process for implementing an application programming interface (API) for partitioning context resources in accordance with at least one embodiment. In at least one embodiment, the process for implementing an API for partitioning context resources illustrated in block diagram 1400 is a process for implementing context resource partitioning API 1302, which is described herein at least in connection with Fig. 13. In at least one embodiment, part or all of the process depicted in block diagram 1400 for performing a context resource partitioning API (or other processes described herein or variations and / or combinations thereof) is performed under the control of one or more computer systems, servers, processors, integrated circuits, and / or other such devices, as described in connection with Fig. 26-58, configured with processor-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors such as those described herein. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, a processor such as the one described herein, at least in conjunction with Fig. 1 performs one or more steps of the process to implement a context resource partitioning API shown in block diagram 1400. In at least one embodiment, one or more other processors, such as those described herein, perform one or more steps of the process to implement a context resource partitioning API shown in block diagram 1400.

[0221] In at least one embodiment, in step 1402 of the process for performing a context resource partitioning API shown in block diagram 1400, a processor performing the process performs one or more operations to receive or otherwise obtain a context resource partitioning API. In at least one embodiment, in step 1402, a context resource partitioning API is received by otherwise providing it to a processor, such as processor 110, which is illustrated here at least in connection with Fig. 1. In at least one embodiment, in step 1402, an API for partitioning context resources is called by a processor, such as processor 102, described here at least in conjunction with Fig. 1 described in detail. In at least one embodiment, in step 1402, an API for partitioning context resources includes one or more arguments as described herein at least in connection with Fig. 13. In at least one embodiment, after step 1402, the process shown in block diagram 1400 for performing an API for partitioning context resources continues in step 1404.

[0222] In at least one embodiment, a processor performing this process performs one or more operations to determine whether a context resource subdivision API (e.g., received in step 1402) is valid in step 1404 of the process illustrated in block diagram 1400. In at least one embodiment, the operations to determine whether a context resource subdivision API is valid include, in step 1404, operations to check the validity of arguments of the context resource subdivision API (e.g., arguments described herein at least in connection with Fig. 13). In at least one embodiment, if a context resource partitioning API is determined to be valid ("YES" branch) in step 1404, the process for performing a context resource partitioning API, as shown in block diagram 1400, continues to step 1406. In at least one embodiment, if a context resource partitioning API is determined to be invalid ("NO" branch) in step 1404, the process for performing a context resource partitioning API, as shown in block diagram 1400, continues to step 1418, which is described below.

[0223] In at least one embodiment, in step 1406 of the process for performing a context resource partitioning API illustrated in block diagram 1400, a processor performing the process performs one or more operations to identify resources in an input resource list (e.g., received as an argument to a context resource partitioning API received in step 1402), as described herein. In at least one embodiment, after step 1406, the process for performing a context resource partitioning API illustrated in block diagram 1400 continues in step 1408.

[0224] In at least one embodiment, in step 1408 of the process depicted in block diagram 1400 for performing a context resource partitioning API, a processor performing this process performs one or more input resource partitioning operations (e.g., identified in step 1406) using flags and / or minimum counts (e.g., as arguments to a context resource partitioning API received in step 1402). In at least one embodiment, in step 1408, one or more input resource partitioning operations are performed using an indication of the available resources. In at least one embodiment, in step 1408, one or more input resource partitioning operations are performed to partition the resources into two or more resource lists, as described herein.In at least one embodiment, if no input resource list was identified in step 1406, no partitioning of the input resources is performed in step 1408. In at least one embodiment, after step 1408, the process continues to perform an API for partitioning context resources, as shown in block diagram 1400, in step 1410.

[0225] In at least one embodiment, in step 1410 of the process for performing a context resource partitioning API shown in block diagram 1400, a processor performing this process performs one or more operations to determine whether an input resource list has been partitioned (e.g., in step 1408). In at least one embodiment, if it is determined in step 1410 that an input resource list has been partitioned ("YES" branch), the process for performing a context resource partitioning API shown in block diagram 1400 continues to step 1412. In at least one embodiment, if it is determined in step 1410 that an input resource list has not been partitioned ("NO" branch), the process for performing a context resource partitioning API shown in block diagram 1400 continues to step 1418, which is described below.

[0226] In at least one embodiment, in step 1412 of the process for performing a context resource partitioning API shown in block diagram 1400, a processor performing the process performs one or more operations to store a partitioned input resource list (e.g., partitioned in step 1408) in a memory location received as an argument to a context resource partitioning API received in step 1402. In at least one embodiment, in step 1412, a partitioned input resource list is an empty list (e.g., contains no resources). In at least one embodiment, after step 1412, the process for performing a context resource partitioning API shown in block diagram 1400 continues in step 1414.

[0227] In at least one embodiment, in step 1414 of the process for performing a context resource partitioning API shown in block diagram 1400, a processor performing the process performs one or more operations to store a remaining input resource list (e.g., resources remaining after an input resource list is partitioned in step 1408) in a memory location received as an argument to a context resource partitioning API received in step 1402. In at least one embodiment, in step 1412, a remaining input resource list is an empty list (e.g., it contains no resources). In at least one embodiment, after step 1412, the process for performing a context resource partitioning API shown in block diagram 1400 continues in step 14164.

[0228] In at least one embodiment, in step 1416 of the process for performing a context resource partitioning API indicated in block diagram 1400, a processor performing the process performs one or more operations to return a success indication (e.g., a success indication 1322, which is illustrated here at least in connection with Fig. 13). In at least one embodiment, in step 1416, a success indication is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 1416, the process for executing an API for partitioning context resources, as shown in block diagram 1400, is terminated. In at least one embodiment described in Fig. 14, after step 1416, the process of executing an API for partitioning context resources shown in block diagram 1400 continues in step 1402 described above.

[0229] In at least one embodiment, in step 1418 of the process for performing a context resource partitioning API indicated in block diagram 1400, a processor performing the process performs one or more operations to return an error indicator (e.g., an error indicator 1324, which is illustrated here at least in connection with Fig. 13). In at least one embodiment, in step 1418, an error indicator is returned to a calling process (e.g., a process executed by a processor such as processor 102, described here at least in connection with Fig. 1). In at least one embodiment, after step 1418, the process for executing an API for partitioning context resources, as shown in block diagram 1400, is terminated. In at least one embodiment described in Fig. 14, after step 1418, the process of executing an API for partitioning context resources shown in block diagram 1400 continues in step 1402 described above.

[0230] In at least one embodiment, the operations of the process for performing an API for partitioning context resources shown in block diagram 1400 are performed in a different order than that shown in FIG. 1400. In at least one embodiment, the operations of the process for performing an API for partitioning context resources shown in block diagram 1400 are performed concurrently or in parallel. In at least one embodiment, operations of the process for performing an API for partitioning context resources shown in block diagram 1400 that are not dependent on one another (e.g., are order independent) are performed concurrently or in parallel.In at least one embodiment, the operations of the process for implementing a context resource partitioning API shown in block diagram 1400 are performed by a plurality of threads executing on a processor such as those described herein.

[0231] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators as described herein) will include one or more circuits to perform operations and / or instructions described herein in connection with Fig. 13 and Fig. 14, such as one or more circuits for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators such as those described herein) include one or more circuits for performing the operations described herein in connection with Fig. 13 and Fig. 14, such as one or more circuits for implementing an application programming interface (API) to generate an indication of one or more subsets of streaming multiprocessors (SMs) of a processor available for use in implementing one or more software kernels and / or otherwise implementing operations described herein. In at least one embodiment, the Fig. 13 and Fig. 14 described operations and / or instructions systems, methods, operations and / or instructions described herein in connection with the Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, components described herein in connection with Fig. 13 and Fig. 14, one or more processes described here in connection with Fig. 1-25 to execute an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment described in Fig. 13 and Fig. 14 is not shown, comprise one or more components which are described here in connection with Fig. 13 and Fig. 14, one or more components described here in connection with Fig. 26-58, to execute an application programming interface (API) to cause a corresponding plurality of identifiers of streaming multiprocessors (SMs) of one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0232] Fig. 15 is a block diagram 1500 illustrating an application programming interface (API) for generating a resource descriptor, according to at least one embodiment. In at least one embodiment, one or more circuits of a processor execute a resource descriptor generation API 1502 to generate a descriptor for one or more resources. In at least one embodiment, Fig. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to perform a resource descriptor generation API 1502 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads. In at least one embodiment, as described in Fig. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a resource descriptor generation API 1502 to execute an application programming interface (API) to generate a data structure containing information to specify one or more streaming multiprocessors (SMs) of one or more GPUs that can be used to execute one or more software kernels. In at least one embodiment, as described in Fig. 15, one or more circuits of a processor, such as those described herein, execute one or more instructions to execute a generate resource descriptor API 1502 to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads in response to receiving a second API, such as those described herein.

[0233] In at least one embodiment, the resource descriptor generation API 1502, when invoked, receives one or more arguments to specify information about operations to be performed using techniques such as those described herein. In at least one embodiment, the resource descriptor generation API 1502, when invoked, receives one or more arguments specifying information about commands to be executed using techniques such as those described herein.

[0234] In at least one embodiment, the Resource Descriptor Generation API 1502 receives as input one or more arguments comprising a Resource Descriptor Return 1504. In at least one embodiment, the Resource Descriptor Return 1504 is a data value comprising information suitable for identifying, indicating, or otherwise specifying a storage location for storing a resource descriptor generated by the Resource Descriptor Generation API 1502. In at least one embodiment, a storage location for a resource descriptor identified, indicated, or otherwise specified by the Resource Descriptor Return 1504 is one of a plurality of parameters that can be used by the Resource Descriptor Generation API 1502 to generate a resource descriptor.In at least one embodiment, the resource descriptor return 1504 is a data value that identifies, indicates, or otherwise specifies to an API such as the resource descriptor generator API 1502 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0235] In at least one embodiment, Resource Descriptor Generation API 1502 receives as input one or more arguments comprising a resource list 1506. In at least one embodiment, Resource List 1506 is a data value comprising information suitable for identifying, indicating, or otherwise specifying a resource list from which a resource descriptor is generated using Resource Descriptor Generation API 1502. In at least one embodiment, a resource list identified, indicated, or otherwise specified by Resource List 1506 is one of a plurality of parameters that can be used by Resource Descriptor Generation API 1502 to generate a resource descriptor.In at least one embodiment, resource list 1506 is a data value for identifying, indicating, or otherwise specifying a set of operations or instructions for an API such as Generate Resource Descriptor API 1502 to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0236] In at least one embodiment, the resource descriptor generation API 1502 receives as input one or more arguments comprising a number of resources 1508. In at least one embodiment, the number of resources 1508 is a data value comprising information suitable for identifying, indicating, or otherwise specifying a number of resources in the resource list 1506 that can be used to generate a resource descriptor using the resource descriptor generation API 1502. In at least one embodiment, the number of resources identified, indicated, or otherwise specified by the number of resources 1508 is one of a plurality of parameters that can be used by the resource descriptor generation API 1502 to generate a resource descriptor.In at least one embodiment, the number of resources 1508 is a data value that identifies, indicates, or otherwise specifies to an API such as the Generate Resource Descriptor API 1502 a set of operations or instructions to be performed by one or more PPUs, such as GPUs, and / or one or more accelerators within a heterogeneous processor as described herein.

[0237] In at least one embodiment, the resource descriptor generation API 1502 receives as input one or more arguments that include one or more other arguments 1518. In at least one embodiment, the other arguments 1518 are data that include information to indicate other information that can be used in executing the resource descriptor generation API 1502 to generate a resource descriptor.

[0238] In at least one embodiment described in Fig. 15, a processor executes one or more instructions to execute one or more APIs, such as generate resource descriptor API 1502, to execute an application programming interface (API) to specify one or more groups of streaming multiprocessors (SMs) of one or more processors on which to schedule one or more corresponding groups of software threads using one or more arguments, including, but not limited to, resource descriptor return 1504, resource list 1506, number of resources 1508, and / or other arguments 1518. In at least one embodiment, as shown in Fig. 15, a processor executes one or more instructions to execute one or more APIs, such as generate resource descriptor API 1502 to execute an application programming interface (API) to generate a data structure containing information to specify one or more streaming multiprocessors (SMs) of one or more GPUs that can be used to execute one or more software kernels using one or more arguments, including, but not limited to, resource descriptor return 1504, resource list 1506, number of resources 1508, and / or other arguments 1518.

[0239] In at least one embodiment, the resource descriptor generation API 1502, when invoked, causes one or more APIs such as one or more APIs 306 described herein at least in connection with Fig. 3, add, insert, or otherwise include one or more operations or instructions in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the Generate Resource Descriptor API 1502, when invoked, causes one or more APIs, such as one or more APIs 306, in a parallel computing environment, such as the parallel computing environment 308, described herein at least in connection with Fig. 3, adding, inserting, or otherwise including one or more operations or instructions into a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0240] In at least one embodiment, one or more APIs 306, when executed in response to Resource Descriptor Generate API 1502, are to cause one or more processors to perform a Resource Descriptor Generate Return API 1520. In at least one embodiment, Resource Descriptor Generate API Return 1520 is a set of instructions that, when executed, generate and / or specify one or more data values in response to Resource Descriptor Generate API 1502. In at least one embodiment, Resource Descriptor Generate API Return 1520 generates a success indication 1522. In at least one embodiment, Success Indicator 1522 is data comprising any value to indicate the success of Resource Descriptor Generate API 1502.In at least one embodiment, success indicator 1522 includes information indicating one or more specific types of successes generated as a result of performing resource descriptor generation API 1502. In at least one embodiment, success indicator 1522 includes information indicating one or more other data values generated as a result of resource descriptor generation API 1502.

[0241] In at least one embodiment, an error indicator 1524 is provided in the resource descriptor generation API return 1520. In at least one embodiment, the error indicator 1524 is data comprising any value indicative of an error of the resource descriptor generation API 1502. In at least one embodiment, the error indicator 1524 comprises information indicative of one or more specific types of errors generated as a result of performing the resource descriptor generation API 1502. In at least one embodiment, the error indicator 1524 comprises information indicative of one or more other data values generated as a result of the resource descriptor generation API 1502.

[0242] In at least one embodiment, parallel computing environment 308 includes one or more APIs 306, including, but not limited to, Generate Resource Descriptor API 1502, that add various operations of different types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include a semaphore acquire operation. In at least one embodiment, stream operations include a semaphore release operation. In at least one embodiment, stream operations include one or more operations to flush and / or invalidate cache memory, for example, L2 cache of a PPU, for example, a GPU, and / or cache memory of one or more accelerators within a heterogeneous processor.In at least one embodiment, stream operations include one or more indications for submitting an operation to an external device, such as one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations indicating submission of an operation to an external device use software code, such as example software code indicating stream operations, as described herein at least in connection with. Fig. 5 described.

[0243] In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306, including, but not limited to, Generate Resource Descriptor API 1502, one or more function signatures that can be used to specify one or more callback functions for operations to be performed by one or more accelerators within heterogeneous processors. In at least one embodiment, one or more operations cause one or more callback functions to be executed. In at least one embodiment, one or more operations that cause one or more callback functions to be executed use software code, such as example software code that specifies a function signature for a callback function, as described herein at least in connection with Fig. 5 described.

[0244] In at least one embodiment, to specify one or more accelerators within heterogeneous processors to perform one or more operations specified by Generate Re...

Claims

[1] Processor comprising: one or more circuits for implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors available to execute one or more software threads. [2] The processor of claim 1, wherein the API is to receive arguments comprising a memory location that can be used to store the indication of the one or more SMs. [3] The processor of claim 1, wherein the API is to receive arguments comprising a context containing information about resources that can be used to execute the one or more software threads. [4] The processor of claim 1, wherein the API is to indicate that the one or more SMs are available to be used to execute the one or more software threads. [5] The processor of claim 1, wherein the API is to return one or more data structures indicating the one or more SMs allocated by a second API to allocate the one or more data structures. [6] The processor of claim 1, wherein the indication of the one or more SMs is included in a context of the one or more processors. [7] The processor of claim 1, wherein the indication of the one or more SMs is included in a subcontext of the one or more processors. [8] Computer-implemented method comprising: Implementing an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads. [9] The computer-implemented method of claim 8, wherein performing the API comprises receiving arguments comprising a memory location usable for storing the indication of the one or more SMs. [10] The computer-implemented method of claim 8, wherein executing the API comprises receiving arguments comprising a set of resources that can be used to execute the one or more software threads. [11] The computer-implemented method of claim 8, wherein performing the API comprises indicating that the one or more SMs are available to be used to execute the one or more software threads. [12] The computer-implemented method of claim 8, wherein performing the API comprises returning one or more data structures indicating the one or more SMs allocated by a second API to allocate the one or more data structures. [13] The computer-implemented method of claim 8, wherein the indication of the one or more SMs is included in a context of the one or more processors. [14] The computer-implemented method of claim 8, wherein the indication of the one or more SMs is included in a subcontext of the one or more processors. [15] Computer system comprising: one or more processors and memory storing executable instructions that, when executed by the one or more processors, execute an application programming interface (API) to specify one or more streaming multiprocessors (SMs) of one or more processors that can be used to execute one or more software threads. [16] The computer system of claim 15, wherein the API is to receive arguments comprising a memory location that can be used to store the indication of the one or more SMs. [17] The computer system of claim 15, wherein the API is to receive arguments comprising a context containing information about resources that can be used to execute the one or more software threads. [18] The computer system of claim 15, wherein the API is to indicate that the one or more SMs are available to be used for executing the one or more software threads. [19] The computer system of claim 15, wherein the API is to return one or more data structures indicating the one or more SMs allocated by a second API for allocating the one or more data structures. [20] The computer system of claim 15, wherein the indication of the one or more SMs is included in a subcontext of the one or more processors.

Citation Information

Patent Citations

  • US-PATENTANMELDUNGNR.18/593,593

  • US-PATENTANMELDUNGNR.18/593,722

  • US-PATENTANMELDUNGNR.18/593,578

  • US-PATENTANMELDUNGNR.18/593,731

  • US-PATENTANMELDUNGNR.18/593,588