Application programming interface for deallocating data structures

By executing application programming interfaces (APIs) on the graphics processing unit (GPU) and managing the resources of software programs, the problem of inefficient memory and time usage in computer programs is solved, and the optimization of resource usage and computing performance is achieved.

CN120371504APending Publication Date: 2025-07-25NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510124484.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-01-26
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has challenges in the inefficient use of memory, time and resource when executing computer programs, which are difficult to effectively improve.

Method used

By executing an application programming interface (API) on a graphics processing unit (GPU), the resources of the software program include creating, destroying, obtaining and subdividing context resources, generating resource descriptors, recording and waiting for context events, to optimize resource usage.

Benefits of technology

It improves the memory and time usage efficiency of computer programs, optimizes resource management, and improves computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371504A_ABST
    Figure CN120371504A_ABST
Patent Text Reader

Abstract

The invention discloses an application programming interface for deallocating a data structure. Apparatuses, systems, and techniques for performing computing operations are disclosed. In at least one embodiment, a processor executes an application programming interface to deallocate one or more data structures for indicating which of one or more streaming multiprocessors of one or more processors are to be used to execute one or more software threads.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 625,278, filed on January 25, 2024, entitled “APPLICATION PROGRAMMING INTERFACE TO MANAGE RESOURCES,” the entire contents of which are incorporated herein by reference.This application also incorporates the entire disclosures of the following applications for all purposes: co-pending U.S. patent application Ser. No. 18 / 593,578, filed contemporaneously herewith, entitled “APPLICATION PROGRAMMING INTERFACE TO ALLOCATE A DATA STRUCTURE,” co-pending U.S. patent application Ser. No. 18 / 593,593, filed contemporaneously herewith, entitled “APPLICATION PROGRAMMING INTERFACE TO STORE AN IDENTIFIEROF A DATA STRUCTURE,” co-pending U.S. patent application Ser. No. 18 / 593,710, filed contemporaneously herewith, entitled “APPLICATION PROGRAMMING INTERFACE TO INDICATE MULTIPROCESSOR AVAILABILITY,” co-pending U.S. patent application Ser. No. 18 / 593,711, filed contemporaneously herewith, entitled “APPLICATION PROGRAMMING INTERFACE TO STORE IDENTIFIERS and Serial No. 18 / 593,726, filed concurrently with the present application, entitled “APPLICATION PROGRAMMING INTERFACE TO COMMUNICATE CONTEXT,” and Serial No. 18 / 593,730, filed concurrently with the present application, entitled “APPLICATION PROGRAMMING INTERFACE TO READ FROM A DATA STRUCTURE,” and Serial No. 18 / 593,746, filed concurrently with the present application, entitled “APPLICATION PROGRAMMING INTERFACE TO COMMUNICATE CONTEXT,” and Serial No. 18 / 593,750, filed concurrently with the present application, entitled “APPLICATION PROGRAMMING INTERFACE TO WAIT FOR CONTEXT.” CONTEXT)" in co-pending U.S. patent application Ser. No. 18 / 593,731. Technical Field

[0003] At least one embodiment relates to processing resources for executing one or more software programs on a graphics processing unit ("GPU").For example, at least one embodiment relates to implementing an application programming interface to manage resources for software programs. Background Art

[0004] Executing a computer program may use significant amounts of memory, time, or resources. The amount of memory, time, and / or resources used to execute a computer program can be improved. Despite some progress in accelerating or otherwise assisting with the execution of various components of a computer program, challenges remain in improving the use of memory, time, and / or resources when executing a computer program. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a block diagram illustrating a computer system for executing a computing resource operation application programming interface (API) according to at least one embodiment;

[0006] Figure 2 is a block diagram illustrating a context and a sub-context according to at least one embodiment;

[0007] Figure 3 is a block diagram illustrating a software program to be executed by one or more processors according to at least one embodiment;

[0008] Figure 4 is a block diagram illustrating a process for executing one or more application programming interfaces (APIs) according to at least one embodiment;

[0009] Figure 5 is a block diagram illustrating an application programming interface (API) for creating a subcontext according to at least one embodiment;

[0010] Figure 6 is a block diagram illustrating a process for executing an application programming interface (API) to create a subcontext according to at least one embodiment;

[0011] Figure 7 is a block diagram illustrating an application programming interface (API) for destroying a subcontext according to at least one embodiment;

[0012] Figure 8 is a block diagram illustrating a process for executing an application programming interface (API) to destroy a subcontext according to at least one embodiment;

[0013] Figure 9is a block diagram illustrating an application programming interface (API) for getting a subcontext from a stream according to at least one embodiment;

[0014] Figure 10 is a block diagram illustrating a process for executing an application programming interface (API) to obtain a subcontext from a stream according to at least one embodiment;

[0015] Figure 11 is a block diagram illustrating an application programming interface (API) for retrieving resources associated with a context according to at least one embodiment;

[0016] Figure 12 is a block diagram illustrating a process for executing an application programming interface (API) to obtain resources associated with a context in accordance with at least one embodiment;

[0017] Figure 13 is a block diagram illustrating an application programming interface (API) for subdivide context resources according to at least one embodiment;

[0018] Figure 14 is a block diagram illustrating a process for executing an application programming interface (API) to segment context resources according to at least one embodiment;

[0019] Figure 15 is a block diagram illustrating an application programming interface (API) for generating a resource descriptor according to at least one embodiment;

[0020] Figure 16 is a block diagram illustrating a process for executing an application programming interface (API) to generate a resource descriptor according to at least one embodiment;

[0021] Figure 17 is a block diagram illustrating an application programming interface (API) for obtaining device resources for context according to at least one embodiment;

[0022] Figure 18 is a block diagram illustrating a process for executing an application programming interface (API) to obtain contextual device resources according to at least one embodiment;

[0023] Figure 19 is a block diagram illustrating an application programming interface (API) for logging contextual events according to at least one embodiment;

[0024] Figure 20 is a block diagram illustrating a process for executing an application programming interface (API) to record contextual events according to at least one embodiment;

[0025] Figure 21 is a block diagram illustrating an application programming interface (API) for waiting for contextual events according to at least one embodiment;

[0026] Figure 22 is a block diagram illustrating a process for executing an application programming interface (API) to wait for a context event according to at least one embodiment;

[0027] Figure 23 is a block diagram illustrating an example software stack in which an application programming interface (API) is processed according to at least one embodiment;

[0028] Figure 24 is a block diagram illustrating a processor and modules according to at least one embodiment;

[0029] Figure 25 is a block diagram illustrating a driver and / or runtime including one or more libraries for providing one or more application programming interfaces (APIs) according to at least one embodiment;

[0030] Figure 26 An exemplary data center is shown in accordance with at least one embodiment;

[0031] Figure 27 A processing system according to at least one embodiment is shown;

[0032] Figure 28 A computer system according to at least one embodiment is shown;

[0033] Figure 29 A system according to at least one embodiment is shown;

[0034] Figure 30 An exemplary integrated circuit according to at least one embodiment is shown;

[0035] Figure 31 A computing system according to at least one embodiment is shown;

[0036] Figure 32 An APU is shown according to at least one embodiment;

[0037] Figure 33 A CPU according to at least one embodiment is shown;

[0038] Figure 34 An exemplary accelerator integrated slice is shown in accordance with at least one embodiment;

[0039] Figure 35A and Figure 35B An exemplary graphics processor is shown in accordance with at least one embodiment;

[0040] Figure 36A A graphics core according to at least one embodiment is shown;

[0041] Figure 36B GPGPU according to at least one embodiment is shown;

[0042] Figure 37A A parallel processor according to at least one embodiment is shown;

[0043] Figure 37B illustrates a processing cluster according to at least one embodiment;

[0044] Figure 37C A graphics multiprocessor is shown in accordance with at least one embodiment;

[0045] Figure 38 A graphics processor according to at least one embodiment is shown;

[0046] Figure 39 A processor according to at least one embodiment is shown;

[0047] Figure 40 A processor according to at least one embodiment is shown;

[0048] Figure 41 illustrates a graphics processor core according to at least one embodiment;

[0049] Figure 42 illustrates a PPU according to at least one embodiment;

[0050] Figure 43 shows a GPC according to at least one embodiment;

[0051] Figure 44 A streaming multiprocessor is shown in accordance with at least one embodiment;

[0052] Figure 45 illustrates a software stack for a programming platform according to at least one embodiment;

[0053] Figure 46 According to at least one embodiment, Figure 45 CUDA implementation of the software stack;

[0054] Figure 47 According to at least one embodiment, Figure 45 ROCm implementation of the software stack;

[0055] Figure 48 According to at least one embodiment, Figure 45 OpenCL implementation of the software stack;

[0056] Figure 49illustrates software supported by a programming platform according to at least one embodiment;

[0057] Figure 50 According to at least one embodiment, Figures 45-48 Compiled code executed on the programming platform;

[0058] Figure 51 According to at least one embodiment, Figures 45-48 More detailed compiled code executed on the programming platform;

[0059] Figure 52 Transforming source code before compiling it according to at least one embodiment is shown;

[0060] Figure 53A A system configured to compile and execute CUDA source code using different types of processing units is shown in accordance with at least one embodiment;

[0061] Figure 53B A system configured to compile and execute CUDA source code for graphics 53A using a CPU and a CUDA-enabled GPU is shown in accordance with at least one embodiment;

[0062] Figure 53C A system configured to compile and execute CUDA source code for graphics 53A using a CPU and a non-CUDA enabled GPU is shown in accordance with at least one embodiment;

[0063] Figure 54 According to at least one embodiment, Figure 53C An example kernel converted by the CUDA to HIP conversion tool;

[0064] Figure 55 More details are shown according to at least one embodiment. Figure 53C a non-CUDA-enabled GPU; and

[0065] Figure 56 shows how threads of an exemplary CUDA grid are mapped to Figure 55 different computational units; and

[0066] Figure 57 shows how to migrate existing CUDA code to data parallel C++ code according to at least one embodiment; and

[0067] Figure 58 Components of a system for accessing large language models in accordance with at least one embodiment are shown. DETAILED DESCRIPTION

[0068] In at least one embodiment, software (e.g., a computer program) executed by one or more processors causes software executed by other processors (e.g., in a data center) to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the software executed by one or more processors executes one or more application programming interfaces (APIs) to create a subcontext that manages a subset of resources available for executing the software program. In at least one embodiment, the software executed by one or more processors executes one or more APIs to create a subcontext that manages a subset of resources available for executing the software program, thereby causing the software executed by the other processors to perform operations to manage resources associated with the software executed by the other processors.

[0069] In at least one embodiment, software executed by one or more processors executes one or more APIs to destroy a subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to destroy a subcontext, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the subcontext destroyed by the application programming interface (API) for destroying a subcontext is a subcontext created by the API for creating a subcontext described above.

[0070] In at least one embodiment, software executed by one or more processors executes one or more APIs to create a subcontext from a stream (e.g., a series of operations to be performed in a specified order by one or more other processors). In at least one embodiment, software executed by one or more processors executes one or more APIs to create a subcontext from a stream, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the subcontext created by the API for creating a subcontext for a stream serves as a subcontext created by the API for creating a subcontext as described above, and is a subcontext that can be destroyed using the API for destroying a subcontext as also described above.

[0071] In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain resources associated with a context or subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain resources associated with a context or subcontext, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the resources obtained when executing the API for obtaining resources associated with a context or subcontext are resources associated with the context or subcontext created as described above (e.g., using an API for creating a subcontext or using an API for creating a subcontext from a stream).

[0072] In at least one embodiment, software executed by one or more processors executes one or more APIs to subdivide context resources according to one or more criteria. In at least one embodiment, software executed by one or more processors executes one or more APIs to subdivide context resources, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the context resources subdivided by the API for subdividing context resources are resources associated with the context or subcontext executed as described above, and are resources associated with the context or subcontext created as described above (e.g., using the API for creating subcontexts or using the API for creating subcontexts from streams). In at least one embodiment, resource subdivisions (e.g., resource subdivisions obtained by the API for subdividing context resources) are used to generate subcontexts as described herein.

[0073] In at least one embodiment, software executed by one or more processors executes one or more APIs to generate resource descriptors from a resource list. In at least one embodiment, software executed by one or more processors executes one or more APIs to generate resource descriptors, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the resource descriptor list generated by the API for generating resource descriptors is used to create and / or manage subcontexts, as described above.

[0074] In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain device resources associated with a context or subcontext. In at least one embodiment, software executed by one or more processors executes one or more APIs to obtain device resources, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the device resources obtained by executing the obtain device resource API can be used to create subcontexts and / or further subdivide resources, as described above.

[0075] In at least one embodiment, software executed by one or more processors executes one or more APIs to record context events (e.g., to emit an event associated with another context so that synchronization between the contexts can be performed). In at least one embodiment, software executed by one or more processors executes one or more APIs to record context events so that software executed by other processors performs operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the context events recorded by the API for recording context events are used by the context associated with the API for waiting for context events, as described below.

[0076] In at least one embodiment, software executed by one or more processors executes one or more APIs to wait for context events (e.g., to wait for an event associated with another context so that synchronization between the contexts can be performed). In at least one embodiment, software executed by one or more processors executes one or more APIs to wait for context events, thereby causing software executed by other processors to perform operations to manage resources associated with the software executed by the other processors. In at least one embodiment, the context events of the API for waiting for context events are used by the context associated with the API for recording context events, as described above.

[0077] Figure 1 1 is a block diagram 100 illustrating a computer system that executes a computing resource operation application programming interface (API) according to at least one embodiment. In at least one embodiment, a processor 102 executes one or more software programs 104. In at least one embodiment, the processor 102 is a processor such as those described below. In at least one embodiment, the processor 102 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general purpose graphics processing unit (GPGPU), a computing cluster, and / or a combination of these and / or other such processors. In at least one embodiment, the processor 102 is part of a computer system, such as the computer system described herein. In at least one embodiment ( Figure 1 1 , processor 102 is a processor of a client computing system. In at least one embodiment, the client computing system includes one or more client devices, such as the client devices described herein. In at least one embodiment, the client computing system includes one or more devices that are clients of a cloud computing environment, such as those described herein.

[0078] In at least one embodiment, software program 104 includes one or more computing operations, such as computing operations for training a neural network, executing a neural network, executing a Compute Unified Device Architecture (CUDA) program, executing a large language model, performing rendering operations, performing data analysis, and / or performing other operations, including but not limited to the operations described herein. In at least one embodiment, software program 104 includes software such as described herein.

[0079] In at least one embodiment, the software program 104 performs resource operations 106 using a resource operation API 108. In at least one embodiment, the resource operations 106 are operations for managing resources associated with the processor 110. In at least one embodiment, the processor 110 is a processor such as those described below. In at least one embodiment, the processor 110 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general-purpose graphics processing unit (GPGPU), a computing cluster, and / or a combination of these and / or other such processors. In at least one embodiment, the processor 110 is part of a computer system, such as the computer system described herein.

[0080] In at least one embodiment ( Figure 1(not shown), processor 110 is one of multiple processors of a high-performance computing system. In at least one embodiment, a high-performance computing system is a computing system that includes multiple processors to perform computing operations, such as those described herein. In at least one embodiment, the high-performance computing system is a distributed computing system. In at least one embodiment, the high-performance computing system is a deep learning computing system. In at least one embodiment, the operations of performing the computing operations are performed by the high-performance computing system using the systems, methods, operations, and / or techniques described herein. In at least one embodiment, the high-performance computing system includes a cloud computing environment, such as that described herein. In at least one embodiment, the high-performance computing system includes one or more processors, such as processor 110. In at least one embodiment, the high-performance computing system includes one or more graphics processors, such as those described herein. In at least one embodiment, the processors of the high-performance computing system include one or more central processing units (CPUs), graphics processing units (GPUs), parallel processing units (PPUs), general-purpose graphics processing units (GPGPUs), computing clusters, and / or combinations of these and / or other such processors, as described herein.

[0081] In at least one embodiment, a high performance computing system includes a plurality of processors (e.g., processor 110) having an associated set of resources that the processors (e.g., processor 102) can use to perform computing operations, including one or more resource operations 106 using one or more resource operation APIs 108. In at least one embodiment, resource operations 106 are computing operations that use systems, methods, and / or operations such as those described herein to manage resources of a process being executed by one or more processors.

[0082] In at least one embodiment, the resource operation API 108 includes one or more APIs for managing resources of a processor (e.g., processor 110) that is operable to execute the software program 104. In at least one embodiment, the resource operation API 108 includes one or more APIs, such as the create subcontext API 502 (described herein at least in conjunction with Figure 5 and Figure 6 Description), destroy subcontext API 702 (this article at least combines Figure 7 and Figure 8 Description), get subcontext API 902 from the stream (this article at least combines Figure 9 and Figure 10 Description), get context resource API 1102 (this article at least combines Figure 11 and Figure 12 Description), Segment Context Resource API 1302 (this article at least combines Figure 13 and Figure 14 Description), generate resource descriptor API 1502 (this article at least combines Figure 15 and Figure 16 Description), get device resource API 1702 (this article at least combines Figure 17 and Figure 18 Description), Record Context Event API 1902 (this article at least combines Figure 19 and Figure 20 Description), wait context event API 2102 (this article at least in conjunction with Figure 21 and Figure 22 description), and / or other such resource operation APIs.

[0083] In at least one embodiment, the processor 110 receives the resource operation API 108 and performs one or more operations to create one or more contexts 112 that include processor resources 114. In at least one embodiment, the context 112 includes a description of resources available to a processor (e.g., processor 110), as described below in conjunction with at least Figure 2 In at least one embodiment, processor resources 114 include resources available to processor 110 for executing one or more workloads (described below). In at least one embodiment, context 112 is a primary context for processor 110, which is a previously created context and / or a default context for processor 110.

[0084] In at least one embodiment, the processor 110 receives the resource operation API 108 and performs one or more operations to create one or more subcontexts (e.g., subcontexts 118A-118N) that can be used to execute one or more workloads (e.g., workloads 120A-120N). In at least one embodiment, the processor 110 receives the resource operation API 108 and performs one or more operations to create one or more subcontexts based on one or more subsets (e.g., subsets 116A-116N) of the processor resources 114. In at least one embodiment, the subcontexts (e.g., subcontexts 118A-118N) are used to execute the workloads (e.g., workloads 120A-120N). In at least one embodiment, the subcontexts are derived from the context 112 using a subset of resources available to the processor 110 (e.g., as described in at least some embodiments herein). Figure 2In at least one embodiment, a subcontext is referred to as a green context. In at least one embodiment, a subset of processor resources 114 (e.g., subsets 116A-116N) includes some or all of processor resources 114 (e.g., a first subset of processor resources 114 may include a first portion of the processor resources, a second subset of processor resources 114 may include a second portion of the processor resources, and so on). In at least one embodiment, a subset of processor resources is not empty (e.g., includes at least a portion of processor resources 114). In at least one embodiment, a workload (e.g., workloads 120A-120N) includes at least a portion of the computing operations of software program 104 to be executed by one or more processors (e.g., processor 110) using systems, methods, and operations such as those described herein. In at least one embodiment, a workload (e.g., workloads 120A-120N) is also referred to as a software workload. In at least one embodiment, a workload (e.g., workloads 120A-120N) is also referred to as a kernel. In at least one embodiment, workloads (eg, workloads 120A-120N) are also referred to as software kernels.

[0085] In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to create a subcontext using the subcontext API 502, which is at least in combination with Figure 5 and Figure 6 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to destroy the subcontext using the destroy subcontext API 702, which is at least in combination with Figure 7 and Figure 8 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to obtain the subcontext from the stream using the Get Subcontext from Stream API 902, which is at least in combination with Figure 9 and Figure 10 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to obtain the resources associated with the context using the obtain context resource API 1102, and this article at least combines Figure 11 and Figure 12 In at least one embodiment ( Figure 1), the processor 110 receives the resource operation API 108 and performs one or more operations to use the subdivision context resource API 1302 to subdivide the context resource, and this article is at least combined with Figure 13 and Figure 14 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to generate a resource descriptor using the generate resource descriptor API 1502, which is at least in combination with Figure 15 and Figure 16 In at least one embodiment ( Figure 1 The processor 110 receives the resource operation API 108 and performs one or more operations to obtain the device's resources using the obtain device resource API 1702 (at least in conjunction with the present invention). Figure 17 and Figure 18 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to record context events using the record context event API 1902, which is at least in combination with Figure 19 and Figure 20 In at least one embodiment ( Figure 1 ), the processor 110 receives the resource operation API 108 and performs one or more operations to wait for context events using the wait context event API 2102, which is at least in combination with Figure 21 and Figure 22 Described.

[0086] In at least one embodiment, multiple subsets of processor resources 114 (e.g., multiple subsets of subsets 116A-116N) are used by and / or associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or associated with multiple subcontexts (e.g., multiple subcontexts of subcontexts 118A-118N). In at least one embodiment, the number of subsets of processor resources 114 (e.g., subsets 116A-116N) is different from the number of subcontexts (e.g., subcontexts 118A-118N). In at least one embodiment, the number of subsets of processor resources 114 (e.g., subsets 116A-116N) is the same as the number of subcontexts (e.g., subcontexts 118A-118N).

[0087] In at least one embodiment, multiple subcontexts (e.g., multiple subcontexts from subcontexts 118A-118N) are used by and / or are associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or is associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, a subset of processor resources 114 (e.g., one of subsets 116A-116N) is used by and / or is associated with one subcontext (e.g., one of subcontexts 118A-118N). In at least one embodiment, the number of subsets of processor resources 114 (e.g., subsets 116A-116N) is different from the number of subcontexts (e.g., subcontexts 118A-118N). In at least one embodiment, the number of subsets of processor resources 114 (e.g., subsets 116A-116N) is the same as the number of subcontexts (e.g., subcontexts 118A-118N).

[0088] In at least one embodiment, as used in any implementation described herein, unless the context clearly indicates otherwise or clearly contradicts, terms such as "module" and nominalized verbs (e.g., subcontext, workload, context, and / or other terms) refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any implementation described herein may include, for example, individually or in any combination, hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware that stores instructions executed by programmable circuitry. In at least one embodiment, modules may be embodied collectively or individually as circuitry that constitutes part of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), etc.).

[0089] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to generate one or more masks for indicating one or more subsets of streaming multiprocessors in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0090] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations described herein, such as operations executing an application programming interface (API) to generate one or more masks indicating one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein.

[0091] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for executing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0092] In at least one embodiment ( Figure 1), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein, for example, operations of executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations described herein include, for example, executing an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate that one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) may be used to execute one or more kernels and / or otherwise perform the operations described herein.

[0093] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform operations described herein.

[0094] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein, for example, executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise performing the operations described herein. In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations are described, for example, by executing an application programming interface (API) to indicate one or more identifiers of one or more masks, where the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that may be used to execute one or more kernels and / or otherwise perform the operations described herein.

[0095] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for performing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more graphics processing units (GPUs) to be available for executing one or more software kernels and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described are configured to implement an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or otherwise perform the operations described herein.

[0096] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein include, for example, executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more graphics processing units (GPUs) to be available for executing one or more software kernels and / or otherwise performing the operations described herein.

[0097] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for performing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to generate instructions for one or more streaming multiprocessor (SM) subsets in a processor that can be used to execute one or more software kernels and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, this document is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, the present invention relates to a method for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein implement an application programming interface (API) to cause identifiers of corresponding streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components are described for implementing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0098] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein include, for example, executing an application programming interface (API) to cause identifiers of corresponding streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations described herein include, for example, executing an application programming interface (API) to generate indications of one or more streaming multiprocessor (SM) subsets of a processor available for executing one or more software kernels and / or otherwise performing operations described herein.

[0099] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for performing the operations and / or instructions described herein in conjunction with Figure 1One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to generate a data structure including information for indicating one or more streaming multiprocessors (SMs) in one or more GPUs that are to be used to execute one or more software kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding sets of software threads thereon and / or otherwise perform the operations described herein.

[0100] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein, for example, executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25The operations described herein include, for example, executing an application programming interface (API) to generate a data structure that includes information for indicating one or more streaming multiprocessors (SMs) in one or more GPUs to be used to execute one or more software kernels and / or otherwise perform the operations described herein.

[0101] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for executing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) so that the number of streaming multiprocessors indicated by one or more masks that will be available for executing one or more software kernels is indicated and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, this document is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.

[0102] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein, for example, executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise performing the operations described herein. In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations described, for example, executing an application programming interface (API) to cause the number of streaming multiprocessors indicated by one or more masks to be available for executing one or more software kernels to be indicated and / or otherwise performing operations described herein.

[0103] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for implementing the operations and / or instructions described herein, for example, for executing an application programming interface (API) so that the context of one or more first software instructions is communicated to one or more second software instructions and / or otherwise performing one or more circuits for implementing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for implementing the operations and / or instructions described herein in conjunction with Figure 1One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of a processor to be directed to one or more resources of a second context of a processor and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described are configured to implement an application programming interface (API) to enable the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein.

[0104] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 The operations described herein include, for example, executing an application programming interface (API) to cause one or more indications of one or more events of one or more resources of a first context of a processor to be directed to one or more resources of a second context of a processor and / or otherwise performing the operations described herein.

[0105] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to cause one or more second instructions to wait for execution until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for performing the operations and / or instructions described herein in conjunction with Figure 1 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of a processor until one or more indications of one or more events of a second context are generated and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 1 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment ( Figure 1 Not shown), this paper combines Figure 1 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0106] In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 In at least one embodiment ( Figure 1 ), a non-transitory machine-readable medium stores a set of instructions, which, when executed by one or more processors, are used to perform at least the combination of the present invention. Figures 1 to 25 Operations described herein, for example, executing an application programming interface (API) to prevent one or more instructions from being executed by one or more resources of a first context of a processor until one or more indications of one or more events of a second context are generated and / or otherwise performing operations described herein.

[0107] Figure 2 is a block diagram 200 illustrating contexts and subcontexts according to at least one embodiment. In at least one embodiment, a device 202 is used to execute one or more workloads (e.g., Figure 1 In at least one embodiment, the device 202 includes one or more processors, such as those described herein in conjunction with at least one embodiment of the present invention. Figure 1 In at least one embodiment, the device 202 includes one or more channels for executing a flow of workloads (e.g., a first workload followed by a second workload, etc.). Figure 2 In at least one embodiment, the device 202 has a hardware context (e.g., a set of hardware resources). In at least one embodiment, the device 202 has multiple hardware contexts. In at least one embodiment, the performance of the device 202 is improved when the device 202 has only one hardware context. In at least one embodiment, the hardware context of the device 202 has one or more primary contexts associated with the hardware context. In at least one embodiment, the primary context includes descriptions and / or references to resources of the device 202 (e.g., the hardware context). In at least one embodiment, the primary context is mapped one-to-one with the hardware context of the device 202 (e.g., includes only references to the hardware resources of the device 202).

[0108] In at least one embodiment, one or more channels of the device 202 are primary context channels 212 (e.g., C1 204, C2 206, C3 208, and C4 210) that are used by the primary context to execute software workloads using the device 202. In at least one embodiment, the primary context channels 212 are used exclusively by the primary context to execute software workloads. In at least one embodiment, one or more channels of the device 202 are preferred select channels 232 (e.g., C5 214, C6 220, C7 226, and C8 230). In at least one embodiment, the preferred selection channel 232 is used by the subcontext to execute the software workload. In at least one embodiment, the subcontext is dynamically assigned to one or more preferred selection channels 232. In at least one embodiment, the subcontext (also known as the green context) is a subset of the context (e.g., a subset of the resources of the device 202) that does not require any context switch (e.g., a switch of the hardware context). In at least one embodiment, the subcontext is a one-to-one mapping of the resource subset of the hardware context. In at least one embodiment, the software program (e.g., at least in conjunction with the present invention) Figure 1 In at least one embodiment, a software program 104 (e.g., described herein) performs computing operations to switch between subcontexts while executing a workload (e.g., workloads 120A-120N) using device 202. In at least one embodiment, one or more subcontexts are managed by a single thread executing on device 202. In at least one embodiment, only one subcontext managed by a thread is current at a time. In at least one embodiment, the thread performs operations to switch between the subcontexts managed by the thread to keep one subcontext current at a time.

[0109] In at least one embodiment, a subcontext is assigned to multiple channels in the preferred selection channel 232 (e.g., subcontext 216 is assigned to both C5 214 and C6 220, and subcontext 218 is also assigned to both C5 214 and C6 220). In at least one embodiment, a subcontext is assigned to a single channel (e.g., subcontext 222 is assigned to C7 226, and subcontext 224 is also assigned to C7 226). In at least one embodiment, a subcontext is assigned exclusively to a channel (e.g., subcontext 228 is assigned to C8 230).

[0110] In at least one embodiment, the main context is referred to as the full context. In at least one embodiment, the full context has access to all execution resources of device 202. In at least one embodiment, the main context is the default context used by the CUDA runtime to manage resources of device 202. In at least one embodiment, the main context is thread-safe. In at least one embodiment, child contexts are mapped to the main context and can use a subset of the resources of device 202. In at least one embodiment, the child contexts can also use other resources (e.g., resources from other devices and / or contexts).

[0111] In at least one embodiment, a subcontext includes a resource group that is used to manage a set of resources used by the subcontext. In at least one embodiment, a subcontext can implicitly set a set of resources that a software library or software framework (e.g., as described herein) can use. In at least one embodiment, a resource group represents one or more resources that a particular context, subcontext, or API can use. In at least one embodiment, a resource group is a descriptor (e.g., a description of a resource) that can be partitioned by one or more APIs (e.g., the APIs described herein) to manage resources (e.g., in a hierarchical manner). In at least one embodiment, a resource group includes a description or list of multiple different resources of device 202, including but not limited to streaming multiprocessors (SMs), device connections (e.g., replication and compute hardware channels), processor bindings, software schedulers, work distribution, etc. In at least one embodiment, other APIs (e.g., CUDA APIs) can be assigned to a subcontext (e.g., the current subcontext) and use the resources of the subcontext. In at least one embodiment, other APIs (e.g., CUDA APIs) can be assigned directly to a resource group (e.g., without using a subcontext).

[0112] In at least one embodiment, one or more processors (e.g., as described herein) include a processor for executing the Figure 2 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein are used to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0113] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for executing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0114] In at least one embodiment, one or more processors (e.g., such as the processors described herein) include a processor for executing the instructions in conjunction with the instructions herein. Figure 2 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58One or more components described herein are configured to execute an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform operations described herein.

[0115] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein.

[0116] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, the present invention relates to a method, system, and / or instruction for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes are described for implementing an application programming interface (API) to cause identifiers of corresponding streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components are described for implementing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0117] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding sets of software threads thereon and / or otherwise perform the operations described herein.

[0118] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described are configured to execute an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.

[0119] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) so that the context of one or more first software instructions is transferred to one or more second software instructions and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described are configured to implement an application programming interface (API) to enable the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein.

[0120] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 2 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 2 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment ( Figure 2 Not shown), this paper combines Figure 2 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0121] Figure 3 3 is a block diagram 300 illustrating a software program to be executed by one or more processors according to at least one embodiment. In at least one embodiment, the block diagram 300 illustrates a software program 304 to be executed by a processor, such as a central processing unit (CPU) 302 and a graphics processing unit (GPU) 310 and an accelerator 314 within a heterogeneous processor. In at least one embodiment, the CPU 302 is a processor, such as a processor that is combined with at least Figure 1 The processor 102 described herein. In at least one embodiment, the CPU 302 is a processor, such as a processor that combines at least Figure 1Processor 110 as described herein. In at least one embodiment, CPU 302 is any processor having any architecture further described herein. In at least one embodiment, CPU 302 is any general-purpose processor having any architecture further described herein. In at least one embodiment, a processor (e.g., CPU 302) includes circuitry for performing one or more computing operations. In at least one embodiment, a processor (e.g., CPU 302) includes circuitry of any configuration for performing one or more computing operations further described herein.

[0122] In at least one embodiment, a processor (e.g., a central processing unit (CPU) 302) executes a parallel computing environment 308. In at least one embodiment, the processor (e.g., a CPU) 302 is. In at least one embodiment, the processor (e.g., a CPU) executes a parallel computing environment 308, such as the Compute Unified Device Architecture (CUDA). In at least one embodiment, the parallel computing environment 308 includes instructions that, if executed by one or more processors (e.g., CPUs 302), facilitate execution of one or more software programs by one or more CPUs 302, one or more parallel processing units (PPUs) (e.g., GPUs 310), and / or one or more accelerators 314 within a heterogeneous processor.

[0123] In at least one embodiment, one or more PPUs are processors that include one or more circuits for performing parallel computing operations, such as GPU 310 and any other parallel processors further described herein. In at least one embodiment, GPU 310 is hardware that includes circuits for performing one or more computing operations, as further described below in conjunction with various embodiments. In at least one embodiment, GPU 310 includes one or more processing cores, each of which is configured to perform one or more computing operations. In at least one embodiment, GPU 310 includes one or more processing cores for performing one or more parallel computing operations. In at least one embodiment, GPU 310 is packaged together with CPU 302 or other processors as a system on a chip (SoC). In at least one embodiment, GPU 310 is packaged together with CPU 302 or other processors on a shared die or other substrate as a system on a chip (SoC). In at least one embodiment, one or more accelerators 314 within a heterogeneous processor are hardware that includes one or more circuits for performing specific computing operations, such as a deep learning accelerator (DLA), a programmable vision accelerator (PVA), a field programmable gate array (FPGA), or any other accelerator further described herein. In at least one embodiment, the accelerator 314 within the heterogeneous processor is packaged with the CPU 302 or other processor as a system on a chip (SoC). In at least one embodiment, the accelerator 314 within the heterogeneous processor is packaged with the CPU 302 or other processor on a shared die or other substrate as a system on a chip (SoC). In at least one embodiment, one or more CPUs 302, one or more GPUs 310 or other PPUs, and / or accelerators 314 within the heterogeneous processor are packaged as a system on a chip (SoC). In at least one embodiment, one or more CPUs 302, one or more GPUs 310 or other PPUs, and / or accelerators 314 within the heterogeneous processor are packaged on a shared die or other substrate as a system on a chip (SoC).

[0124] In at least one embodiment, the parallel computing environment 308 (e.g., CUDA) includes libraries and other software programs for performing one or more computing operations using one or more PPUs (e.g., GPUs 310) and / or one or more accelerators 314 within a heterogeneous processor. In at least one embodiment, the parallel computing environment 308 includes libraries and other software programs that, when executed by one or more processors (e.g., one or more CPUs 302), cause one or more PPUs (e.g., GPUs 310) and / or one or more accelerators 314 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, the parallel computing environment 308 includes libraries that, if executed, cause one or more PPUs (e.g., GPUs 310) and / or one or more accelerators 314 within a heterogeneous processor to perform mathematical operations. In at least one embodiment, the parallel computing environment 308 includes libraries that, if executed, cause one or more PPUs (e.g., GPUs 310) and / or one or more accelerators 314 within a heterogeneous processor to perform any other operations further described herein.

[0125] In at least one embodiment, one or more PPUs (e.g., GPU 310) and / or one or more accelerators 314 within a heterogeneous processor perform one or more computing operations in response to one or more application programming interfaces (APIs). In at least one embodiment, an API is a set of software instructions that, if executed by one or more processors (e.g., CPU 302), causes one or more PPUs (e.g., GPU 310) and / or one or more accelerators 314 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, the parallel computing environment 308 includes one or more APIs 306 that, if executed by one or more processors (e.g., CPU 302), causes one or more PPUs (e.g., GPU 310) and / or one or more accelerators 314 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, the one or more APIs 306 include one or more functions that, if executed, cause one or more processors (e.g., CPU 302) to perform one or more operations, such as computational operations, error reporting, scheduling other operations to be performed by GPU 310 and / or accelerators 314 within heterogeneous processors, or any other operations further described herein. In at least one embodiment, the one or more APIs 306 include one or more functions that, if executed, cause one or more PPUs (e.g., GPU 310) to perform one or more operations, such as computational operations, error reporting, or any other operations further described herein. In at least one embodiment, the one or more APIs 306 include one or more functions, such as the following in conjunction with Figures 5 to 22 The one or more APIs 306 include functions that, if executed, cause one or more accelerators 314 within the heterogeneous processor to perform one or more operations, such as computational operations, error reporting, or any other operations further described herein. In at least one embodiment, the one or more APIs 306 include one or more functions that cause the CPU 302 to perform one or more computational operations in response to information or events generated by one or more PPUs (e.g., GPU 310) and / or one or more accelerators 314 within the heterogeneous processor. In at least one embodiment, the one or more APIs 306 include one or more functions that, when called, cause the CPU 302 to perform one or more computational operations in response to information or events generated by one or more PPUs (e.g., GPU 310) and / or one or more accelerators 314 within the heterogeneous processor.

[0126] In at least one embodiment, a processor (e.g., CPU 302) executes one or more software programs 304. In at least one embodiment, the one or more software programs are sets of instructions that, if executed, cause one or more processors (e.g., CPU 302, PPUs (e.g., GPU 310), and / or accelerators 314 in heterogeneous processors) to perform computing operations. In at least one embodiment, the software programs 304 include instructions and / or operations to be executed by one or more PPUs (e.g., GPU 310). In at least one embodiment, the one or more software programs 304 include GPU-specific code 312 and / or accelerator-specific code 316. In at least one embodiment, the instructions and / or operations to be executed by one or more PPUs (e.g., GPU 310) are PPU-specific or GPU-specific code 312. In at least one embodiment, the GPU-specific code 312 is a set of software instructions and / or other operations to be executed by one or more GPUs 310 (as further described herein). In at least one embodiment, the software program 304 includes instructions and / or operations to be executed by one or more accelerators 314 in a heterogeneous processor. In at least one embodiment, the instructions and / or operations to be executed by one or more accelerators 314 in a heterogeneous processor are accelerator-specific code 316. In at least one embodiment, the accelerator-specific code 316 is a set of software instructions and / or other operations to be executed by one or more accelerators 314, as further described herein. In at least one embodiment, the PPU-specific or GPU-specific code 312 and / or the accelerator-specific code 316 are executed in response to one or more APIs 306, as described below in conjunction with Figures 5 to 22 As stated.

[0127] In at least one embodiment, one or more processors (e.g., the processors described herein) include a processor for executing the instructions in conjunction with Figure 3 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein are used to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0128] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to deallocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0129] In at least one embodiment, one or more processors (e.g., such as the processors described herein) include a processor for executing the instructions in conjunction with the instructions herein. Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58One or more components described herein for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform operations described herein.

[0130] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, this document is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein.

[0131] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, the present invention relates to a method, system, or process for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes are described for implementing an application programming interface (API) to cause identifiers of corresponding streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components are described that implement an application programming interface (API) to cause identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0132] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding sets of software threads thereon and / or otherwise perform the operations described herein.

[0133] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more indicators of one or more quantities in one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein that implement an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.

[0134] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) so that the context of one or more first software instructions is transferred to one or more second software instructions and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components are described for implementing an application programming interface (API) to enable the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein.

[0135] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 3 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause one or more second instructions to wait for execution until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 3 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment ( Figure 3 Not shown), this paper combines Figure 3 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0136] Figure 4 4 is a block diagram 400 illustrating a process for executing one or more application programming interfaces (APIs) according to at least one embodiment. In at least one embodiment, the process for executing one or more APIs shown in block diagram 400 utilizes one or more accelerators within a heterogeneous processor via a parallel computing environment (e.g., parallel computing environment 308), as described herein in conjunction with at least one embodiment. Figure 3In at least one embodiment, the process of executing one or more APIs shown in block diagram 400 begins 402 at step 404, whereby one or more processors execute a software program comprising one or more instructions that, if executed, cause the one or more processors and / or one or more other processors (e.g., a graphics processing unit (GPU) and / or a heterogeneous processor or one or more accelerators within a heterogeneous processor) to perform one or more computing operations. In at least one embodiment, at step 404, the software program executed by the one or more processors includes one or more instructions that, if executed, cause one or more APIs 306 of a parallel computing environment 308 to be executed, as described above. In at least one embodiment, after step 404, the process of executing one or more APIs shown in block diagram 400 continues at step 406.

[0137] In at least one embodiment, at step 406 of the process for executing one or more APIs shown in block diagram 400, the processor executing the process for executing one or more APIs shown in block diagram 400 determines whether to execute a process such as described herein in conjunction with at least Figures 5 to 22 In at least one embodiment, if it is determined at step 406 that the API is not to be executed (the "No" branch), the process of executing one or more of the APIs shown in block diagram 400 continues at step 416. In at least one embodiment, if it is determined at step 406 that the API is to be executed (the "Yes" branch), the process of executing one or more of the APIs shown in block diagram 400 continues at step 408.

[0138] In at least one embodiment, at step 408 of the process for executing one or more APIs shown in block diagram 400, the processor executing the process for executing one or more APIs shown in block diagram 400 executes the APIs, for example, in conjunction with at least Figures 5 to 22 In at least one embodiment, at step 408, one or more processors execute one or more instructions to execute one or more API calls by the one or more processors and / or by one or more other processors (e.g., a GPU and / or an accelerator within a heterogeneous processor), e.g., as described herein in at least one embodiment. Figures 5 to 22 The API calls described above (e.g., create subcontext API 502, destroy subcontext API 702, get subcontext from stream API 902, get context resources API 1102, refine context resources API 1302, generate resource descriptor API 1502, get device resources API 1702, record context events API 1902, and / or wait for context events API 2102) are executed as described above. In at least one embodiment, after step 408, the process of executing one or more APIs shown in block diagram 400 continues at step 410.

[0139] In at least one embodiment, at step 410 of the process for executing one or more APIs shown in block diagram 400, the processor executing the process for executing one or more APIs shown in block diagram 400 determines whether a return value is returned as a result of executing one or more instructions for executing one or more API calls by the one or more processors and / or one or more other processors (e.g., a GPU and / or an accelerator within a heterogeneous processor as described above), such as described herein in conjunction with at least one embodiment. Figures 5 to 22 In at least one embodiment, at step 410, a processor executing the process for executing one or more of the APIs shown in block diagram 400 determines whether a return value will be returned using an API return, such as described herein in conjunction with at least one embodiment of the present invention. Figures 5 to 22The API described in step 410 returns (e.g., create subcontext API return 520, destroy subcontext API return 720, get subcontext from stream API return 920, get context resource API return 1120, refine context resource API return 1320, generate resource descriptor API return 1520, get device resource API return 1720, record context event API return 1920, and / or wait for context event API return 2120). In at least one embodiment, if it is determined that a return value is returned ("yes" branch), the process of executing one or more APIs shown in block diagram 400 continues at step 412. In at least one embodiment, if it is determined that a return value is not returned ("no" branch) at step 410, the process of executing one or more APIs shown in block diagram 400 continues at step 414.

[0140] In at least one embodiment, at step 412 of the process for executing one or more APIs shown in block diagram 400, a processor executing the process for executing one or more APIs shown in block diagram 400 sets a return value. In at least one embodiment, at step 412, the return value is stored in a memory used by the API (e.g., as described herein in conjunction with at least one embodiment). Figures 5 to 22 The return value is set in a memory location specified by the APIs described herein (e.g., create subcontext API 502, destroy subcontext API 702, get subcontext from stream API 902, get context resource API 1102, subdivide context resource API 1302, generate resource descriptor API 1502, get device resource API 1702, record context event API 1902, and / or wait for context event API 2102). In at least one embodiment, at step 412, the return value is stored in a memory location specified by the API return (e.g., at least in conjunction with the API described herein). Figures 5 to 22 The return value is set in a memory location included in any of the API returns described in the block diagram 400 (e.g., create subcontext API return 520, destroy subcontext API return 720, get subcontext from stream API return 920, get context resources API return 1120, refine context resources API return 1320, generate resource descriptor API return 1520, get device resources API return 1720, record context events API return 1920, and / or wait for context events API return 2120). In at least one embodiment, after step 412, the process of executing one or more APIs shown in block diagram 400 continues at step 414.

[0141] In at least one embodiment, at step 414 of the process for executing one or more APIs shown in block diagram 400, the processor executing the process for executing one or more APIs shown in block diagram 400 uses an API return (e.g., as described herein in at least some embodiments) to generate a response. Figures 5 to 22 The process of executing one or more of the APIs shown in block diagram 400 may return success or failure (e.g., an error) by executing one or more of the APIs shown in block diagram 400 (e.g., create subcontext API return 520, destroy subcontext API return 720, get subcontext from stream API return 920, get context resources API return 1120, refine context resources API return 1320, generate resource descriptor API return 1520, get device resources API return 1720, record context events API return 1920, and / or wait for context events API return 2120). In at least one embodiment, after step 414, the process of executing one or more of the APIs shown in block diagram 400 continues at step 416.

[0142] In at least one embodiment, at step 416 of the process for executing one or more APIs shown in block diagram 400, the processor executing the process for executing one or more APIs shown in block diagram 400 determines whether execution of the software program (e.g., started at step 404) is complete. In at least one embodiment, at step 416, the processor executing the process for executing one or more APIs shown in block diagram 400 determines whether execution of the software program (e.g., started at step 404) is complete based at least in part on whether the one or more processors are executing instructions of the software program (e.g., started at step 404). In at least one embodiment, at step 416, if it is determined that execution of the software program (e.g., started at step 404) is complete, the process for executing one or more APIs shown in block diagram 400 ends 418. In at least one embodiment, at step 416, if it is determined that execution of the software program (e.g., started at step 404) is not complete, the process for executing one or more APIs shown in block diagram 400 continues at step 404 to continue executing one or more instructions of the software program.

[0143] In at least one embodiment, the operations of the process of one or more APIs shown in block diagram 400 are performed to communicate with Figure 4In at least one embodiment, the operations of the process for executing one or more APIs shown in block diagram 400 are performed simultaneously or in parallel. In at least one embodiment, one or more operations that are not mutually dependent (e.g., are order-independent) for executing the process for executing one or more APIs shown in block diagram 400 are performed simultaneously or in parallel. In at least one embodiment, the operations of the process for executing one or more APIs shown in block diagram 400 are performed by multiple threads (e.g., threads described herein) executing on a processor.

[0144] In at least one embodiment, one or more processors (e.g., the processors described herein) include a processor for executing the instructions in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0145] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for executing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to deallocate one or more data structures indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0146] In at least one embodiment, one or more processors (e.g., such as the processors described herein) include a processor for executing the instructions in conjunction with the instructions herein. Figure 4 One or more circuits for performing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 4The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform operations described herein.

[0147] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, this document is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25One or more processes described herein that implement an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors to be available for executing one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to direct one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein.

[0148] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment, the present invention relates to a method, system, or process for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein. Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein implement an application programming interface (API) to cause identifiers of corresponding streaming multiprocessors (SMs) in one or more processors to be stored according to groups in which the corresponding SMs are to be used and / or otherwise perform operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58One or more components are described that implement an application programming interface (API) to cause identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform operations described herein.

[0149] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors on which to schedule one or more corresponding software threads and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding software threads thereon and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to schedule one or more corresponding sets of software threads thereon and / or otherwise perform the operations described herein.

[0150] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more indicators of one or more quantities of one or more streaming multiprocessors (SMs) in one or more processors to be read from one or more data structures storing the one or more indicators and / or otherwise perform operations described herein.

[0151] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, such as for executing an application programming interface (API) so that the context of one or more first software instructions is transferred to one or more second software instructions and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause the context of one or more first software instructions to be transferred to one or more second software instructions and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 4 Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described are configured to implement an application programming interface (API) to enable the context of one or more first software instructions to be transferred to one or more second software instructions and / or to otherwise perform the operations described herein.

[0152] In at least one embodiment, one or more processors (e.g., such as those described herein) include a processor for executing the instructions herein in conjunction with Figure 4 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 4 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to the one or more first software instructions and / or otherwise perform operations described herein. In at least one embodiment ( Figure 4Not shown), this paper combines Figure 4 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to execute an application programming interface (API) to cause one or more second instructions to wait to be executed until the one or more second instructions receive a context corresponding to one or more first software instructions and / or otherwise perform operations described herein.

[0153] Figure 5 5 is a block diagram illustrating an application programming interface (API) for creating a subcontext according to at least one embodiment. In at least one embodiment, one or more circuits of a processor are configured to execute a create subcontext API 502 to create a subcontext of a main context using resources of the main context. In at least one embodiment ( Figure 5 ), one or more circuits of a processor (such as those described herein) execute one or more instructions to execute a create subcontext API 502 to execute an application programming interface (API) to allocate one or more data structures indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads. In at least one embodiment ( Figure 5 ), one or more circuits of a processor (such as those described herein) execute one or more instructions to execute a create subcontext API 502 to execute an application programming interface (API) to generate one or more masks indicating one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment ( Figure 5 ), one or more circuits of a processor (such as those described herein) execute one or more instructions to execute a create subcontext API 502 to execute an application programming interface (API) to allocate one or more data structures indicating which of one or more streaming multiprocessors (SMs) of the one or more processors are to be used to execute one or more software threads in response to receiving a second API (such as the API described herein).

[0154] In at least one embodiment, the create subcontext API 502, when called, receives one or more arguments indicating information about an operation to be performed using techniques such as those described herein. In at least one embodiment, the create subcontext API 502, when called, receives one or more arguments indicating information about an instruction to be performed using techniques such as those described herein.

[0155] In at least one embodiment, create subcontext API 502 receives as input one or more parameters, including subcontext return 504. In at least one embodiment, subcontext return 504 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location for storing a created subcontext (e.g., a subcontext created using create subcontext API 502). In at least one embodiment, the subcontext return location identified, indicated, or otherwise specified by subcontext return 504 is one of a plurality of parameters available to create subcontext API 502 for creating a subcontext. In at least one embodiment, subcontext return 504 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., create subcontext API 502) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0156] In at least one embodiment, the create subcontext API 502 receives as input one or more parameters including a resource descriptor 506. In at least one embodiment, the resource descriptor 506 is a data value that includes information that can be used to identify, indicate, or otherwise specify the storage location of a resource descriptor used to create a subcontext (e.g., a subcontext created using the create subcontext API 502). In at least one embodiment, the resource descriptor identified, indicated, or otherwise specified by the resource descriptor 506 is one of a plurality of parameters available to the create subcontext API 502 for creating a subcontext. In at least one embodiment, the resource descriptor 506 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the create subcontext API 502) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0157] In at least one embodiment, create subcontext API 502 receives as input one or more parameters including device 508. In at least one embodiment, device 508 is a data value that includes information that can be used to identify, indicate, or otherwise specify a device associated with a subcontext (e.g., a subcontext created using create subcontext API 502). In at least one embodiment, the device identified, indicated, or otherwise specified by device 508 is one of a plurality of parameters available to create subcontext API 502 for creating a subcontext. In at least one embodiment, device 508 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., create subcontext API 502) a set of operations or instructions to be performed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0158] In at least one embodiment, create subcontext API 502 receives as input one or more parameters including flags 510. In at least one embodiment, flags 510 are data values that include information that can be used to identify, indicate, or otherwise specify flags to be used when creating a subcontext (e.g., using create subcontext API 502). In at least one embodiment, flags 510 are data values that indicate which SMs of a device (e.g., indicated by device 508) are to be used for the subcontext created by create subcontext API 502. In at least one embodiment, the flags identified, indicated, or otherwise specified by flags 510 are one of a plurality of parameters available to create subcontext API 502 for creating a subcontext. In at least one embodiment, flags 510 are data values that are used to identify, indicate, or otherwise specify to an API (e.g., create subcontext API 502) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0159] In at least one embodiment, create subcontext API 502 receives as input one or more parameters, including one or more other parameters 518. In at least one embodiment, other parameters 518 are data including information indicating any other information available when executing create subcontext API 502 to create a subcontext.

[0160] In at least one embodiment ( Figure 5), the processor executes one or more instructions to execute one or more APIs (e.g., create subcontext API 502) to execute an application programming interface (API) to allocate one or more data structures indicating which of one or more streaming multiprocessors (SMs) of the one or more processors are to be used to execute one or more software threads using one or more parameters, including but not limited to subcontext return 504, resource descriptor 506, device 508, flag 510, and / or other parameters 518. In at least one embodiment ( Figure 5 ), the processor executes one or more instructions to execute one or more APIs (e.g., create subcontext API 502) to execute an application programming interface (API) to generate one or more masks indicating one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more parameters, the one or more parameters including, but not limited to, a subcontext return 504, a resource descriptor 506, a device 508, a flag 510, and / or other parameters 518.

[0161] In at least one embodiment, the create subcontext API 502, if called, causes one or more APIs (e.g., Figure 3 In at least one embodiment, the create subcontext API 502, if called, causes one or more APIs (e.g., one or more APIs 306) to add one or more operations or instructions to be added, inserted, or otherwise included in a flow or instruction set executed by one or more accelerators within a heterogeneous processor. Figure 3 One or more operations or instructions are added to the parallel computing environment 308 described herein to add, insert, or otherwise include the operations or instructions in a stream or instruction set to be executed by one or more accelerators within the heterogeneous processor.

[0162] In at least one embodiment, in response to create subcontext API 502, one or more APIs 306, if executed, cause one or more processors to execute create subcontext API return 520. In at least one embodiment, create subcontext API return 520 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to create subcontext API 502. In at least one embodiment, create subcontext API return 520 indicates a success indicator 522. In at least one embodiment, success indicator 522 is data including any value indicating the success of create subcontext API 502. In at least one embodiment, success indicator 522 includes information indicating one or more specific types of success generated as a result of executing create subcontext API 502. In at least one embodiment, success indicator 522 includes information indicating one or more other data values generated as a result of executing create subcontext API 502.

[0163] In at least one embodiment, the create subcontext API return 520 indicates an error indicator 524. In at least one embodiment, the error indicator 524 is data including any value indicating a failure of the create subcontext API 502. In at least one embodiment, the error indicator 524 includes information indicating one or more specific types of errors generated as a result of executing the create subcontext API 502. In at least one embodiment, the error indicator 524 includes information indicating one or more other data values generated as a result of executing the create subcontext API 502.

[0164] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (which include but are not limited to a create subcontext API 502) adds various operations of various types to a stream for execution by one or more accelerators within a heterogeneous processor. In at least one embodiment, a stream operation includes an acquire semaphore operation. In at least one embodiment, a stream operation includes a release semaphore operation. In at least one embodiment, a stream operation includes one or more operations for refreshing cache memory and / or invalidating cache memory, such as the L2 cache memory of a PPU (e.g., a GPU) and / or the cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, a stream operation includes one or more operations for indicating that an operation is to be submitted to an external device (e.g., one or more accelerators within a heterogeneous processor). In at least one embodiment, example software code indicating a stream operation type is as follows:

[0165]

[0166] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to a create subcontext API 502) includes one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more operations cause the execution of the one or more callback functions. In at least one embodiment, example software code indicating a function signature for a callback function is as follows:

[0167] / **

[0168] * The callback function signature used to submit to the external device.

[0169] * /

[0170] typedef unsigned int(*cuSocketExternalDeviceSubmitCallback)(void*submitArgs);

[0171] In at least one embodiment, to specify one or more accelerators within a heterogeneous processor to perform one or more operations indicated by the create subcontext API 502 to the one or more APIs 306, one or more data structures of the one or more APIs 306 may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, example software code indicating a data structure representing a device node of one or more accelerators within a heterogeneous processor is as follows:

[0172]

[0173] In at least one embodiment, to specify the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor, one or more data structures of one or more APIs 306 are used. In at least one embodiment, example software code indicating a data structure for specifying the type and data of one or more operations to be performed by one or more accelerators within a heterogeneous processor is as follows:

[0174]

[0175]

[0176] In at least one embodiment, one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be added to a stream or other instruction set for execution by one or more accelerators within a heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set are executed in response to the create subcontext API 502, as described above. In at least one embodiment, example software code indicating a stream operation API call in a parallel computing environment 308 (e.g., CUDA) is as follows:

[0177]

[0178] In at least one embodiment, the one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs, similar to how one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor are added to one or more flows or instruction sets in response to the create subcontext API 502. In at least one embodiment, example software code that directs one or more APIs 306 of the parallel computing environment 308 to add one or more operations or instructions to one or more executable graphs is as follows:

[0179]

[0180]

[0181] Figure 6 6 is a block diagram 600 illustrating a process for executing an application programming interface (API) to create a subcontext according to at least one embodiment. In at least one embodiment, the process for executing the API to create a subcontext shown in block diagram 600 is a process for executing at least one embodiment of the present invention in conjunction with Figure 5 In at least one embodiment, part or all of the process of executing the API to create a subcontext (or any other process described herein, or variations and / or combinations thereof) shown in block diagram 600 is implemented on one or more computer systems, servers, processors, integrated circuits, and / or a combination thereof. Figures 26 to 58 The computer system, server, processor, integrated circuit and / or combination thereof is executed under the control of the other such devices described Figures 26 to 58Other such devices described are configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions that can be executed by one or more processors (e.g., the processors described herein). In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, the processor (e.g., at least one processor described herein) is a processor that is configured to execute computer-executable instructions. Figure 1 The processor 110 described herein (e.g., the processor 110 described herein) performs one or more steps of the process shown in block diagram 600 to execute the API to create a subcontext. In at least one embodiment, one or more other processors (e.g., the processors described herein) perform one or more steps of the process shown in block diagram 600 to execute the API to create a subcontext.

[0182] In at least one embodiment, in step 602 of the process of executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to receive or otherwise obtain the API for creating a subcontext. In at least one embodiment, in step 602, the processor receives or otherwise obtains the API for creating a subcontext. Figure 1 In at least one embodiment, in step 602, by a processor (e.g., processor 102, at least in conjunction with Figure 1 In at least one embodiment, in step 602, the API for creating a subcontext includes at least one embodiment of the present invention in conjunction with Figure 5 In at least one embodiment, after step 602 , the process of executing an API to create a subcontext shown in block diagram 600 proceeds to step 604 .

[0183] In at least one embodiment, in step 604 of the process of executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to determine whether the API for creating a subcontext (e.g., received in step 602) is valid. In at least one embodiment, in step 604, the operation of determining whether the API for creating a subcontext is valid includes: verifying the parameters of the API for creating a subcontext (e.g., the parameters of the API for creating a subcontext, as described herein at least in conjunction with the example of FIG. 1 ); Figure 5In at least one embodiment, if, at step 604, it is determined that the API for creating a subcontext is valid ("yes" branch), the process of executing the API to create a subcontext shown in block diagram 600 proceeds to step 606. In at least one embodiment, if, at step 604, it is determined that the API for creating a subcontext is invalid ("no" branch), the process of executing the API to create a subcontext shown in block diagram 600 proceeds to step 620, as described below.

[0184] In at least one embodiment, at step 606 of the process of executing an API to create a child context shown in block diagram 600, the processor executing the process performs one or more operations to obtain resources from the device. In at least one embodiment, at step 606, the one or more operations to obtain resources from the device include querying the context (e.g., primary context) of the device (e.g., device 202) (as described herein in at least conjunction with Figure 2 The invention further comprises one or more operations for determining a set of resources available for or otherwise associated with the device, as described herein in conjunction with at least Figure 2 In at least one embodiment, the device is a parameter of an API for creating a subcontext, as described herein in conjunction with at least Figure 5 In at least one embodiment, after step 606 , the process of executing an API to create a subcontext shown in block diagram 600 proceeds to step 608 .

[0185] In at least one embodiment, in step 608 of the process of executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to determine whether a subcontext can be created based at least in part on the resource list and / or one or more flags (e.g., received as parameters to the API for creating a subcontext, as described herein in at least some embodiments). Figure 5 In at least one embodiment, the one or more operations that determine whether a subcontext can be created are based at least in part on available resources, a source context (e.g., from which the subcontext is to be created), a device for creating the context, etc. In at least one embodiment, after step 608, the process of executing an API to create a subcontext shown in block diagram 600 proceeds to step 6104.

[0186] In at least one embodiment, in step 610 of the process for executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to determine whether a subcontext can be created (e.g., based on the determination in step 608). In at least one embodiment, in step 604, if it is determined that a subcontext can be created (the "yes" branch), the process for executing an API to create a subcontext shown in block diagram 600 proceeds to step 612. In at least one embodiment, in step 610, if it is determined that a subcontext cannot be created (the "no" branch), the process for executing an API to create a subcontext shown in block diagram 600 proceeds to step 620, as described below.

[0187] In at least one embodiment, in step 612 of the process of executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to create a subcontext based on the resources and flags received as parameters of the API for creating a subcontext. In at least one embodiment, in step 612, the subcontext is created, as described herein in at least some embodiments. Figure 1 and Figure 2 In at least one embodiment, after step 612 , the process of executing an API to create a subcontext shown in block diagram 600 proceeds to step 614 .

[0188] In at least one embodiment, in step 614 of the process of executing an API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to assign a subset of a resource set (e.g., of a main context and / or a subcontext) to the created subcontext. In at least one embodiment, in step 614, the subset of the resource set is exclusively assigned to the subcontext (e.g., to a single context). In at least one embodiment, the subset is assigned to the subcontext by, for example, providing a list of subsets of the resource set to the subcontext. In at least one embodiment, after step 614, the process of executing an API to create a subcontext shown in block diagram 600 proceeds to step 616.

[0189] In at least one embodiment, in step 616 of the process of executing the API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to return a success indicator (e.g., success indicator 522, which is described herein in at least some embodiments). Figure 5 In at least one embodiment, a success indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, described herein in at least one embodiment) at step 616. Figure 1In at least one embodiment, after step 616, the process of executing the API to create a subcontext shown in block diagram 600 terminates. Figure 6 (not shown), after step 616, the process of executing the API to create a subcontext shown in block diagram 600 continues to the above-mentioned step 602.

[0190] In at least one embodiment, in step 618 of the process of executing the API to create a subcontext shown in block diagram 600, the processor executing the process performs one or more operations to return an error indicator (e.g., error indicator 524, which is described herein in at least some embodiments). Figure 5 In at least one embodiment, at step 618, an error indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, as described herein in at least one embodiment). Figure 1 In at least one embodiment, after step 618, the process of executing the API to create a subcontext shown in block diagram 600 terminates. Figure 6 (not shown), after step 618, the process of executing the API to create a subcontext shown in block diagram 600 continues to the above-mentioned step 602.

[0191] In at least one embodiment, the operations of the process of executing an API to create a subcontext shown in block diagram 600 are performed in a different order than that shown in block diagram 600. In at least one embodiment, the operations of the process of executing an API to create a subcontext shown in block diagram 600 are performed simultaneously or in parallel. In at least one embodiment, the operations of the process of executing an API to create a subcontext shown in block diagram 600 that are independent of each other (e.g., order-independent) are performed simultaneously or in parallel. In at least one embodiment, the operations of the process of executing an API to create a subcontext shown in block diagram 600 are performed by multiple threads executing on a processor (e.g., the processor described herein).

[0192] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 5 and Figure 6One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations and / or instructions described herein in conjunction with Figure 5 and Figure 6 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to generate one or more masks for indicating one or more subsets of streaming multiprocessors in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 5 and Figure 6 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 5 and Figure 6 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 5 and Figure 6 Not shown), this paper combines Figure 5 and Figure 6 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to allocate one or more data structures for indicating which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0193] Figure 7 FIG7 is a block diagram 700 illustrating an application programming interface (API) for destroying a subcontext according to at least one embodiment. In at least one embodiment, one or more circuits of a processor are configured to execute a destroy subcontext API 702 to destroy a subcontext created by executing an API for creating a subcontext (e.g., as described herein in conjunction with at least one embodiment). Figure 5 In at least one embodiment ( Figure 7 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a destroy subcontext API 702 to execute an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) in the one or more processors are to be used to execute one or more software threads. In at least one embodiment ( Figure 7 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a destroy subcontext API 702 to execute an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that may be used to execute one or more kernels. In at least one embodiment ( Figure 7 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a destroy subcontext API 702 to execute an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) in the one or more processors are to be used to execute the one or more software threads in response to receiving a second API (e.g., an API described herein).

[0194] In at least one embodiment, the destroy subcontext API 702, when called, receives one or more arguments indicating information about an operation to be performed using techniques such as those described herein. In at least one embodiment, the destroy subcontext API 702, when called, receives one or more arguments indicating information about an instruction to be performed using techniques such as those described herein.

[0195] In at least one embodiment, the destroy subcontext API 702 receives as input one or more parameters including a subcontext 704. In at least one embodiment, the subcontext 704 is a data value that includes information that can be used to identify, indicate, or otherwise specify a subcontext to be destroyed using the destroy subcontext API 702. In at least one embodiment, the subcontext identified, indicated, or otherwise specified by the subcontext 704 is one of a plurality of parameters that the destroy subcontext API 702 can use to destroy a subcontext. In at least one embodiment, the subcontext 704 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the destroy subcontext API 702) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0196] In at least one embodiment, the destroy subcontext API 702 receives as input one or more parameters, including one or more other parameters 718. In at least one embodiment, the other parameters 718 are data that include information indicating any other information that may be used to execute the destroy subcontext API 702 to destroy the subcontext.

[0197] In at least one embodiment ( Figure 7 ), the processor executes one or more instructions to execute one or more APIs (e.g., destroy subcontext API 702), thereby executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) in the one or more processors are to be used to execute one or more software threads using one or more parameters including, but not limited to, subcontext 704 and / or other parameters 718. In at least one embodiment ( Figure 7 ), the processor executes one or more instructions to execute one or more APIs (e.g., destroy subcontext API 702), thereby executing an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more parameters including, but not limited to, subcontext 704 and / or other parameters 718.

[0198] In at least one embodiment, the destroy subcontext API 702, if called, causes one or more APIs (e.g., Figure 3In at least one embodiment, the destroy subcontext API 702, if called, causes one or more APIs (e.g., one or more APIs 306) to be added, inserted, or otherwise included in a flow or instruction set to be executed by one or more accelerators within a heterogeneous processor. Figure 3 One or more operations or instructions are added to the parallel computing environment 308 described herein to add, insert, or otherwise include the operations or instructions in a stream or instruction set to be executed by one or more accelerators within the heterogeneous processor.

[0199] In at least one embodiment, in response to the destroy subcontext API 702, one or more APIs 306, if executed, cause one or more processors to execute a destroy subcontext API return 720. In at least one embodiment, the destroy subcontext API return 720 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the destroy subcontext API 702. In at least one embodiment, the destroy subcontext API return 720 indicates a success indicator 722. In at least one embodiment, the success indicator 722 is data including any value indicating the success of the destroy subcontext API 702. In at least one embodiment, the success indicator 722 includes information indicating one or more specific types of success generated as a result of executing the destroy subcontext API 702. In at least one embodiment, the success indicator 722 includes information 702 indicating one or more other data values generated as a result of executing the destroy subcontext API 702.

[0200] In at least one embodiment, the destroy subcontext API return 720 indicates an error indicator 724. In at least one embodiment, the error indicator 724 is data including any value indicating a failure of the destroy subcontext API 702. In at least one embodiment, the error indicator 724 includes information indicating one or more specific types of errors generated as a result of executing the destroy subcontext API 702. In at least one embodiment, the error indicator 724 includes information indicating one or more other data values generated as a result of executing the destroy subcontext API 702.

[0201] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to a destroy subcontext API 702) adds various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes an acquire semaphore operation. In at least one embodiment, the stream operation includes a release semaphore operation. In at least one embodiment, the stream operation includes one or more operations for flushing and / or invalidating cache memory, such as the L2 cache memory of a PPU (e.g., a GPU) and / or the cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes one or more operations for indicating submission of an operation to an external device (e.g., one or more accelerators within a heterogeneous processor). In at least one embodiment, the one or more operations for indicating submission of an operation to an external device use software code, such as at least in combination with Figure 5 Example software code indicating stream operations is described.

[0202] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to a destroy subcontext API 702) includes one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more operations cause the one or more callback functions to be executed. In at least one embodiment, the one or more operations that cause the one or more callback functions to be executed use software code, such as example software code indicating a function signature of a callback function, as described herein in at least conjunction with Figure 5 Described.

[0203] In at least one embodiment, to specify that one or more accelerators within a heterogeneous processor perform one or more operations indicated by the destroy subcontext API 702 to the one or more APIs 306, one or more data structures of the one or more APIs 306 may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, the one or more data structures of the one or more APIs 306 that may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations use software code, such as example software code indicating a data structure representing a device node of one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 Described.

[0204] In at least one embodiment, to specify the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 for specifying the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor use software code, such as example software code indicating data structures for specifying the type and data of one or more operations to be performed by one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 Described.

[0205] In at least one embodiment, one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be added to a stream or other instruction set to be executed by one or more accelerators within the heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set are executed in response to the destroy subcontext API 702, as described above. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set that are executed in response to the destroy subcontext API 702 utilize software code, such as example software code that directs stream operation API calls in the parallel computing environment 308, as described at least in conjunction with the present disclosure. Figure 5 Described.

[0206] In at least one embodiment, the one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs are similar to the manner in which one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor are added to one or more streams or instruction sets in response to destroying a subcontext API 702 as described herein. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs utilize software code, such as example software code that instructs one or more APIs 306 of the parallel computing environment 308 to add one or more operations or instructions to one or more executable graphs, as described herein in conjunction with at least one embodiment. Figure 5 Described.

[0207] Figure 8 800 is a block diagram illustrating a process for executing an application programming interface (API) to destroy a subcontext according to at least one embodiment. In at least one embodiment, the process for executing an API to destroy a subcontext shown in block diagram 800 is a process for executing the destroy subcontext API 702, which is described herein in conjunction with at least one embodiment. Figure 7 In at least one embodiment, part or all of the process shown in block diagram 800 for executing an API to destroy a subcontext (or any other process described herein, or variations and / or combinations thereof) is implemented on one or more computer systems, servers, processors, integrated circuits, and / or a combination thereof. Figures 26 to 58 The computer system, server, processor, integrated circuit and / or combination thereof is executed under the control of the other such devices described Figures 26 to 58 Other such devices described are configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions that can be executed by one or more processors (e.g., the processors described herein). In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, the processor (e.g., at least one processor described herein) is a processor that executes the computer-readable instructions. Figure 1 The processor 110 described herein performs one or more steps of the process for executing an API to destroy a subcontext shown in block diagram 800. In at least one embodiment, one or more other processors (e.g., the processors described herein) perform one or more steps of the process for executing an API to destroy a subcontext shown in block diagram 800.

[0208] In at least one embodiment, in step 802 of the process for executing an API to destroy a subcontext shown in block diagram 800, a processor executing the process performs one or more operations to receive or otherwise obtain an API for destroying a subcontext. In at least one embodiment, in step 802, the API is provided to the processor in other ways (e.g., as described herein in at least one embodiment). Figure 1 In at least one embodiment, in step 802, by the processor (e.g., at least in conjunction with the present invention) receiving an API for destroying the subcontext. Figure 1In at least one embodiment, in step 802, the API for destroying the subcontext includes one or more parameters (e.g., at least one of the parameters described herein). Figure 7 In at least one embodiment, after step 802 , the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 804 .

[0209] In at least one embodiment, in step 804 of the process for executing an API to destroy a subcontext shown in block diagram 800, a processor executing the process performs one or more operations to determine whether the API for destroying a subcontext (e.g., received in step 802) is valid. In at least one embodiment, in step 804, determining whether the API for destroying a subcontext is valid includes: verifying parameters of the API for destroying a subcontext (e.g., Figure 7 In at least one embodiment, if it is determined at step 804 that the API for destroying the subcontext is valid ("yes" branch), the process for executing the API to destroy the subcontext shown in block diagram 800 proceeds to step 806. In at least one embodiment, if it is determined at step 804 that the API for destroying the subcontext is invalid ("no" branch), the process for executing the API to destroy the subcontext shown in block diagram 800 proceeds to step 816 described below.

[0210] In at least one embodiment, at step 806 of the process for executing an API to destroy a subcontext shown in block diagram 800, a processor executing the process performs one or more operations to identify a subcontext to be destroyed. In at least one embodiment, the subcontext is identified as a parameter of the API for destroying a subcontext received in step 802. In at least one embodiment, after step 806, the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 808.

[0211] In at least one embodiment, in step 808 of the process for executing an API to destroy a subcontext shown in block diagram 800, the processor executing the process performs one or more operations to determine whether a subcontext to be destroyed is identified (e.g., identified in step 806). In at least one embodiment, in step 808, if it is determined that a subcontext to be destroyed is identified (the "yes" branch), the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 810. In at least one embodiment, in step 808, if it is determined that a subcontext to be destroyed is not identified (the "no" branch), the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 816, as described below.

[0212] In at least one embodiment, in step 810 of the process for executing an API for destroying a subcontext shown in block diagram 800, the processor executing the process performs one or more operations to stop one or more flows associated with the context to be destroyed (e.g., the subcontext identified in step 806), release any resources associated with the context to be destroyed, and destroy the context to be destroyed (e.g., indicating that the context is invalid). In at least one embodiment, the flows are allowed to complete their current operations before being stopped. In at least one embodiment, the flows are stopped immediately (e.g., without being allowed to complete their current operations). In at least one embodiment, multiple flows associated with the subcontext are stopped. In at least one embodiment, after step 810, the process for executing an API for destroying a subcontext shown in block diagram 800 proceeds to step 812.

[0213] In at least one embodiment, in step 812 of the process for executing an API to destroy a subcontext shown in block diagram 800, the processor executing the process performs one or more operations to determine whether the subcontext to be destroyed has been destroyed (e.g., by executing step 710). In at least one embodiment, in step 812, if it is determined that the context to be destroyed has been destroyed (the "yes" branch), the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 814. In at least one embodiment, in step 812, if it is determined that the context to be destroyed has not been destroyed (the "no" branch), the process for executing an API to destroy a subcontext shown in block diagram 800 proceeds to step 816, as described below.

[0214] In at least one embodiment, in step 814 of the process for executing the API to destroy the subcontext shown in block diagram 800, the processor executing the process performs one or more operations to return a success indicator (e.g., success indicator 722, which is described herein in conjunction with at least one embodiment). Figure 7 In at least one embodiment, in step 814, a success indicator is returned to the calling process (e.g., a process executed by a processor, such as processor 102, herein at least in conjunction with Figure 1 In at least one embodiment, after step 814, the process for executing the API to destroy the subcontext shown in block diagram 800 terminates. Figure 8 (not shown), after step 814, the process for executing the API to destroy the subcontext shown in block diagram 800 continues to the above-mentioned step 802.

[0215] In at least one embodiment, in step 816 of the process for executing the API to destroy the subcontext shown in block diagram 800, the processor executing the process performs one or more operations to return an error indicator (e.g., error indicator 724, which is described herein in at least some embodiments). Figure 7 In at least one embodiment, at step 816, an error indicator is returned to the calling process (e.g., a process executed by a processor, such as processor 102, herein at least in conjunction with Figure 1 In at least one embodiment, after step 816, the process for executing the API to destroy the subcontext shown in block diagram 800 terminates. Figure 8 (not shown), after step 816, the process for executing the API to destroy the subcontext shown in block diagram 800 continues with the above-mentioned step 802.

[0216] In at least one embodiment, the operations shown in block diagram 800 for executing the process for an API to destroy a subcontext are executed in a different order than that shown in block diagram 800. In at least one embodiment, the operations shown in block diagram 800 for executing the process for an API to destroy a subcontext are executed simultaneously or in parallel. In at least one embodiment, the operations shown in block diagram 800 for executing the process for an API to destroy a subcontext that are independent of each other (e.g., order-independent) are executed simultaneously or in parallel. In at least one embodiment, the operations shown in block diagram 800 for executing the process for an API to destroy a subcontext are executed by multiple threads executing on a processor (e.g., a processor described herein).

[0217] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 7 and Figure 8One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) in one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for executing the operations and / or instructions described herein in conjunction with Figure 7 and Figure 8 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more masks to be disabled, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 7 and Figure 8 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 Described herein are systems, methods, operations, and / or instructions for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform operations described herein. In at least one embodiment, the present invention is combined with Figure 7 and Figure 8 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) of one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 7 and Figure 8 Not shown), this paper combines Figure 7 and Figure 8 One or more components described herein include Figures 26 to 58 One or more components described herein are configured to implement an application programming interface (API) to deallocate one or more data structures used to indicate which of one or more streaming multiprocessors (SMs) among one or more processors are to be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0218] Figure 9 900 is a block diagram illustrating an application programming interface (API) for obtaining a subcontext from a flow according to at least one embodiment. In at least one embodiment, one or more circuits of a processor are configured to execute the obtain subcontext from flow API 902, using a channel to obtain a subcontext associated with an identified flow being executed, as described herein in conjunction with at least Figure 1 and Figure 2 In at least one embodiment ( Figure 9 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute the Get Subcontext from Stream API 902, thereby executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in the one or more processors to be stored. In at least one embodiment ( Figure 9 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a get subcontext from stream API 902, thereby executing an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels. In at least one embodiment ( Figure 9 ), one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a get subcontext from stream API 902, thereby executing an application programming interface (API) such that one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in the one or more processors are stored in response to receiving a second API (e.g., an API described herein).

[0219] In at least one embodiment, the get subcontext from stream API 902, when called, receives one or more arguments indicating information about an operation to be performed using techniques such as those described herein. In at least one embodiment, the get subcontext from stream API 902, when called, receives one or more arguments indicating information about an instruction to be performed using techniques such as those described herein.

[0220] In at least one embodiment, the get subcontext from stream API 902 receives as input one or more parameters, including a stream 904. In at least one embodiment, the stream 904 is a data value that includes information that can be used to identify, indicate, or otherwise specify a stream from which a subcontext is to be obtained using the get subcontext from stream API 902. In at least one embodiment, the stream identified, indicated, or otherwise specified by the stream 904 is one of a plurality of parameters that can be used by the get subcontext from stream API 902 to obtain a subcontext from a stream. In at least one embodiment, the stream 904 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the get subcontext from stream API 902) a set of operations or instructions to be performed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0221] In at least one embodiment, the get subcontext from stream API 902 receives as input one or more parameters, including a subcontext return 906. In at least one embodiment, the subcontext return 906 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location for storing a subcontext identified using the get subcontext from stream API 902. In at least one embodiment, the storage location for storing the subcontext identified, indicated, or otherwise specified by the subcontext return 906 is one of a plurality of parameters that can be used by the get subcontext from stream API 902 to obtain a subcontext from a stream. In at least one embodiment, the subcontext return 906 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the get subcontext from stream API 902) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0222] In at least one embodiment, the get subcontext from stream API 902 receives as input one or more parameters including one or more other parameters 918. In at least one embodiment, the other parameters 918 are data including information indicating any other information available when executing the get subcontext from stream API 902 to obtain a subcontext from a stream.

[0223] In at least one embodiment ( Figure 9), the processor executes one or more instructions to execute one or more APIs (e.g., get subcontext from stream API 902), thereby executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in the one or more processors to be stored using one or more parameters (including but not limited to stream 904, subcontext return 906, and / or other parameters 918). In at least one embodiment ( Figure 9 ), the processor executes one or more instructions to execute one or more APIs (e.g., get subcontext from stream API 902), thereby executing an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels using one or more parameters including, but not limited to, stream 904, subcontext return 906, and / or other parameters 918.

[0224] In at least one embodiment, the Get Subcontext from Stream API 902, if called, causes one or more APIs (e.g., Figure 3 In at least one embodiment, the get subcontext from stream API 902, if called, causes one or more APIs (e.g., one or more APIs 306) to be added, inserted, or otherwise included in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. Figure 3 One or more operations or instructions are added to the parallel computing environment 308 described herein to add, insert, or otherwise include the operations or instructions in a stream or instruction set to be executed by one or more accelerators within the heterogeneous processor.

[0225] In at least one embodiment, in response to the Get Subcontext from Stream API 902, one or more APIs 306, if executed, cause one or more processors to execute the Get Subcontext from Stream API return 920. In at least one embodiment, the Get Subcontext from Stream API return 920 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the Get Subcontext from Stream API 902. In at least one embodiment, the Get Subcontext from Stream API return 920 indicates a success indicator 922. In at least one embodiment, the success indicator 922 is data including any value indicating the success of the Get Subcontext from Stream API 902. In at least one embodiment, the success indicator 922 includes information indicating one or more specific types of success generated as a result of executing the Get Subcontext from Stream API 902. In at least one embodiment, the success indicator 922 includes information indicating one or more other data values generated as a result of executing the Get Subcontext from Stream API 902.

[0226] In at least one embodiment, the get subcontext from stream API return 920 indicates an error indicator 924. In at least one embodiment, the error indicator 924 is data including any value indicating a failure of the get subcontext from stream API 902. In at least one embodiment, the error indicator 924 includes information indicating one or more specific types of errors generated as a result of the get subcontext from stream API 902. In at least one embodiment, the error indicator 924 includes information indicating one or more other data values generated as a result of executing the get subcontext from stream API 902.

[0227] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to a get subcontext from stream API 902) adds various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes an acquire semaphore operation. In at least one embodiment, the stream operation includes a release semaphore operation. In at least one embodiment, the stream operation includes one or more operations for flushing and / or invalidating cache memory, such as the L2 cache memory of a PPU (e.g., a GPU) and / or the cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes one or more operations for indicating submission of an operation to an external device (e.g., one or more accelerators within a heterogeneous processor). In at least one embodiment, the one or more operations for indicating submission of an operation to an external device use software code, such as described in at least one embodiment in conjunction with Figure 5 Example software code for the indicated flow operations.

[0228] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to the get subcontext from stream API 902) includes one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more operations cause the execution of the one or more callback functions. In at least one embodiment, the one or more operations that cause the execution of the one or more callback functions use software code, such as example software code indicating a function signature of the callback function, as described herein in at least conjunction with Figure 5 As stated.

[0229] In at least one embodiment, to specify that one or more accelerators within a heterogeneous processor perform one or more operations indicated by the get subcontext from stream API 902 to one or more APIs 306, one or more data structures of one or more APIs 306 may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, the one or more data structures of one or more APIs 306 that may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations utilize software code, such as example software code indicating data structures representing device nodes of one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 As stated.

[0230] In at least one embodiment, to specify the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 for specifying the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor use software code, for example, example software code indicating data structures for specifying the type and data of one or more operations to be performed by one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 As stated.

[0231] In at least one embodiment, one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be added to a stream or other instruction set to be executed by one or more accelerators within the heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set are executed in response to the get subcontext from stream API 902, as described above. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set that are executed in response to the get subcontext from stream API 902 utilize software code, such as example software code that directs stream operation API calls in the parallel computing environment 308, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0232] In at least one embodiment, the one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs are similar to the manner in which one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor are added to one or more streams or instruction sets in response to the get subcontext from stream API 902, as described herein. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs utilize software code, for example, example software code that instructs the one or more APIs 306 of the parallel computing environment 308 to add one or more operations or instructions to one or more executable graphs, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0233] Figure 10 1000 is a block diagram illustrating a process for executing an application programming interface (API) to obtain a subcontext from a stream according to at least one embodiment. In at least one embodiment, the process for executing the API to obtain a subcontext from a stream illustrated in block diagram 1000 is a process for executing the obtain subcontext from stream API 902, which is described herein in conjunction with at least one embodiment. Figure 9 In at least one embodiment, part or all of the process shown in block diagram 1000 for executing an API to obtain a subcontext from a stream (or any other process described herein, or variations and / or combinations thereof) is implemented on one or more computer systems, servers, processors, integrated circuits, and / or a combination thereof. Figures 26 to 58One or more computer systems, servers, processors, integrated circuits and / or combinations thereof executed under the control of such other devices Figures 26 to 58 Other such devices described are configured with computer executable instructions and implemented as code (e.g., computer executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions that can be executed by one or more processors (e.g., the processors described herein). In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, the processor (e.g., at least one processor described herein) is a computer-readable storage medium that is configured to execute computer executable instructions. Figure 1 The processor 110 described herein performs one or more steps of the process for executing an API to obtain a subcontext from a stream as shown in block diagram 1000. In at least one embodiment, one or more other processors (e.g., the processors described herein) perform one or more steps of the process for executing an API to obtain a subcontext from a stream as shown in block diagram 1000.

[0234] In at least one embodiment, in step 1002 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, a processor executing the process performs one or more operations to receive or otherwise obtain an API for obtaining a subcontext from a stream. In at least one embodiment, in step 1002, the API for obtaining a subcontext from a stream is provided to the processor in other ways (e.g., as described herein in at least one embodiment). Figure 1 In at least one embodiment, in step 1002, the processor (e.g., at least one embodiment of the present invention) receives the received Figure 1 In at least one embodiment, in step 1002, the API for obtaining a subcontext from a stream includes one or more parameters (e.g., at least one of the parameters described herein). Figure 9 In at least one embodiment, after step 1002 , the process for executing an API to obtain a subcontext from a stream as shown in block diagram 1000 proceeds to step 1004 .

[0235] In at least one embodiment, in step 1004 of the process for obtaining an API for a subcontext from a stream shown in block diagram 1000, a processor executing the process performs one or more operations to determine whether the API for obtaining a subcontext from a stream (e.g., received in step 1002) is valid. In at least one embodiment, in step 1004, the operation of determining whether the API for obtaining a subcontext from a stream is valid includes validating parameters of the API for obtaining a subcontext from a stream (e.g., Figure 9 In at least one embodiment, if, at step 1004, it is determined that the API for obtaining a subcontext from a stream is valid (the "yes" branch), the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1006. In at least one embodiment, if, at step 1004, it is determined that the API for obtaining a subcontext from a stream is invalid (the "no" branch), the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1014, described below.

[0236] In at least one embodiment, at step 1006 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, a processor executing the process performs one or more operations to identify a subcontext from a stream (e.g., the stream indicated by the argument of the API for obtaining a subcontext from a stream). In at least one embodiment, the subcontext maintains a list of streams associated with the context. In at least one embodiment, each stream maintains a list of subcontexts associated with the stream. In at least one embodiment, after step 1006, the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1008.

[0237] In at least one embodiment, in step 1008 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, the processor executing the process performs one or more operations to determine whether a subcontext is identified (e.g., in step 1006). In at least one embodiment, in step 1008, if it is determined that a subcontext is identified (the "yes" branch), the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1010. In at least one embodiment, in step 1008, if it is determined that a subcontext is not identified (the "no" branch), the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1014, as described below.

[0238] In at least one embodiment, in step 1010 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, a processor executing the process performs one or more operations to store the identified subcontext (e.g., identified in step 1006) in a storage location indicated by the API for obtaining a subcontext from a stream (e.g., received as a parameter of the API for obtaining a subcontext from a stream, received in step 1002). In at least one embodiment, after step 1010, the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000 proceeds to step 1012.

[0239] In at least one embodiment, in step 1012 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, the processor executing the process performs one or more operations to return a success indicator (e.g., success indicator 922, which is described herein in at least some embodiments). Figure 9 In at least one embodiment, at step 1012, a success indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, as described herein in at least one embodiment). Figure 1 In at least one embodiment, after step 1012, the process for executing the API to obtain a subcontext from a stream as shown in block diagram 1000 terminates. Figure 6 (not shown), after step 1012, the process for executing the API to obtain the subcontext from the stream shown in block diagram 1000 continues to the above-mentioned step 1002.

[0240] In at least one embodiment, in step 1014 of the process for executing an API to obtain a subcontext from a stream shown in block diagram 1000, the processor executing the process performs one or more operations to return an error indicator (e.g., error indicator 924, which is described herein in at least some embodiments). Figure 9 In at least one embodiment, at step 1014, an error indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, as described herein). Figure 1 In at least one embodiment, after step 1014, the process for executing the API to obtain a subcontext from a stream as shown in block diagram 1000 terminates. Figure 10 ), after step 1014, the process for executing the API to obtain the subcontext from the stream shown in block diagram 1000 continues to the above-mentioned step 1002.

[0241] In at least one embodiment, the operations of the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 are performed in a different order than that shown in block diagram 1000. In at least one embodiment, the operations of the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 are performed simultaneously or in parallel. In at least one embodiment, the mutually independent (e.g., order-independent) operations of the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 are performed simultaneously or in parallel. In at least one embodiment, the operations of the process for executing the API to obtain a subcontext from a stream shown in block diagram 1000 are performed by multiple threads executing on a processor (e.g., a processor described herein).

[0242] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 9 and Figure 10 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or one or more circuits for otherwise performing the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations and / or instructions described herein in conjunction with Figure 9 and Figure 10 One or more circuits for performing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to indicate one or more identifiers of one or more masks, wherein the one or more masks indicate one or more streaming multiprocessor subsets in one or more graphics processing units (GPUs) that can be used to execute one or more kernels and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 9 and Figure 10 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or otherwise include Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 9 and Figure 10 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein that implement an application programming interface (API) to cause one or more identifiers of one or more data structures indicating the number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform the operations described herein. In at least one embodiment ( Figure 9 and Figure 10 Not shown), this paper combines Figure 9 and Figure 10 One or more components described herein include Figures 26 to 58 One or more components described herein for executing an application programming interface (API) to cause one or more identifiers of one or more data structures indicating a number of one or more streaming multiprocessors (SMs) in one or more processors to be stored and / or otherwise perform operations described herein.

[0243] Figure 11 1 is a block diagram illustrating an application programming interface (API) for obtaining resources associated with a context according to at least one embodiment. In at least one embodiment, one or more circuits of a processor are configured to execute an obtain context resource API 1102 to obtain resources associated with a context. In at least one embodiment, Figure 11 Not shown, one or more circuits of a processor (such as the processors described herein) are configured to execute one or more instructions to execute the Get Context Resource API 1102, thereby executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more processors available for executing one or more software threads. In at least one embodiment, Figure 11 100 , one or more circuits of a processor (e.g., a processor described herein) are configured to execute one or more instructions to execute a get context resource API 1102, thereby executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more graphics processing units (GPUs) that are available for executing one or more software kernels. In at least one embodiment, Figure 11 Also not shown, one or more circuits of a processor (such as the processor described herein) are used to execute one or more instructions to execute the obtain context resource API 1102, thereby executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads in response to receiving a second API (such as those described herein).

[0244] In at least one embodiment, the Get Context Resource API 1102, when called, receives one or more parameters indicating information about an operation to be performed using techniques such as those described herein. In at least one embodiment, the Get Context Resource API 1102, when called, receives one or more parameters indicating information about an instruction to be performed using techniques such as those described herein.

[0245] In at least one embodiment, the get context resources API 1102 receives as input one or more parameters, including a resource return 1104. In at least one embodiment, the resource return 1104 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location for storing resources (e.g., resource descriptors) associated with the context using the get context resources API 1102. In at least one embodiment, the storage location for storing the resources identified, indicated, or otherwise specified by the resource return 1104 is one of a plurality of parameters that can be used by the get context resources API 1102 to obtain resources associated with the context. In at least one embodiment, the resource return 1104 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the get context resources API 1102) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0246] In at least one embodiment, the get context resources API 1102 receives as input one or more parameters including a context 1106. In at least one embodiment, the context 1106 is a data value that includes information that can be used to identify, indicate, or otherwise specify a context or subcontext from which a set of resources can be identified and returned using the get context resources API 1102. In at least one embodiment, the context or subcontext identified, indicated, or otherwise specified by the context 1106 is one of a plurality of parameters that can be used by the get context resources API 1102 to obtain resources associated with the context. In at least one embodiment, the context 1106 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the get context resources API 1102) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0247] In at least one embodiment, the get context resource API 1102 receives as input one or more parameters, including one or more other parameters 1118. In at least one embodiment, the other parameters 1118 are data containing information indicating any other information available when executing the get context resource API 1102 to obtain resources associated with the context.

[0248] In at least one embodiment, Figure 11 Not shown, the processor executes one or more instructions to execute one or more APIs (e.g., get context resource API 1102), thereby executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads using one or more parameters (including but not limited to resource return 1104, context 1106, and / or other parameters 1118). In at least one embodiment, Figure 11 Not shown, the processor executes one or more instructions to execute one or more APIs (e.g., get context resource API 1102), thereby executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) available for executing one or more software kernels using one or more parameters (including but not limited to resource return 1104, context 1106 and / or other parameters 1118).

[0249] In at least one embodiment, the Get Context Resource API 1102, if called, causes one or more APIs (e.g., Figure 3 In at least one embodiment, the get context resource API 1102, if called, causes one or more APIs (e.g., one or more APIs 306) to be added, inserted, or otherwise included in a flow or instruction set executed by one or more accelerators within a heterogeneous processor. Figure 3 One or more operations or instructions are added to the parallel computing environment 308 described herein to be added, inserted into, or otherwise included in a flow or instruction set executed by one or more accelerators within a heterogeneous processor.

[0250] In at least one embodiment, in response to the Get Context Resource API 1102, one or more APIs 306 (if executed) are used to cause one or more processors to execute the Get Context Resource API return 1120. In at least one embodiment, the Get Context Resource API return 1120 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the Get Context Resource API 1102. In at least one embodiment, the Get Context Resource API return 1120 indicates a success indicator 1122. In at least one embodiment, the success indicator 1122 is data including any value indicating the success of the Get Context Resource API 1102. In at least one embodiment, the success indicator 1122 includes information indicating one or more specific types of success generated as a result of executing the Get Context Resource API 1102. In at least one embodiment, the success indicator 1122 includes information indicating one or more other data values generated as a result of executing the Get Context Resource API 1102.

[0251] In at least one embodiment, the get context resource API return 1120 indicates an error indicator 1124. In at least one embodiment, the error indicator 1124 is data including any value indicating a failure of the get context resource API 1102. In at least one embodiment, the error indicator 1124 includes information indicating one or more specific types of errors generated as a result of executing the get context resource API 1102. In at least one embodiment, the error indicator 1124 includes information indicating one or more other data values generated as a result of executing the get context resource API 1102.

[0252] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to the Get Context Resource API 1102) adds various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes an Get Semaphore operation. In at least one embodiment, the stream operation includes a Release Semaphore operation. In at least one embodiment, the stream operation includes one or more operations for flushing and / or invalidating cache memory, such as the L2 cache of a PPU (e.g., a GPU) and / or the cache of one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes one or more operations for indicating submission of an operation to an external device (e.g., one or more accelerators within a heterogeneous processor). In at least one embodiment, the one or more operations for indicating submission of an operation to an external device use software code, such as at least in combination with Figure 5Example software code indicating stream operations is described.

[0253] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to the get context resource API 1102) includes one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more operations cause the execution of the one or more callback functions. In at least one embodiment, the one or more operations that cause the execution of the one or more callback functions use software code, such as example software code indicating a function signature of the callback function, as described herein in at least conjunction with Figure 5 As stated.

[0254] In at least one embodiment, to specify one or more accelerators within a heterogeneous processor to perform one or more operations indicated by the get context resource API 1102 to the one or more APIs 306, one or more data structures of the one or more APIs 306 may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, the one or more data structures of the one or more APIs 306 that may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations use software code, such as example software code indicating a data structure representing a device node of one or more accelerators within a heterogeneous processor, as described at least herein in conjunction with the figures.

[0255] In at least one embodiment, to specify the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 for specifying the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor use software code, such as example software code indicating data structures for specifying the type and data of one or more operations to be performed by one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 As stated.

[0256] In at least one embodiment, one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be added to a stream or other instruction set to be executed by one or more accelerators within the heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set are executed in response to the Get Context Resource API 1102, as described above. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set that are executed in response to the Get Context Resource API 1102 utilize software code, such as example software code that directs stream operation API calls in the parallel computing environment 308, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0257] In at least one embodiment, the one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs. In at least one embodiment, the instructions, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs similar to the manner in which one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor are added to one or more streams or instruction sets in response to the get context resource API 1102 described herein. In at least one embodiment, the instructions, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs utilize software code, such as example software code that instructs one or more APIs 306 of the parallel computing environment 308 to add one or more operations or instructions to one or more executable graphs, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0258] Figure 12 1200 is a block diagram illustrating a process for executing an application programming interface (API) to obtain resources associated with a context according to at least one embodiment. In at least one embodiment, the process for executing an API to obtain resources associated with a context shown in block diagram 1200 is a process described herein in conjunction with at least one embodiment of the present invention. Figure 11 In at least one embodiment, part or all of the process for executing an API to obtain resources associated with a context (or any other process described herein, or variations and / or combinations thereof) shown in block diagram 1200 is implemented on one or more computer systems, servers, processors, integrated circuits, and / or a combination thereof. Figures 26 to 58The present invention relates to a computer program that is executed under the control of such devices described herein, which are configured with computer executable instructions and implemented as code (e.g., computer executable instructions, one or more computer programs, or one or more applications) that is jointly executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes multiple computer-readable instructions that can be executed by one or more processors (e.g., the processors described herein). In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, the processor (e.g., at least one processor described herein) is a computer program that is executed by one or more processors (e.g., the processors described herein). Figure 1 The processor 110 described herein performs one or more steps of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context. In at least one embodiment, one or more other processors (e.g., the processors described herein) perform one or more steps of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context.

[0259] In at least one embodiment, in step 1202 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, a processor executing the process performs one or more operations to receive or otherwise obtain an API for obtaining a resource associated with a context. In at least one embodiment, in step 1202, the processor receives or otherwise obtains an API for obtaining a resource associated with a context. Figure 1 In at least one embodiment, at step 1202, the processor (e.g., the processor 110 described herein) receives an API for obtaining resources associated with the context. Figure 1 The processor 102 described herein is provided to receive an API for obtaining resources associated with a context. In at least one embodiment, in step 1202, the API for obtaining resources associated with a context includes one or more parameters, such as at least one of the parameters described herein. Figure 11 In at least one embodiment, after step 1202 , the process shown in block diagram 1200 for executing an API to obtain a resource associated with a context continues at step 1204 .

[0260] In at least one embodiment, in step 1204 of the process of executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether the API for obtaining a resource associated with the context (e.g., received in step 1202) is valid. In at least one embodiment, in step 1204, the operation of determining whether the API for obtaining a resource associated with the context is valid includes verifying parameters of the API for obtaining a resource associated with the context (e.g., Figure 11 In at least one embodiment, if, in step 1204, it is determined that the API for obtaining the resource associated with the context is valid (the "yes" branch), the process for executing the API to obtain the resource associated with the context shown in block diagram 1200 proceeds to step 1206. In at least one embodiment, if, in step 1204, it is determined that the API for obtaining the resource associated with the context is invalid (the "no" branch), the process for executing the API to obtain the resource associated with the context shown in block diagram 1200 proceeds to step 1214, as described below.

[0261] In at least one embodiment, in step 1206 of the process for executing an API to obtain resources associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine the request type of the API for obtaining resources associated with the context (e.g., received in step 1202). In at least one embodiment, the request type may be to obtain resources associated with the context. In at least one embodiment, the request type may be to obtain resources associated with a subcontext. In at least one embodiment, the request type may be to obtain resources associated with a device. In at least one embodiment, the request type is received as a parameter of the API for obtaining resources associated with the context (e.g., received in step 1202). In at least one embodiment, after step 1206, the process for executing an API for obtaining resources associated with the context shown in block diagram 1200 continues at step 1208.

[0262] In at least one embodiment, in step 1208 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether the request type (e.g., determined in step 1206) is to obtain a resource associated with the context. In at least one embodiment, in step 1208, if it is determined that the request type is to obtain a resource associated with the context ("yes" branch), the process for executing an API to obtain a resource associated with the context shown in block diagram 1200 continues at step 1216. In at least one embodiment, in step 1208, if it is determined that the request type is not to obtain a resource associated with the context ("no" branch), the process for executing an API to obtain a resource associated with the context shown in block diagram 1200 continues at step 1210.

[0263] In at least one embodiment, at step 1210 of the process for executing an API to obtain resources associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether the request type (e.g., determined at step 1206) is to obtain resources associated with a subcontext. In at least one embodiment, at step 1210, if it is determined that the request type is to obtain resources associated with a subcontext ("yes" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1222. In at least one embodiment, at step 1210, if it is determined that the request type is not to obtain resources associated with a subcontext ("no" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1212.

[0264] In at least one embodiment, at step 1212 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether the request type (e.g., determined at step 1206) is to obtain a resource associated with a device. In at least one embodiment, at step 1212, if it is determined that the request type is to obtain a resource associated with a device ("yes" branch), the process for executing an API to obtain a resource associated with a context shown in block diagram 1200 continues at step 1228. In at least one embodiment, at step 1212, if it is determined that the request type is not to obtain a resource associated with a device ("no" branch), the process for executing an API to obtain a resource associated with a context shown in block diagram 1200 continues at step 1214.

[0265] In at least one embodiment, in step 1214 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to return an error indicator (e.g., error indicator 1124, which is described herein in at least some embodiments). Figure 11 In at least one embodiment, in step 1214, an error indicator is returned to the calling process (e.g., a process executed by a processor, such as processor 102, as described herein in at least one embodiment). Figure 1 In at least one embodiment, after step 1214, the process for executing an API to obtain a resource associated with a context as shown in block diagram 1200 terminates. Figure 12 (not shown), after step 614, the process for executing the API to obtain resources associated with the context shown in block diagram 1200 continues to execute the above step 1202.

[0266] In at least one embodiment, in step 1216 of the process of executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations using the systems, methods, and operations described herein to identify a context (e.g., the context received as a parameter of the API for obtaining a resource associated with a context received at step 1202). In at least one embodiment, after step 1216, the process of executing an API to obtain a resource associated with a context shown in block diagram 1200 continues at step 1218.

[0267] In at least one embodiment, in step 1218 of the process of executing an API to obtain resources associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether a context is identified (e.g., at step 1216). In at least one embodiment, at step 1218, if it is determined that a context is identified (the "yes" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1220. In at least one embodiment, at step 1204, if it is determined that a context is not identified (the "no" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 continues at step 1214, as described above.

[0268] In at least one embodiment, at step 1220 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to store the identified context (e.g., the context identified in step 1216). In at least one embodiment, at step 1220, the processor performs one or more operations to store the identified context in a storage location received as a parameter of the API for obtaining a resource associated with the context (e.g., received in step 1202). In at least one embodiment, after step 1220, the process for executing an API to obtain a resource associated with the context shown in block diagram 1200 continues at step 1230.

[0269] In at least one embodiment, at step 1222 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, a processor executing the process performs one or more operations to identify a subcontext (e.g., the subcontext received as a parameter of the API for obtaining a resource associated with a context received at step 1202), as described herein. In at least one embodiment, after step 1222, the process for executing an API to obtain a resource associated with a context shown in block diagram 1200 continues at step 1224.

[0270] In at least one embodiment, in step 1224 of the process for executing an API to obtain resources associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to determine whether a subcontext is identified (e.g., in step 1222). In at least one embodiment, in step 1224, if it is determined that a subcontext is identified (the "yes" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 proceeds to step 1226. In at least one embodiment, in step 1224, if it is determined that a subcontext is not identified (the "no" branch), the process for executing an API to obtain resources associated with a context shown in block diagram 1200 proceeds to step 1214 described above.

[0271] In at least one embodiment, in step 1226 of the process for executing an API to obtain resources associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to store the identified subcontext (e.g., the subcontext identified in step 1222). In at least one embodiment, in step 1226, the processor performs one or more operations to store the identified subcontext in a storage location received as a parameter of the API for obtaining resources associated with the context (e.g., received in step 1202). In at least one embodiment, after step 1226, the process for executing an API to obtain resources associated with the context shown in block diagram 1200 continues at step 1230.

[0272] In at least one embodiment, in step 1228 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to store a device context (e.g., a primary context or current context of a device, such as device 202, in conjunction with at least Figure 2 1228). In at least one embodiment, at step 1228, the processor performs one or more operations to store the device context in a storage location received as an argument to an API for obtaining resources associated with the context (e.g., received at step 1202). In at least one embodiment, at step 1228, the stored device context is a primary context, a default context, or a current context for the device. In at least one embodiment, after step 1228, the process shown in block diagram 1200 for executing an API to obtain resources associated with the context continues at step 1230.

[0273] In at least one embodiment, in step 1230 of the process for executing an API to obtain a resource associated with a context shown in block diagram 1200, the processor executing the process performs one or more operations to return a success indicator (e.g., success indicator 1122, which is described herein in at least some embodiments). Figure 11 In at least one embodiment, at step 1230, a success indicator is returned to the calling process (e.g., a process executed by a processor, such as processor 102, herein at least in conjunction with Figure 1 In at least one embodiment, after step 1230, the process for executing an API to obtain a resource associated with a context as shown in block diagram 1200 terminates. Figure 12 Not shown, after step 1230 , the process for executing an API to obtain resources associated with a context shown in block diagram 1200 continues to execute step 1202 described above.

[0274] In at least one embodiment, the operations of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context are performed in an order different from the order shown in block diagram 1200. In at least one embodiment, the operations of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context are performed simultaneously or in parallel. In at least one embodiment, the operations of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context that are independent of each other (e.g., order-independent) are performed simultaneously or in parallel. In at least one embodiment, the operations of the process shown in block diagram 1200 for executing an API to obtain resources associated with a context are performed by multiple threads executing on a processor (e.g., the processor described herein).

[0275] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 11 and Figure 12 One or more circuits for executing the operations and / or instructions described herein, for example, for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) in one or more processors to be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for executing the operations and / or instructions described herein in conjunction with Figure 11 and Figure 12 One or more circuits for executing the operations and / or instructions described herein, for example, one or more circuits for executing an application programming interface (API) to instruct one or more streaming multiprocessors (SMs) of one or more graphics processing units (GPUs) to be available for executing one or more software kernels and / or otherwise performing the operations described herein. In at least one embodiment, the present invention is combined with Figure 11 and 12 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or including the Figures 1 to 25 The systems, methods, operations and / or instructions described herein are for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors that can be used to execute one or more software threads and / or otherwise perform the operations described herein. In at least one embodiment, the present invention is combined with Figure 11 and12 The components described herein are implemented in conjunction with Figures 1 to 25 One or more processes described herein for executing an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors available for executing one or more software threads and / or otherwise performing the operations described herein. In at least one embodiment, Figure 11 and 12 Not shown in this paper Figure 11 and 12 One or more components described herein include Figures 26 to 58 One or more components described are configured to implement an application programming interface (API) to indicate one or more streaming multiprocessors (SMs) in one or more processors that can be used to execute one or more software threads and / or otherwise perform the operations described herein.

[0276] Figure 13 1300 is a block diagram illustrating an application programming interface (API) for segmenting context resources according to at least one embodiment. In at least one embodiment, one or more circuits of a processor are configured to execute the segmented context resource API 1302 to segment context resources according to one or more received segmentation criteria. In at least one embodiment, Figure 13 1 , one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a segmented context resource API 1302, thereby executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used. In at least one embodiment, Figure 13 1302, one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a segmented context resource API 1302, thereby executing an application programming interface (API) to generate an indication of one or more streaming multiprocessor (SM) subsets in the processor that are available for executing one or more software kernels. In at least one embodiment, Figure 13 , one or more circuits of a processor (e.g., a processor described herein) execute one or more instructions to execute a segmented context resource API 1302, thereby executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in the one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used in response to receiving a second API (e.g., an API described herein).

[0277] In at least one embodiment, the segment context resource API 1302, when called, receives one or more arguments indicating information about an operation to be performed using techniques such as those described herein. In at least one embodiment, the segment context resource API 1302, when called, receives one or more arguments indicating information about an instruction to be performed using techniques such as those described herein.

[0278] In at least one embodiment, the segment context resource API 1302 receives as input one or more parameters including a segment resource list return 1304. In at least one embodiment, the segment resource list return 1304 is a data value that includes information that can be used to identify, indicate, or otherwise specify a storage location where a segment resource list is to be stored as a result of using the segment context resource API 1302. In at least one embodiment, the segment resource list return identified, indicated, or otherwise specified by the segment resource list return 1304 is one of a plurality of parameters that can be used by the segment context resource API 1302 to segment context resources. In at least one embodiment, the segment resource list return 1304 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the segment context resource API 1302) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0279] In at least one embodiment, the segmented context resources API 1302 receives as input one or more parameters including a remaining resource list return 1306. In at least one embodiment, the remaining resource list return 1306 is a data value that includes information that can be used to store a location where a remaining resource list (e.g., a list of resources remaining after segmentation) is stored as a result of using the segmented context resources API 1302. In at least one embodiment, the remaining resource list return identified, indicated, or otherwise specified by the remaining resource list return 1306 is one of a plurality of parameters that can be used by the segmented context resources API 1302 to segment context resources. In at least one embodiment, the remaining resource list return 1306 is a data value used to identify, indicate, or otherwise specify to an API (e.g., the segmented context resources API 1302) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0280] In at least one embodiment, the segmented context resources API 1302 receives as input one or more parameters including an input resource list 1308. In at least one embodiment, the input resource list 1308 is a data value that includes information that can be used to identify, indicate, or otherwise specify a list of resources to be segmented using the segmented context resources API 1302. In at least one embodiment, the input resource list identified, indicated, or otherwise specified by the input resource list 1308 is one of a plurality of parameters that can be used by the segmented context resources API 1302 to segment context resources. In at least one embodiment, the input resource list 1308 is a data value that is used to identify, indicate, or otherwise specify to an API (e.g., the segmented context resources API 1302) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0281] In at least one embodiment, the segmented context resource API 1302 receives as input one or more parameters including flags 1310. In at least one embodiment, flags 1310 are data values that include information that can be used to identify, indicate, or otherwise specify one or more flags that indicate at least which resources (e.g., SMs) in the input resource list 1308 are to be assigned to the segmented resource list (e.g., stored in the segmented resource list return 1304) and which resources are remaining (e.g., stored in the remaining resource list return 1306) when using the segmented context resource API 1302. In at least one embodiment, the flag identified, indicated, or otherwise specified by flags 1310 is one of a plurality of parameters that can be used by the segmented context resource API 1302 to segment context resources. In at least one embodiment, flags 1310 are data values that are used to identify, indicate, or otherwise specify to an API (e.g., the segmented context resource API 1302) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0282] In at least one embodiment, the segment context resource API 1302 receives as input one or more parameters including a minimum count 1312. In at least one embodiment, the minimum count 1312 is a data value that includes information that can be used to identify, indicate, or otherwise specify a minimum count of segment resources (e.g., a minimum number or minimum count of segment resources in a segment resource list) obtained using the segment context resource API 1302. In at least one embodiment, the minimum count identified, indicated, or otherwise specified by the minimum count 1312 is one of a plurality of parameters that can be used by the segment context resource API 1302 to segment context resources. In at least one embodiment, the minimum count 1312 is a data value used to identify, indicate, or otherwise specify to an API (e.g., the segment context resource API 1302) a set of operations or instructions to be executed by one or more PPUs (e.g., GPUs) and / or one or more accelerators within a heterogeneous processor, as described herein.

[0283] In at least one embodiment, the segment context resource API 1302 receives as input one or more parameters including one or more other parameters 1318. In at least one embodiment, the other parameters 1318 are data including information indicating any other information available when executing the segment context resource API 1302 to segment the context resource.

[0284] In at least one embodiment, Figure 13 130, the processor executes one or more instructions to execute one or more APIs, such as a segmented context resource API 1302, thereby executing an application programming interface (API) so that a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors are stored according to a plurality of groups in which the corresponding plurality of SMs are to be used using one or more parameters, including but not limited to a segmented resource list return 1304, a remaining resource list return 1306, an input resource list 1308, a flag 1310, a minimum count 1312, and / or other parameters 1318. In at least one embodiment, Figure 13 Not shown, the processor executes one or more instructions to execute one or more APIs, such as a subdivision context resource API 1302, thereby executing an application programming interface (API) to generate an indication of one or more streaming multiprocessor (SM) subsets in the processor that are available for executing one or more software kernels using one or more parameters, including but not limited to a subdivision resource list return 1304, a remaining resource list return 1306, an input resource list 1308, a flag 1310, a minimum count 1312 and / or other parameters 1318.

[0285] In at least one embodiment, the segment context resource API 1302, if called, causes one or more APIs (e.g., Figure 3 In at least one embodiment, the segmented context resource API 1302, if called, causes one or more APIs (e.g., one or more APIs 306) to be added, inserted, or otherwise included in a flow or instruction set to be executed by one or more accelerators within a heterogeneous processor. Figure 3 The method further includes adding one or more operations or instructions to a parallel computing environment 308 described herein that are to be added, inserted, or otherwise included in a flow or instruction set executed by one or more accelerators within a heterogeneous processor.

[0286] In at least one embodiment, in response to the segment context resource API 1302, one or more APIs 306 (if executed) will cause one or more processors to execute the segment context resource API return 1320. In at least one embodiment, the segment context resource API return 1320 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the segment context resource API 1302. In at least one embodiment, the segment context resource API return 1320 indicates a success indicator 1322. In at least one embodiment, the success indicator 1322 is data including any value indicating the success of the segment context resource API 1302. In at least one embodiment, the success indicator 1322 includes information indicating one or more specific types of success generated as a result of executing the segment context resource API 1302. In at least one embodiment, the success indicator 1322 includes information indicating one or more other data values generated as a result of executing the segment context resource API 1302.

[0287] In at least one embodiment, the segment context resource API return 1320 indicates an error indicator 1324. In at least one embodiment, the error indicator 1324 is data including any value indicating a failure of the segment context resource API 1302. In at least one embodiment, the error indicator 1324 includes information indicating one or more specific types of errors generated as a result of executing the segment context resource API 1302. In at least one embodiment, the error indicator 1324 includes information indicating one or more other data values generated as a result of executing the segment context resource API 1302.

[0288] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to the subdivision context resource API 1302) adds various operations of various types to a stream to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes an acquire semaphore operation. In at least one embodiment, the stream operation includes a release semaphore operation. In at least one embodiment, the stream operation includes one or more operations for flushing and / or invalidating cache memory, such as the L2 cache memory of a PPU (e.g., a GPU) and / or the cache memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation includes one or more operations for indicating submission of an operation to an external device (e.g., one or more accelerators within a heterogeneous processor). In at least one embodiment, the one or more operations for indicating submission of an operation to an external device use software code, such as at least in conjunction with the present invention. Figure 5 Example software code indicating stream operations is described.

[0289] In at least one embodiment, a parallel computing environment 308 including one or more APIs 306 (including but not limited to the segmented context resource API 1302) includes one or more function signatures that can be used to indicate one or more callback functions for operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more operations cause the one or more callback functions to be executed. In at least one embodiment, the one or more operations that cause the one or more callback functions to be executed use software code, such as example software code indicating a function signature of the callback function, as described herein in at least conjunction with Figure 5 As stated.

[0290] In at least one embodiment, to specify one or more accelerators within a heterogeneous processor to perform one or more operations indicated by the segmented context resource API 1302 to the one or more APIs 306, one or more data structures of the one or more APIs 306 may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations. In at least one embodiment, the one or more data structures of the one or more APIs 306 that may be used to specify one or more external devices to which the one or more APIs 306 are to submit the one or more operations utilize software code, such as example software code indicating a data structure representing a device node of one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 As stated.

[0291] In at least one embodiment, to specify the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor, one or more data structures of one or more APIs 306 are used. In at least one embodiment, one or more data structures of one or more APIs 306 for specifying the type and data of one or more operations indicated by one or more operations to be performed by one or more accelerators within a heterogeneous processor use software code, such as example software code indicating data structures for specifying the type and data of one or more operations to be performed by one or more accelerators within a heterogeneous processor, as described herein in at least conjunction with Figure 5 As stated.

[0292] In at least one embodiment, one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be added to a stream or other instruction set to be executed by one or more accelerators within the heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set are executed in response to the segment context resource API 1302, as described above. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set, which are executed in response to the segment context resource API 1302, use software code, such as example software code that directs stream operation API calls in the parallel computing environment 308, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0293] In at least one embodiment, the one or more APIs 306 include instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators within the heterogeneous processor to be added to one or more executable graphs. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators in the heterogeneous processor to be added to one or more executable graphs are similar to the manner in which one or more operations or instructions to be performed by one or more accelerators in the heterogeneous processor are added to one or more streams or instruction sets in response to the segmented context resource API 1302, as described herein. In at least one embodiment, the instructions that, if executed, cause one or more operations or instructions to be performed by one or more accelerators in the heterogeneous processor to be added to one or more executable graphs use software code, such as example software code that instructs one or more APIs 306 of the parallel computing environment 308 to add one or more operations or instructions to one or more executable graphs, as described herein in conjunction with at least one embodiment. Figure 5 As stated.

[0294] Figure 14 1400 is a block diagram illustrating a process of executing an application programming interface (API) to subdivide context resources according to at least one embodiment. In at least one embodiment, the process of executing an API to subdivide context resources shown in block diagram 1400 is a block diagram illustrating a process of executing an application programming interface (API) to subdivide context resources according to at least one embodiment. Figure 13 In at least one embodiment, some or all of the process of executing the API to segment context resources shown in block diagram 1400 (or any other process described herein, or variations and / or combinations thereof) is implemented on one or more computer systems, servers, processors, integrated circuits, and / or a combination thereof. Figures 26 to 58 The present invention relates to a computer program that is executed under the control of such devices described herein, which are configured with computer executable instructions and implemented as code (e.g., computer executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes multiple computer-readable instructions that can be executed by one or more processors (e.g., the processors described herein). In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, the processor (e.g., at least one processor described herein) is a computer program that is executed by one or more processors (e.g., the processors described herein). Figure 1 The processor 110 described herein performs one or more steps of the process for executing an API to subdivide context resources shown in block diagram 1400. In at least one embodiment, one or more other processors (e.g., the processors described herein) perform one or more steps of the process for executing an API to subdivide context resources shown in block diagram 1400.

[0295] In at least one embodiment, at step 1402 of the process for executing an API to subdivide context resources shown in block diagram 1400, a processor executing the process performs one or more operations to receive or otherwise obtain an API for subdividing context resources. In at least one embodiment, at step 1402, the API is provided to the processor in other ways (e.g., as described herein in at least one embodiment). Figure 1 In at least one embodiment, at step 1402, by the processor (e.g., at least in conjunction with the present invention) in other ways, Figure 1 In at least one embodiment, in step 1402, the API for subdividing context resources includes one or more parameters, such as at least Figure 13In at least one embodiment, after step 1402 , the process shown in block diagram 1400 for executing an API to subdivide context resources continues at step 1404 .

[0296] In at least one embodiment, in step 1404 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to determine whether the API for subdividing context resources (e.g., received in step 1402) is valid. In at least one embodiment, in step 1404, the operation for determining whether the API for subdividing context resources is valid includes validating parameters of the API for subdividing context resources (e.g., at least in combination with Figure 13 In at least one embodiment, at step 1404, if it is determined that the API for segmenting the context resource is valid (the "yes" branch), the process for executing the API to segment the context resource shown in block diagram 1400 proceeds to step 1406. In at least one embodiment, at step 1404, if it is determined that the API for segmenting the context resource is invalid (the "no" branch), the process for executing the API to segment the context resource shown in block diagram 1400 proceeds to step 1418 described below.

[0297] In at least one embodiment, at step 1406 of the process for executing an API to refine context resources shown in block diagram 1400, a processor executing the process performs one or more operations to identify resources in the input resource list (e.g., received as parameters of the API for refine context resources received at step 1402), as described herein. In at least one embodiment, after step 1406, the process for executing an API to refine context resources shown in block diagram 1400 continues at step 1408.

[0298] In at least one embodiment, in step 1408 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to subdivide input resources (e.g., identified in step 1406) using flags and / or minimum counts (e.g., received as parameters of the API for subdividing context resources received at step 1402). In at least one embodiment, in step 1408, the one or more operations for subdividing the input resources are performed using an indication of available resources. In at least one embodiment, in step 1408, the one or more operations for subdividing the input resources are performed to divide the resources into two or more resource lists, as described herein. In at least one embodiment, in step 1408, if no input resource list is identified in step 1406, step 1408 will not subdivide the input resources. In at least one embodiment, after step 1408, the process for executing an API to subdivide context resources shown in block diagram 1400 continues at step 1410.

[0299] In at least one embodiment, in step 1410 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to determine whether the input resource list is subdivided (e.g., at step 1408). In at least one embodiment, in step 1410, if it is determined that the input resource list is subdivided (the "yes" branch), the process for executing an API to subdivide context resources shown in block diagram 1400 continues at step 1412. In at least one embodiment, in step 1410, if it is determined that the input resource list is not subdivided (the "no" branch), the process for executing an API to subdivide context resources shown in block diagram 1400 continues at step 1418, as described below.

[0300] In at least one embodiment, in step 1412 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to store the subdivided input resource list (e.g., subdivided in step 1408) in the storage location received as a parameter of the API for subdividing context resources received in step 1402. In at least one embodiment, in step 1412, the subdivided input resource list is an empty list (e.g., does not contain any resources). In at least one embodiment, after step 1412, the process for executing an API to subdivide context resources shown in block diagram 1400 continues at step 1414.

[0301] In at least one embodiment, in step 1414 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to store the remaining input resource list (e.g., the resources remaining after subdividing the input resource list at step 1408) in the storage location received as a parameter of the API for subdividing context resources received in step 1402. In at least one embodiment, at step 1412, the remaining input resource list is an empty list (e.g., does not contain any resources). In at least one embodiment, after step 1412, the process for executing an API to subdivide context resources shown in block diagram 1400 continues at step 1414.

[0302] In at least one embodiment, in step 1416 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to return a success indicator (e.g., success indicator 1322, which is described herein in conjunction with at least one embodiment). Figure 13 In at least one embodiment, at step 1416, a success indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, as described herein in at least one embodiment). Figure 1 In at least one embodiment, after step 1416, the process for executing an API to subdivide context resources shown in block diagram 1400 terminates. Figure 14 Not shown, after step 1416, the process shown in block diagram 1400 for executing an API to subdivide context resources continues at step 1402 described above.

[0303] In at least one embodiment, in step 1418 of the process for executing an API to subdivide context resources shown in block diagram 1400, the processor executing the process performs one or more operations to return an error indicator (e.g., error indicator 1324, which is described herein in conjunction with at least one embodiment). Figure 13 In at least one embodiment, in step 1418, an error indicator is returned to the calling process (e.g., a process executed by a processor such as processor 102, as described herein in at least some embodiments). Figure 1 In at least one embodiment, after step 1418, the process for executing an API to subdivide context resources shown in block diagram 1400 terminates. In at least one embodiment ( Figure 14 ), after step 1418, the process for executing the API to subdivide the context resources shown in block diagram 1400 continues to execute the above step 1402.

[0304] In at least one embodiment, the operations of the process for executing the API to subdivide context resources shown in block diagram 1400 are performed in an order different from the order shown in block diagram 1400. In at least one embodiment, the operations of the process for executing the API to subdivide context resources shown in block diagram 1400 are performed simultaneously or in parallel. In at least one embodiment, the operations of the process for executing the API to subdivide context resources shown in block diagram 1400 that are independent of each other (e.g., order-independent) are performed simultaneously or in parallel. In at least one embodiment, the operations of the process for executing the API to subdivide context resources shown in block diagram 1400 are performed by multiple threads executing on a processor (e.g., the processor described herein).

[0305] In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include a processor for executing the operations described herein in conjunction with Figure 13 and Figure 14 One or more circuits for performing the operations and / or instructions described herein, such as one or more circuits for executing an application programming interface (API) to cause a plurality of identifiers of a corresponding plurality of streaming multiprocessors (SMs) in one or more processors to be stored according to a plurality of groups in which the corresponding plurality of SMs are to be used and / or otherwise perform the operations described herein. In at least one embodiment, one or more processors (e.g., processor 102, processor 110, and / or other processors and / or accelerators, such as those described herein) include circuits for performing the operations and / or instructions described herein in conjunction with Figure 13 and Figure 14 One or more circuits that execute the operations and / or instructions described herein, for example, one or more circuits that execute an application programming interface (API) to generate instructions for one or more streaming multiprocessor (SM) subsets in a processor that can be used to execute one or more software kernels and / or otherwise perform the operations described herein. In at least one embodiment, this document is combined with Figure 13 and 14 The operations and / or instructions described herein are included in conjunction with Figures 1 to 25 The systems, methods, operations and / or instructions described herein and / or including the Figures ...

Claims

1. A processor, comprising: one or more circuits for executing an application programming interface (API) to deallocate one or more data structures for indicating which of one or more streaming multi-processors (SMs) in one or more processors will be used to execute one or more software threads.

2. The processor according to claim 1, wherein the API is for receiving a parameter including a sub-context indicating the one or more data structures to be deallocated.

3. The processor according to claim 1, wherein the API is for indicating that a sub-context including the one or more data structures has been deallocated.

4. The processor according to claim 1, wherein the one or more data structures indicate that a plurality of SMs of the one or more devices are at least divided into the one or more SMs and a second one or more SMs not including the first one or more SMs.

5. The processor according to claim 1, wherein the API is for deallocating the one or more data structures allocated by a second API for allocating one or more data structures.

6. The processor according to claim 1, wherein the one or more data structures are included in a context of the one or more processors.

7. The processor according to claim 1, wherein the one or more data structures indicate a sub-context of a context of the one or more processors.

8. A computer-implemented method, comprising: executing an application programming interface (API) to deallocate one or more data structures for indicating which of one or more streaming multi-processors (SMs) in one or more processors will be used to execute one or more software threads.

9. The computer-implemented method according to claim 8, wherein performing the API comprises: receiving a parameter including a sub-context indicating the one or more data structures to be deallocated.

10. The computer-implemented method according to claim 8, wherein performing the API comprises: indicating that a sub-context including the one or more data structures has been deallocated.

11. The computer-implemented method according to claim 8, wherein the one or more data structures indicate that a plurality of SMs of the one or more devices are at least divided into the one or more SMs and a second one or more SMs not including the first one or more SMs.

12. The computer-implemented method according to claim 8, wherein performing the API comprises: deallocating the one or more data structures allocated by a second API for allocating one or more data structures.

13. The computer-implemented method according to claim 8, wherein the one or more data structures are included in a context of the one or more processors.

14. The computer-implemented method according to claim 8, wherein the one or more data structures are included in a sub-context of a context of the one or more processors.

15. A computer system, comprising: One or more processors and a memory, the memory storing executable instructions that, when executed by the one or more processors, execute an application programming interface (API) to deallocate one or more data structures for indicating which of one or more streaming multi-processors (SMs) in the one or more processors will be used to execute one or more software threads.

16. The computer system according to claim 15, wherein the API is for receiving a parameter including a sub-context indicating the one or more data structures to be deallocated.

17. The computer system according to claim 15, wherein the API is for indicating that a sub-context including the one or more data structures has been deallocated.

18. The computer system according to claim 15, wherein the one or more data structures indicate that a plurality of SMs of the one or more devices are at least divided into the one or more SMs and a second one or more SMs that do not include the first one or more SMs.

19. The computer system according to claim 15, wherein the API is for deallocating the one or more data structures allocated by a second API for allocating one or more data structures.

20. The computer system according to claim 15, wherein the one or more data structures are included in a sub-context of a context of the one or more processors.