Application programming interface for causing execution of accelerator operations

By providing an application programming interface (API), the complexity of PPU and accelerator interfaces in heterogeneous processors is solved, and more efficient resource management and system efficiency are achieved.

CN120283224APending Publication Date: 2025-07-08NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082289.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-28
Filing Date
2023-11-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When developing accelerators that utilize heterogeneous processors, software programmers need to use expensive technologies to interface PPUs with accelerators, resulting in increased development complexity and cost.

Method used

A set of application programming interfaces (APIs) is provided to simplify the operation between the PPU and the accelerator in the heterogeneous processor, including the stream operation API, the memory operation API, the error information management API, etc., to achieve more efficient communication and resource management.

Benefits of technology

By simplifying the interface between the PPU and the accelerator, development complexity and cost are reduced, and the efficiency and scalability of heterogeneous processor systems are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283224A_ABST
    Figure CN120283224A_ABST
Patent Text Reader

Abstract

Apparatus, systems, and techniques for executing one or more application programming interfaces (APIs) to perform one or more operations for one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more processors are to execute one or more instructions in response to one or more APIs to indicate one or more operations in a sequence of operations to be performed by one or more accelerators within the heterogeneous processor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] Priority claim

[0003] This application claims the benefit of priority of U.S. Patent Application No. 18 / 070,084, filed on November 28, 2022, entitled "APPLICATION PROGRAMMING INTERFACE TO CAUSE PERFORMANCE OF ACCELERATOR OPERATIONS", and this application is a continuation application of the above application for the United States, the entire content of which is incorporated herein by reference. Technical Field

[0004] At least one embodiment relates to processing resources for executing one or more application programming interfaces (APIs) to perform one or more operations for one or more accelerators within a heterogeneous processor. For example, at least one embodiment relates to a processor or computer system for executing one or more application programming interfaces that enable the performance of various accelerator functions described herein. Background Art

[0005] Parallel computing environments, such as Compute Unified Device Architecture (CUDA), allow software programmers to develop software programs that run in whole or in part on one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)). Software programmers are increasingly leveraging accelerators in heterogeneous processors to further improve performance. When software programmers develop software programs that combine PPUs with accelerators in heterogeneous processors (e.g., deep learning accelerators (DLAs)), various expensive techniques are required to interface these PPUs with these accelerators. Brief Description of the Drawings

[0006] Figure 1 is a block diagram showing a software program to be executed by a processor (e.g., a central processing unit (CPU) and a graphics processing unit (GPU)) and an accelerator in a heterogeneous processor according to at least one embodiment;

[0007] Figure 2 shows an application programming interface (API) for indicating one or more operations in an operation sequence to be performed by one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0008] Figure 3Shows an API for indicating memory to be transferred between one or more parallel processing units (PPUs) (e.g., GPUs) and one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0009] Figure 4 Shows an API for indicating one or more memory regions available for storing error information generated by one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0010] Figure 5 Shows an API for indicating one or more memory regions that are no longer available for storing error information generated by one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0011] Figure 6 Shows an API for indicating one or more instruction sequences to be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0012] Figure 7 Shows an API for indicating one or more instruction sequences that are no longer to be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor according to at least one embodiment;

[0013] Figure 8 Shows a process for executing one or more APIs of one or more accelerators within a heterogeneous processor by a parallel computing environment according to at least one embodiment;

[0014] Fig. 9 Shows an exemplary data center according to at least one embodiment;

[0015] Fig.10 Shows a processing system according to at least one embodiment;

[0016] Fig.11 Shows a computer system according to at least one embodiment;

[0017] Fig.12 Shows a system according to at least one embodiment;

[0018] Fig.13 Shows an exemplary integrated circuit according to at least one embodiment;

[0019] Fig.14 Shows a computing system according to at least one embodiment;

[0020] Fig.15Shows an APU according to at least one embodiment;

[0021] Fig.16 Shows a CPU according to at least one embodiment;

[0022] Fig.17 Shows an exemplary accelerator integration slice according to at least one embodiment;

[0023] Figure 18A-18B Shows an exemplary graphics processor according to at least one embodiment;

[0024] Fig.19A Shows a graphics core according to at least one embodiment;

[0025] Fig.19B Shows a GPGPU according to at least one embodiment;

[0026] Fig. 20A Shows a parallel processor according to at least one embodiment;

[0027] Fig. 20B Shows a processing cluster according to at least one embodiment;

[0028] Fig. 20C Shows a graphics multiprocessor according to at least one embodiment;

[0029] Fig.21 Shows a graphics processor according to at least one embodiment;

[0030] Fig. 22 Shows a processor according to at least one embodiment;

[0031] Fig.23 Shows a processor according to at least one embodiment;

[0032] Fig.24 Shows a graphics processor core according to at least one embodiment;

[0033] Fig.25 Shows a PPU according to at least one embodiment;

[0034] Fig.26 Shows a GPC according to at least one embodiment;

[0035] Fig. 27 Shows a streaming multiprocessor according to at least one embodiment;

[0036] Fig.28 Shows the software stack of a programming platform according to at least one embodiment;

[0037] Fig.29 Shows according to at least one embodiment of Fig.28 CUDA implementation of the software stack;

[0038] Fig.30 illustrates according to at least one embodiment Fig.28 ROCm implementation of the software stack;

[0039] Fig.31 illustrates according to at least one embodiment Fig.28 OpenCL implementation of the software stack;

[0040] Fig.32 illustrates software supported by a programming platform according to at least one embodiment;

[0041] Fig.33 illustrates according to at least one embodiment in Figure 28-Figure 31 compiled code executed on the programming platform;

[0042] Fig.34 illustrates according to at least one embodiment in Figure 28-Figure 31 more detailed compiled code executed on the programming platform;

[0043] Fig.35 illustrates converting source code before compiling the source code according to at least one embodiment;

[0044] Fig.36A illustrates a system configured to compile and execute CUDA source code using different types of processing units according to at least one embodiment;

[0045] Fig.36B illustrates according to at least one embodiment a system configured to compile and execute Fig.36A CUDA source code using a CPU and a CUDA-enabled GPU;

[0046] Fig.36C illustrates according to at least one embodiment a system configured to compile and execute Fig.36A CUDA source code using a CPU and a non-CUDA-enabled GPU;

[0047] Fig.37 illustrates according to at least one embodiment by Fig.36C exemplary kernels converted by the CUDA to HIP conversion tool;

[0048] Fig.38 more particularly illustrates according to at least one embodiment Fig.36C a non-CUDA-enabled GPU; and

[0049] Fig.39shows how threads of an exemplary CUDA grid according to at least one embodiment are mapped to Fig.38 different computing units; and

[0050] Fig.40 shows how existing CUDA code can be migrated to data parallel C++ code according to at least one embodiment. DETAILED DESCRIPTION

[0051] Figure 1 is a block diagram showing a software program 104 to be executed by a processor (such as a central processing unit (CPU) 102 and a graphics processing unit (GPU) 110) and an accelerator 114 within a heterogeneous processor according to at least one embodiment. In at least one embodiment, the CPU 102 is any processor having any architecture further described herein. In at least one embodiment, the CPU 102 is any general-purpose processor having any architecture further described herein. In at least one embodiment, the processor (such as the CPU 102) includes circuitry for performing one or more computing operations. In at least one embodiment, the processor (such as the CPU 102) includes any circuit configuration for performing one or more computing operations further described herein.

[0052] In at least one embodiment, the processor (such as the central processing unit (CPU) 102) executes a parallel computing environment 106. In at least one embodiment, the processor (such as the CPU 102) is. In at least one embodiment, the processor (such as the CPU) executes a parallel computing environment 106, such as Compute Unified Device Architecture (CUDA). In at least one embodiment, the parallel computing environment 106 is instructions that, if executed by one or more processors (such as the CPU 102), can facilitate one or more CPUs 102, one or more parallel processing units (PPUs) (such as the GPU 110), and / or one or more accelerators 114 within a heterogeneous processor to execute one or more software programs.

[0053] In at least one embodiment, one or more PPUs are processors that include one or more circuits for performing parallel computing operations, such as GPU 110 and any other parallel processors further described herein. In at least one embodiment, GPU 110 is hardware that includes circuits for performing one or more computing operations, as further described below in connection with various embodiments. In at least one embodiment, GPU 110 includes one or more processing cores, each for performing one or more computing operations. In at least one embodiment, GPU 110 includes one or more processing cores for performing one or more parallel computing operations. In at least one embodiment, GPU 110 is packaged with CPU 102 or other processors as a system-on-chip (SoC). In at least one embodiment, GPU 110 is packaged with CPU 102 or other processors on a shared die or other substrate as a system-on-chip (SoC).

[0054] In at least one embodiment, one or more accelerators 114 within a heterogeneous processor are hardware that includes one or more circuits for performing specific computing operations, such as a deep learning accelerator (DLA), a programmable vision accelerator (PVA), a field-programmable gate array (FPGA), or any other accelerator further described herein. In at least one embodiment, one or more accelerators 114 within a heterogeneous processor are one or more accelerators 114 within a heterogeneous processor. In at least one embodiment, one or more accelerators 114 within a heterogeneous processor are one or more systems that include one or more processors. In at least one embodiment, one or more accelerators 114 within a heterogeneous processor are one or more systems that include a GPU and other accelerators (such as those accelerators described above and / or any other accelerators further described herein).

[0055] In at least one embodiment, accelerator 114 within a heterogeneous processor is packaged with CPU 102 or other processors as a system-on-chip (SoC). In at least one embodiment, accelerator 114 within a heterogeneous processor is packaged with CPU 102 or other processors on a shared die or other substrate as a system-on-chip (SoC). In at least one embodiment, one or more CPUs 102, one or more GPUs 110 or other PPUs, and / or accelerator 114 within a heterogeneous processor are packaged as a system-on-chip (SoC). In at least one embodiment, one or more CPUs 102, one or more GPUs 110 or other PPUs, and / or accelerator 114 within a heterogeneous processor are packaged on a shared die or other substrate as a system-on-chip (SoC).

[0056] In at least one embodiment, the parallel computing environment 106 (e.g., CUDA) includes libraries and other software programs for performing one or more computing operations using one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor. In at least one embodiment, the parallel computing environment 106 includes libraries and other software programs that, if executed by one or more processors (e.g., one or more CPUs 102), cause one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, the parallel computing environment 106 includes libraries that, if executed, cause one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor to perform mathematical operations. In at least one embodiment, the parallel computing environment 106 includes libraries that, if executed, cause one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor to perform any other operations further described herein.

[0057] In at least one embodiment, one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor perform one or more computing operations in response to one or more application programming interfaces (APIs). In at least one embodiment, an API is a set of software instructions that, if executed by one or more processors (e.g., CPU 102), cause one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, the parallel computing environment 106 includes one or more APIs 108 that, if executed by one or more processors (e.g., CPU 102), cause one or more PPUs (e.g., GPU 110) and / or one or more accelerators 114 within a heterogeneous processor to perform one or more computing operations. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c) that, if executed, cause one or more processors (e.g., CPU 102) to perform one or more operations such as computing operations, error reporting, scheduling other operations to be performed by GPU 110 and / or accelerators 114 within a heterogeneous processor, or any other operations further described herein. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c) that, if executed, cause one or more PPUs (e.g., GPU 110) to perform one or more operations such as computing operations, error reporting, or any other operations further described herein. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c), such as those described below in connection with Figures 2 to 7Those described, if these functions or APIs are executed, cause one or more accelerators 114 within the heterogeneous processor to perform one or more operations, such as computing operations, error reporting, or any other operations further described herein. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c) for causing the CPU 102 to perform one or more computing operations in response to information or events generated by one or more PPUs (such as GPU 110) and / or one or more accelerators 114 within the heterogeneous processor. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c) that, when called, cause the CPU 102 to perform one or more computing operations in response to information or events generated by one or more PPUs (such as GPU 110) and / or one or more accelerators 114 within the heterogeneous processor. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c), such as socket API 108(a), as described below in connection with Figure 2 Those described. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c), such as memory API 108(b), as described below in connection with Figure 3 Those described. In at least one embodiment, one or more APIs 108 include one or more functions or APIs 108(a), 108(b), 108(c), such as one or more error APIs 108(c), as described below in connection with Figures 4 to 7 Those described.

[0058] In at least one embodiment, a processor (e.g., CPU 102) executes one or more software programs 104. In at least one embodiment, the one or more software programs are an instruction set that, if executed, causes one or more processors (e.g., CPU 102, PPU (e.g., GPU 110)) and / or an accelerator 114 in a heterogeneous processor to perform computational operations. In at least one embodiment, the software program 104 includes instructions and / or operations to be executed by one or more PPUs (e.g., GPU 110). In at least one embodiment, the one or more software programs 104 include GPU-specific code 112 and / or accelerator-specific code 116. In at least one embodiment, the instructions and / or operations to be executed by one or more PPUs (e.g., GPU 110) are PPU-specific or GPU-specific code 112. In at least one embodiment, the GPU-specific code 110 is a set of software instructions and / or other operations (as further described herein) to be executed by one or more GPUs 110. In at least one embodiment, the software program 104 includes instructions and / or operations to be executed by one or more accelerators 114 in a heterogeneous processor. In at least one embodiment, the instructions and / or operations to be executed by one or more accelerators 114 in a heterogeneous processor are accelerator-specific code 116. In at least one embodiment, the accelerator-specific code 116 is a set of software instructions and / or other operations to be executed by one or more accelerators 116, as further described herein. In at least one embodiment, the PPU-specific or GPU-specific code 112 and / or the accelerator-specific code 116 will be executed in response to one or more APIs 108, as described below in connection with Figures 2 to 7 described.

[0059] Figure 2An application programming interface (API) 202 is shown that indicates one or more operations 206 in a stream 204 or sequence of operations to be performed by one or more accelerators within a heterogeneous processor, according to at least one embodiment. In at least one embodiment, the API 202 is a set of instructions that, if executed, cause one or more processors to perform one or more functions in response to one or more API calls 202. In at least one embodiment, an API call 202 is a set of instructions that, if executed, cause one or more processors to execute the API 202. In at least one embodiment, an API call 202 is a function call. In at least one embodiment, an API call 202 is a software function to be called by one or more software programs. In at least one embodiment, in response to an API call 202, one or more processors are used to execute a set of instructions and then return 212. In at least one embodiment, the return 212 is a change in the control flow from the API 202 to the software program after calling the API 202. In at least one embodiment, the return 212 causes one or more data values 214, 216 to be transferred to a memory accessible by one or more software programs.

[0060] In at least one embodiment, the API call is a stream operation API call 202. In at least one embodiment, the stream operation API call 202 is a set of software instructions that, if executed by one or more processors, cause one or more operations in an operation stream to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation API call 202 is a set of software instructions that, if executed by one or more processors, cause one or more other instructions to be added to a set of other instructions to be executed, where the one or more other instructions are to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operation API call 202 is a set of software instructions that, if executed by one or more processors, cause one or more other instructions to be added to a set of other instructions to be executed by a parallel processing unit (PPU) (such as a graphics processing unit (GPU)), where the one or more other instructions are to be performed by one or more accelerators within a heterogeneous processor.

[0061] In at least one embodiment, if the stream operation API call 202 is invoked by one or more software programs, the API is caused to direct one or more accelerators within the heterogeneous processor to execute one or more operations. In at least one embodiment, if the stream operation API call 202 is invoked by one or more software programs, the API is caused to direct one or more accelerators within the heterogeneous processor to execute one or more operations in a stream. In at least one embodiment, if the stream operation API call 202 is invoked by one or more software programs, the API is caused to direct one or more accelerators within the heterogeneous processor to execute one or more instructions. In at least one embodiment, if the stream operation API call 202 is invoked by one or more software programs, the API is caused to direct one or more accelerators within the heterogeneous processor to execute one or more instructions out of a set of instructions to be executed by one or more PPUs (e.g., GPUs).

[0062] In at least one embodiment, if the stream operation API call 202 is invoked, the API is caused to enqueue one or more operations into a stream for execution partially by one or more accelerators within the heterogeneous processor and partially by one or more GPUs. In at least one embodiment, if the stream operation API call 202 is invoked, the API is caused to enqueue one or more instructions into a set of instructions for execution partially by one or more accelerators within the heterogeneous processor and partially by one or more GPUs.

[0063] In at least one embodiment, the stream operation API call 202 is used to cause one or more circuits in a processor to execute the API to direct one or more accelerators in the heterogeneous processor to execute one or more instructions. In at least one embodiment, one or more circuits of the processor are used to execute the API in response to the stream operation API call 202 to direct one or more accelerators within the heterogeneous processor to execute one or more instructions. In at least one embodiment, the stream operation API call 202 is used to cause one or more circuits in a processor to execute the API to direct one or more accelerators within the heterogeneous processor to execute one or more operations in a stream. In at least one embodiment, one or more circuits of the processor are used to execute the API in response to the stream operation API call 202 to direct one or more accelerators within the heterogeneous processor to execute one or more stream operations and / or portions of a stream.

[0064] In at least one embodiment, the stream operation API call 202 is used to cause one or more processors in the system to execute an API to direct one or more accelerators within the heterogeneous processor to execute one or more instructions. In at least one embodiment, one or more processors in the system are used to execute an API in response to the stream operation API call 202 to direct one or more accelerators within the heterogeneous processor to execute one or more instructions. In at least one embodiment, the stream operation API call 202 is used to cause one or more processors in the system to execute an API to direct one or more accelerators within the heterogeneous processor to execute one or more operations in a stream. In at least one embodiment, one or more processors in the system are used to execute an API in response to the stream operation API call 202 to direct one or more accelerators within the heterogeneous processor to execute one or more stream operations and / or respective portions of a stream.

[0065] In at least one embodiment, the stream operation API call 202, when called, receives one or more parameters 204, 206, 208, 210 to indicate information about the operation to be performed. In at least one embodiment, the stream operation API call 202, when called, receives one or more parameters 204, 206, 208, 210 to indicate information about the instruction to be executed.

[0066] In at least one embodiment, the stream operation API call 202 receives as input a parameter 204, 206, 208, 210 that includes a stream identifier 204. In at least one embodiment, the stream identifier 204 is a data value that includes information that can be used to identify a set of operations to be executed by a PPU (e.g., GPU) and / or one or more accelerators within the heterogeneous processor. In at least one embodiment, the stream identifier 204 is a data value that includes information that can be used to identify a set of instructions to be executed by a PPU (e.g., GPU) and / or one or more accelerators within the heterogeneous processor. In at least one embodiment, the stream identifier 204 is a data value that is used to indicate to the API a set of operations or instructions to be executed by one or more PPUs (e.g., GPU) and / or one or more accelerators within the heterogeneous processor. In at least one embodiment, the stream identifier 204 is a pointer to a stream. In at least one embodiment, the stream identifier 204 is a data structure for identifying a stream, such as described below. In at least one embodiment, the stream identifier 204 is a pointer to a data structure for identifying a stream. In at least one embodiment, the stream identifier 204 is a CUStream, which is defined as follows:

[0067] CUstream usrStream

[0068] In at least one embodiment, the stream identifier 204 is any other data or data type that can be used to identify the stream or other sequence of computational operations described herein.

[0069] In at least one embodiment, the stream operation API call 202 receives as input parameters 204, 206, 208, 210 that include an operation list 206. In at least one embodiment, the operation list 206 parameter is a data value that includes information for indicating one or more operations to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the operation list 206 parameter is a data value that includes information for indicating one or more instructions to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the operation list 206 is a set of data values that is used to indicate to the API a set of operations or instructions to be performed by one or more accelerators within a heterogeneous processor. In at least one embodiment, the operation list 206 is a pointer to a set of operations to be performed. In at least one embodiment, the operation list 206 is a data structure that is used to identify one or more operations to be performed and / or characteristics of the operations to be performed, such as those described below. In at least one embodiment, the operation list 206 is a pointer to a data structure that is used to identify one or more operations to be performed and / or characteristics of the operations to be performed. In at least one embodiment, the operation list is a cuSocketStreamOp, as described below. In at least one embodiment, the stream identifier 204 is any other data or data type that can be used to identify the stream or other sequence of computational operations described herein.

[0070] In at least one embodiment, the stream operation API call 202 receives as input parameters 204, 206, 208, 210 that include the number of operations 208. In at least one embodiment, the number of operations 208 parameter is a data value that includes information for indicating the number of operations indicated by the operation list 206 parameter. In at least one embodiment, the number of operations 208 parameter is a data value that includes information for indicating the number of instructions indicated by the operation list 206 parameter. In at least one embodiment, the number of operations 208 parameter is a positive integer value. In at least one embodiment, the number of operations 208 parameter is any other type of numerical value.

[0071] In at least one embodiment, the stream operation API call 202 receives as input parameters 204, 206, 208, 210 that include additional parameters 210. In at least one embodiment, the additional parameters 210 are data that includes information for indicating any other information that the API may use in response to the stream operation API call 202.

[0072] In at least one embodiment, if the stream operation API call 202 is invoked, it causes the API 108 to add one or more operations or instructions indicated by the operation list 206 parameter, which will be added, inserted, or otherwise included in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, if the stream operation API call 202 is invoked, it causes the API 108 in a parallel computing environment 106 (e.g., Compute Unified Device Architecture (CUDA)) to add one or more operations or instructions indicated by the operation list 206 parameter, which will be added, inserted, or otherwise included in a stream or instruction set to be executed by one or more accelerators within a heterogeneous processor.

[0073] In at least one embodiment, in response to the stream operation API call 202, if the API 108 is executed, it causes one or more processors to execute the stream operation API return 212. In at least one embodiment, the stream operation API return 212 is a set of instructions that, if executed, generate and / or indicate one or more data values in response to the stream operation API call 202. In at least one embodiment, the stream operation API return 212 indicates a success identifier 214. In at least one embodiment, the success identifier 214 is data including any value used to indicate the success of the stream operation API call 202. In at least one embodiment, the stream operation API return 212 indicates an error identifier 216. In at least one embodiment, the error identifier 216 is data including any value used to indicate the failure of the stream operation API call 202. In at least one embodiment, the error identifier 216 includes information indicating one or more specific types of errors generated as a result of the stream operation API call 202. In at least one embodiment, the error identifier 216 includes information indicating one or more other data values generated as a result of the stream operation API call 202.

[0074] In at least one embodiment, a parallel computing environment 106 that includes an API 108 (the API 108 includes a stream operation API 202) adds various types of operations to a stream for execution by one or more accelerators within a heterogeneous processor. In at least one embodiment, the stream operations include an acquire semaphore operation. In at least one embodiment, the stream operations include a release semaphore operation. In at least one embodiment, the stream operations include one or more operations for flushing and / or invalidating a cache memory, such as an L2 cache memory of a PPU (e.g., GPU), and / or a cache memory of one or more accelerators within the heterogeneous processor. In at least one embodiment, the stream operations include one or more operations for indicating the submission of an operation to an external device (e.g., one or more accelerators within the heterogeneous processor). In at least one embodiment, an example software code indicating the type of stream operation is as follows:

[0075]

[0076]

[0077] In at least one embodiment, the parallel computing environment 106 includes an API 108, the API 108 includes a stream operation API 202, and the stream operation API 202 includes one or more function signatures that can be used to indicate one or more callback functions for operations to be executed by one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more operations in the operation list 206 cause one or more callback functions to be executed. In at least one embodiment, an example software code indicating the function signature of a callback function is as follows:

[0078]

[0079] In at least one embodiment, to specify that one or more accelerators within a heterogeneous processor execute an operation list 206 indicated to the API 108 by a stream operation API call 202, a data structure of the API 108 can be used to specify one or more external devices to which the API 108 is to submit the operation list for which. In at least one embodiment, an example software code indicating a data structure of a device node representing one or more accelerators within a heterogeneous processor is as follows:

[0080]

[0081] In at least one embodiment, to specify the types and data of one or more operations indicated by the operation list 206 and to be executed by one or more accelerators within the heterogeneous processor, a data structure of API 108 will be used. In at least one embodiment, an example software code indicating a data structure for specifying the types and data of one or more operations to be executed by one or more accelerators within the heterogeneous processor is as follows:

[0082]

[0083]

[0084] In at least one embodiment, API 108 includes instructions which, if executed, cause one or more operations or instructions to be added to a stream or other instruction set for execution by one or more accelerators within the heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be added to a stream or other instruction set will be executed in response to a stream operation API call 202 as described above. In at least one embodiment, an example software code indicating a stream operation API call in the parallel computing environment 106 (e.g., CUDA) is as follows:

[0085]

[0086] In at least one embodiment, API 108 includes instructions which, if executed, cause one or more operations or instructions to be added to one or more executable graphs to be executed by one or more accelerators within the heterogeneous processor, in a manner similar to how one or more operations or instructions to be executed by one or more accelerators within the heterogeneous processor are added to one or more streams or instruction sets in response to a stream operation API call 202. In at least one embodiment, API 108 includes instructions which, if executed, cause one or more operations or instructions to be added to one or more executable graphs in response to a stream operation API call 202 for submission to one or more streams or instruction sets for execution. In at least one embodiment, an example software code indicating that API 108 of the parallel computing environment 106 adds one or more operations or instructions to one or more executable graphs is as follows:

[0087]

[0088] Figure 3An application programming interface (API) according to at least one embodiment is shown for performing a memory operation 302 to indicate memory that will be transferred between one or more parallel processing units (PPUs) (e.g., a graphics processing unit (GPU)) and one or more accelerators within a heterogeneous processor as described above. In at least one embodiment, the API 302 is a set of instructions that, if executed, cause one or more processors to execute one or more functions in response to one or more API calls 302. In at least one embodiment, an API call 302 is a set of instructions that, if executed, cause one or more processors to execute the API. In at least one embodiment, an API call 302 is a function call. In at least one embodiment, an API call 302 is a software function to be called by one or more software programs. In at least one embodiment, in response to an API call 302, one or more processors are used to execute a set of instructions and then return 310. In at least one embodiment, the return 310 is a change in the control flow from the API 302 to the software program after calling the API 302. In at least one embodiment, the return 310 causes one or more data values 312, 314 to be transferred into a memory accessible by one or more software programs.

[0089] In at least one embodiment, the API and / or API call 302 is a memory operation API call 302. In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to indicate data stored in the memory of one or more first accelerators that is to be copied between the memory of the one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to indicate data stored in a set or more sets of memory addresses that are to be copied between the one or more first accelerators and the one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to indicate a set or more sets of memory addresses that are to be copied between the one or more first accelerators and the one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to transfer data between the memory of the one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to transfer any other information between the memory of the one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor).

[0090] In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to indicate pointers that can be used to access data to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (such as GPUs and / or accelerators within a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to generate and / or indicate a data structure that includes one or more memory addresses that can be used to access data to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (such as GPUs and / or accelerators within a heterogeneous processor).

[0091] In at least one embodiment, if the memory operation API call 302 is invoked by one or more software programs, the API is caused to indicate one or more memory regions to be transferred from one or more first accelerators to one or more second accelerators. In at least one embodiment, if the memory operation API call 302 is invoked by one or more software programs, the API is caused to indicate one or more data sets to be transferred from one or more first accelerators to one or more second accelerators. In at least one embodiment, if the memory operation API call 302 is invoked by one or more software programs, the API is caused to indicate one or more memory addresses to which data is to be stored and that are to be transferred from one or more first accelerators to one or more second accelerators. In at least one embodiment, if the memory operation API call 302 is invoked by one or more software programs, the API is caused to indicate one or more accelerators within a heterogeneous processor that include memory storing data to be transferred, copied, or otherwise moved to the memory of one or more other accelerators (such as GPUs and / or accelerators within a heterogeneous processor).

[0092] In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to transfer data between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, one or more circuits of the processor execute an API in response to the memory operation API call 302 to indicate one or more regions of the memory of one or more first accelerators that are to be copied to the memory of one or more second accelerators. In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to indicate data stored in one or more sets of memory addresses that are to be copied between one or more first accelerators and one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to indicate one or more sets of memory addresses that are to be copied between one or more first accelerators and one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to transfer data between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to transfer any other information between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to indicate a pointer that can be used to access data that is to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 causes or otherwise executes an API in one or more circuits in a processor to generate and / or indicate a data structure that includes one or more memory addresses that can be used to access data that is to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., accelerators within a GPU and / or a heterogeneous processor).

[0093] In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to transfer data between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, one or more processors in the system are used to execute an API in response to the memory operation API call 302 to transfer data between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to indicate one or more regions of the memory of one or more first accelerators that are to be copied into the memory of one or more second accelerators. In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to indicate data stored in one or more sets of memory addresses that are to be copied between one or more first accelerators and one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to indicate one or more sets of memory addresses that are to be copied between one or more first accelerators and one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to transfer data between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to transfer any other information between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to indicate a pointer that can be used to access data that is to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor).In at least one embodiment, the memory operation API call 302 is used to cause one or more processors in the system to execute an API to generate and / or indicate a data structure including one or more memory addresses that can be used to access data to be transferred between the memory of one or more first accelerators and the memory of one or more second accelerators (e.g., an accelerator within a GPU and / or a heterogeneous processor).

[0094] In at least one embodiment, when the memory operation API call 302 is invoked, it receives one or more parameters 304, 306, 308 to indicate information about the operation to be performed. In at least one embodiment, when the memory operation API call 302 is invoked, it receives one or more parameters 304, 306, 308 to indicate information about the instruction to be executed.

[0095] In at least one embodiment, the memory operation API call 202 receives as input a parameter 304 that includes an input pointer 304. In at least one embodiment, the input pointer 304 is data that includes information that can be used to identify a device (e.g., one or more accelerators within a GPU and / or a heterogeneous processor) that contains the memory to be transferred. In at least one embodiment, the input pointer 304 is data that includes one or more memory addresses that are used to indicate information to be transferred from one or more first accelerators to one or more second accelerators. In at least one embodiment, the input pointer 304 is data that includes one or more memory addresses that are used to indicate information to be transferred from the memory of one or more first accelerators to the memory of one or more second accelerators. In at least one embodiment, the input pointer 304 is data for indicating a pointer for which memory information is to be determined. In at least one embodiment, the input pointer 304 is data for indicating the memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the input pointer 304 is a Compute Unified Device Architecture (CUDA) pointer or any other type of pointer further described herein. In at least one embodiment, the input pointer 304 is a pointer. In at least one embodiment, the input pointer 304 is a data structure that includes one or more pointers to memory. In at least one embodiment, the input pointer 304 is a pointer to a data structure for identifying one or more locations in memory and / or data at one or more locations in memory, as described below. In at least one embodiment, the input pointer 304 is any other data or data type that can be used to identify data in memory and / or one or more locations in memory that include data, as described below.

[0096] In at least one embodiment, the memory operation API call 302 receives as inputs parameters 304, 306, 308 including the parameters of the output structure 306. In at least one embodiment, the parameters of the output structure 306 are data including information indicating a data structure that contains a handle and offset information regarding one or more data values in the memory. In at least one embodiment, the parameters of the output structure 306 are data including information indicating a data structure that contains a handle and offset information regarding one or more data values in the memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the handle information is one or more memory addresses, such as a pointer. In at least one embodiment, the offset information is one or more numerical values for indicating a position in the memory. In at least one embodiment, the handle information is one or more memory addresses, such as a pointer, for indicating a position in the memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the offset information is one or more numerical values for indicating a position in the memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the parameters of the output structure 306 are a data structure including one or more data values for indicating a handle and offset information regarding one or more positions in the memory. In at least one embodiment, the parameters of the output structure 306 are a data structure including one or more data values for indicating any other information regarding the memory. In at least one embodiment, the parameters of the output structure 306 are a data structure including one or more data values for indicating any other information regarding the memory of one or more accelerators within a heterogeneous processor. In at least one embodiment, the parameters of the output structure 306 are a pointer. In at least one embodiment, the parameters of the output structure 306 are a data structure including one or more pointers to the memory. In at least one embodiment, the parameters of the output structure 306 are a pointer to a data structure for identifying one or more positions in the memory and / or data at one or more positions in the memory, as described below. In at least one embodiment, the parameters of the output structure 306 are any other data or data type that can be used to identify data in the memory and / or one or more positions in the memory that include the data, as further described herein.

[0097] In at least one embodiment, the memory operation API call 302 receives parameters 304, 306, 308 including additional parameter 308 as input. In at least one embodiment, the additional parameter 308 is data including information for indicating any additional information available to the API in response to the memory operation API call 302. In at least one embodiment, the additional parameter 308 includes information for indicating one or more flags that can be used to configure one or more operations in response to the memory operation API call 302. In at least one embodiment, the additional parameter 308 includes information for indicating any additional information that can be used to perform and / or configure to perform one or more operations in response to the memory operation API call 302.

[0098] In at least one embodiment, if the memory operation API call 302 is invoked, it causes the API 108 to transfer information between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor). In at least one embodiment, if the memory operation API call 302 is invoked, it causes the API 108 in a parallel computing environment 106 (such as Compute Unified Device Architecture (CUDA)) to transfer information between the memory of one or more first accelerators and the memory of one or more second accelerators (such as accelerators within a GPU and / or a heterogeneous processor).

[0099] In at least one embodiment, in response to a memory operation API call 302, API 108, if executed, causes one or more processors to perform a memory operation API return 310. In at least one embodiment, the memory operation API return 310 is a set of instructions that, if executed, generate and / or indicate one or more data values in response to the memory operation API call 302. In at least one embodiment, the memory operation API return 310 indicates a success identifier 312. In at least one embodiment, the success identifier 312 is data that includes any value used to indicate the success or successful operation of the memory operation API call 302. In at least one embodiment, the memory operation API return 310 indicates an error identifier 314. In at least one embodiment, the error identifier 314 is data that includes any value used to indicate the failure or failed operation of the memory operation API call 302. In at least one embodiment, the error identifier 314 includes information indicating one or more specific types of errors generated as a result of or in response to the memory operation API call 302. In at least one embodiment, the error identifier 314 includes information indicating one or more other data values generated in response to or as a result of the memory operation API call 302.

[0100] In at least one embodiment, the output structure 306 is a data structure that includes information for indicating one or more memory locations and / or the contents of one or more memory locations. In at least one embodiment, the output structure 306 is a data structure that includes information for indicating one or more memory locations and / or the contents of one or more memory locations of one or more accelerators within a heterogeneous processor. In at least one embodiment, an example software code of a data structure for indicating one or more memory locations and / or the contents of one or more memory locations of one or more accelerators within a heterogeneous processor is as follows:

[0101]

[0102] In at least one embodiment, API 108 includes instructions that, if executed, cause information to be transferred between the memories of one or more first accelerators and the memories of one or more second accelerators (such as accelerators within a heterogeneous processor). In at least one embodiment, the instructions for causing information to be transferred between the memories of one or more first accelerators and the memories of one or more second accelerators are to be executed in response to the memory operation API call 302 as described above. In at least one embodiment, an example software code for indicating the memory operation API call 302 in a parallel computing environment 106 (such as CUDA) is as follows:

[0103] / **

[0104] * Obtain the memory information of a CUDA pointer.

[0105] *

[0106] * - param[in] ptr - The input CUDA pointer for which memory information is being requested.

[0107] * - param[out] memPtrInfo - The output structure where the handle and offset information will be available

[0108] *

[0109] * - Returns CUDA_SUCCESS if successful, otherwise returns an appropriate error.

[0110] * /

[0111] CUresult cuSocketMemPtrGetInfo(

[0112] void* ptr,

[0113] cuSocketMemPtrInfo* memPtrInfo,

[0114] unsigned int flags );

[0116] Figure 4An application programming interface (API) 402 according to at least one embodiment is shown, which is used to indicate one or more memory regions that will be available for storing error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, if the API 402 is executed, it causes one or more processors to poll for one or more errors from the one or more accelerators within the heterogeneous processor by checking the error information in the one or more memory regions indicated in response to the API 402 and generated by the one or more accelerators within the heterogeneous processor. In at least one embodiment, if the API 402 is executed, it causes one or more processors to poll for one or more errors from the one or more accelerators within the heterogeneous processor by checking the error information in the one or more memory regions indicated using the API 402 and generated by the one or more accelerators within the heterogeneous processor. In at least one embodiment, if the API 402 is executed, it causes one or more processors to indicate one or more memory regions that will be polled for error information indicating one or more errors from the one or more accelerators within the heterogeneous processor.

[0117] In at least one embodiment, the API 402 is a set of instructions that, if executed, causes one or more processors to execute one or more functions in response to one or more API calls 402. In at least one embodiment, an API call 402 is a set of instructions that, if executed, causes one or more processors to execute the API. In at least one embodiment, an API call 402 is a function call. In at least one embodiment, an API call 402 is a software function that will be called by one or more software programs. In at least one embodiment, in response to an API call 402, one or more processors are used to execute a set of instructions and then return 410. In at least one embodiment, the return 410 is a change in the control flow from the API 402 to the software program after calling the API 402 (e.g., by the API call 402). In at least one embodiment, the return 410 causes one or more data values 412, 414 to be transferred to a memory accessible by one or more software programs.

[0118] In at least one embodiment, the API and / or API call 402 is a register error notification buffer API call 402. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to poll to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to poll to obtain one or more errors from the one or more accelerators within the heterogeneous processor by checking error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by the one or more accelerators within the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to indicate one or more memory regions that are to be polled to obtain error information indicating one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to indicate one or more buffers in memory that are to be polled to obtain error information indicating one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to indicate one or more buffers in the memory of a central processing unit (CPU) that are to be polled to obtain error information indicating one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed, cause the API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (such as a graphics processing unit (GPU)) that are to be polled to obtain error information indicating one or more errors from one or more accelerators within the heterogeneous processor.

[0119] In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to poll to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to poll to obtain one or more errors from one or more accelerators within a heterogeneous processor by examining error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by the one or more accelerators within the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to indicate one or more memory regions that will be polled to obtain error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to indicate one or more buffers in memory that will be polled to obtain error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to indicate one or more buffers in the memory of a central processing unit (CPU) that will be polled to obtain error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to execute the API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)) that will be polled to obtain error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0120] In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to poll to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to poll for one or more errors from one or more accelerators within a heterogeneous processor by checking error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by one or more accelerators in the heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to indicate one or more memory regions that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to indicate one or more buffers in memory that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to indicate one or more buffers in the memory of a central processing unit (CPU) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification buffer API call 402, if invoked by one or more software programs, causes the API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (such as a graphics processing unit (GPU)) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0121] In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to poll for one or more errors from one or more accelerators within a heterogeneous processor by checking error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by one or more accelerators in the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to indicate one or more memory regions that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to indicate one or more buffers in memory that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to indicate one or more buffers in the memory of a central processing unit (CPU) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 is used to cause or otherwise execute, by one or more circuits in a processor, an API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (such as a graphics processing unit (GPU)) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0122] In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to poll to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to poll for one or more errors from one or more accelerators within a heterogeneous processor by checking error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by one or more accelerators in the heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to indicate one or more memory regions that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to indicate one or more buffers in memory that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to indicate one or more buffers in the memory of a central processing unit (CPU) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification buffer API call 402 causes one or more processors in the system to execute an API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (such as a graphics processing unit (GPU)) that will be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0123] In at least one embodiment, the register error notification buffer API call 402 receives, when called, one or more parameters 404, 406, 408 for indicating information about an operation to be performed. In at least one embodiment, the register error notification buffer API call 402 receives, when called, one or more parameters 404, 406, 408 for indicating information about an instruction to be executed. In at least one embodiment, the register error notification buffer API call 402 receives, when called, one or more parameters 404, 406, 408 for indicating information about a memory to be polled. In at least one embodiment, the register error notification buffer API call 402 receives, when called, one or more parameters 404, 406, 408 for indicating information about a memory that will be polled to obtain one or more errors from one or more accelerators within a heterogeneous processor.

[0124] In at least one embodiment, the register error notification buffer API call 402 receives as input parameters 404, 406, 408 that include a buffer array 404. In at least one embodiment, the buffer array 404 is a set of one or more indicators that include one or more buffers in memory, as described above. In at least one embodiment, the buffer array 404 is data that includes information indicating memory (e.g., CPU and / or GPU memory) for storing error information. In at least one embodiment, the buffer array 404 is data that includes information indicating memory (e.g., CPU and / or GPU memory) for storing error information regarding one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 404 is data that includes information indicating memory (e.g., CPU and / or GPU memory) for storing one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 404 is data that includes information indicating one or more buffers in memory that will be polled to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 404 is data that includes an array of memory locations for storing error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 404 is data that includes an array of memory locations that will be polled to identify one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 404 parameter is a pointer. In at least one embodiment, the buffer array 404 parameter is a data structure that includes one or more buffers. In at least one embodiment, the buffer array 404 parameter is a pointer to a data structure, as described below. In at least one embodiment, the buffer array 404 parameter is any other data or data type that can be used to identify one or more buffers, as further described herein.

[0125] In at least one embodiment, the register error notification buffer API call 402 receives parameters 404, 406, 408 including the buffer count parameter 406 as input. In at least one embodiment, the buffer count parameter 406 is data including information for indicating the number of elements indicated by the buffer array parameter 404. In at least one embodiment, the buffer count parameter 406 is data including information for indicating the number of buffers indicated by the buffer array parameter 404. In at least one embodiment, the buffer count parameter 406 is data including information for indicating the number of buffers to be polled for error information. In at least one embodiment, the buffer count parameter 406 is data including information for indicating the number of buffers to be polled for error information generated by one or more accelerators within a heterogeneous processor.

[0126] In at least one embodiment, the register error notification buffer API call 402 receives parameters 404, 406, 408 including other parameter 408 as input. In at least one embodiment, the other parameter 408 is data including information for indicating any other information available to the API in response to the register error notification buffer API call 402. In at least one embodiment, the other parameter 408 includes information for indicating any other information available to perform one or more operations and / or configured to perform one or more operations in response to the register error notification buffer API call 402.

[0127] In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 to poll for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 to indicate the memory that will be polled for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 to indicate one or more buffers in the memory that will be polled for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 in the parallel computing environment 106 (such as Compute Unified Device Architecture (CUDA)) to poll for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 in the parallel computing environment 106 (such as CUDA) to indicate the memory that will be polled for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, if the registered error notification buffer API call 402 is invoked, it causes the API 108 in the parallel computing environment 106 (such as CUDA) to indicate one or more buffers in the memory that will be polled for one or more errors from one or more accelerators within the heterogeneous processor.

[0128] In at least one embodiment, in response to a registration error notification buffer API call 402, if API 108 is executed, it causes one or more processors to perform a registration error notification buffer API return 410. In at least one embodiment, the registration error notification buffer API return 410 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the registration error notification buffer API call 402. In at least one embodiment, the registration error notification buffer API return 410 indicates a success identifier 412. In at least one embodiment, the success identifier 412 is data including any value used to indicate the success or successful operation of the registration error notification buffer API call 402. In at least one embodiment, the registration error notification buffer API return 410 indicates an error identifier 414. In at least one embodiment, the error identifier 414 is data including any value used to indicate the failure or failed operation of the registration error notification buffer API call 402. In at least one embodiment, the error identifier 414 includes information indicating one or more specific types of errors generated as a result of or in response to the registration error notification buffer API call 402. In at least one embodiment, the error identifier 414 includes information indicating one or more other data values generated in response to or as a result of the registration error notification buffer API call 402.

[0129] In at least one embodiment, to specify the type of accelerator within a heterogeneous processor for generating errors that will be polled by API 108, a data type will be declared. In at least one embodiment, an example software code indicating the data type for specifying the type of accelerator in a heterogeneous processor is as follows:

[0130]

[0131] In at least one embodiment, to specify one or more buffers or other regions of memory for storing error information generated by one or more accelerators within a heterogeneous processor and being polled for that error information, a data structure of API 108 will be used. In at least one embodiment, an example software code indicating the data structure for indicating one or more buffers or other regions of memory for storing error information generated by one or more accelerators within a heterogeneous processor and being polled for that error information is as follows:

[0132]

[0133] In at least one embodiment, API 108 includes instructions which, if executed, cause one or more buffers to store error information to be generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions which, if executed, cause one or more buffers to be polled for error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, instructions for indicating a memory to be polled for error information generated by one or more accelerators in a heterogeneous processor are to be executed in response to a registered error notification buffer API call 402 as described above. In at least one embodiment, an example software code indicating a registered error notification buffer API call 402 in a parallel computing environment 106 (e.g., CUDA) is as follows:

[0134]

[0135] Figure 5 An application programming interface (API) 502 is shown in accordance with at least one embodiment, which is for indicating one or more memory regions that are no longer available for storing error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, API 502 is for causing one or more processors to unregister or otherwise indicate one or more buffers that are no longer available for storing error information from one or more accelerators within a heterogeneous processor (e.g., executed in response to API 402 described above). In at least one embodiment, API 502 is for causing one or more processors to stop polling for one or more errors from one or more accelerators within a heterogeneous processor, e.g., executed in response to API 402 as described above. Figure 4 described API 402). Figure 4 executed.

[0136] In at least one embodiment, if executed, API 502 causes one or more processors to stop polling for one or more errors from one or more accelerators within a heterogeneous processor by unregistering one or more buffers storing error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, if executed, API 502 is for causing one or more processors to stop polling for one or more errors from one or more accelerators within a heterogeneous processor by unregistering one or more buffers in a parallel computing environment, as described above. Figure 1As described above, one or more buffers store error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, if the API 502 is executed, it causes one or more processors to indicate one or more memory regions that are to be de-registered from the parallel computing environment, such that the one or more memory regions no longer receive error information indicating one or more errors from one or more accelerators within the heterogeneous processor.

[0137] In at least one embodiment, the API 502 is a set of instructions that, if executed, causes one or more processors to execute one or more functions in response to one or more API calls 502. In at least one embodiment, an API call 502 is a set of instructions that, if executed, causes one or more processors to execute the API. In at least one embodiment, an API call 502 is a function call. In at least one embodiment, an API call 502 is a software function to be called by one or more software programs. In at least one embodiment, in response to an API call 502, one or more processors are used to execute a set of instructions and then return 510. In at least one embodiment, the return 510 is a change in the control flow from the API 502 to the software program after calling the API 502 (e.g., by the API call 502). In at least one embodiment, the return 510 causes one or more data values 512, 514 to be transferred to a memory accessible by one or more software programs.

[0138] In at least one embodiment, the API and / or the API call 502 is a de-register error notification buffer API call 502. In at least one embodiment, the de-register error notification buffer API call 502 is a set of software instructions that, if executed, causes the API to de-register one or more buffers from the parallel computing environment, such as those described above in connection with Figure 1As described above, one or more buffers are used to receive information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to stop polling for one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to unregister one or more memory regions (e.g., buffers) available for polling for one or more errors from one or more accelerators within the heterogeneous processor by checking for error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by one or more accelerators within the heterogeneous processor, as described above in connection with Figure 4 As described above. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to indicate that one or more memory regions for polling for error information indicating one or more errors from one or more accelerators within the heterogeneous processor are no longer to be polled. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to indicate to a parallel computing environment (e.g., Compute Unified Device Architecture (CUDA)) that one or more buffers in memory for polling for error information indicating one or more errors from one or more accelerators within the heterogeneous processor are no longer to be polled. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to indicate that one or more buffers in the memory of a central processing unit (CPU) for polling for error information indicating one or more errors from one or more accelerators within the heterogeneous processor are no longer to be polled. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed, cause the API to indicate that one or more buffers in the memory of one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)) for polling for error information indicating one or more errors from one or more accelerators within the heterogeneous processor are no longer to be polled.

[0139] In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to unregister one or more memory regions (e.g., buffers) from a parallel computing environment, where the one or more memory regions are available for polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to stop polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to unregister one or more buffers that are available for polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to indicate one or more memory regions that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to indicate one or more buffers in memory that are to be unregistered and made unavailable for polling error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to indicate one or more buffers in the memory of a central processing unit (CPU) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to execute an API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0140] In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to stop polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to unregister one or more buffers from the parallel computing environment, as described above in connection with Figure 1 which the one or more buffers may be used to poll for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to indicate one or more memory regions that will be unregistered from the parallel computing environment to no longer receive and / or store error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to indicate one or more buffers in memory that can no longer be polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to indicate one or more buffers in the memory of a central processing unit (CPU) that are no longer polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502, if invoked by one or more software programs, causes the API to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)) that are no longer polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0141] In at least one embodiment, the unregister error notification buffer API call 502 causes or otherwise executes an API by one or more circuits in a processor to unregister one or more memory regions (e.g., buffers) from a parallel computing environment. In at least one embodiment, the unregister error notification buffer API call 502 causes or otherwise executes an API by one or more circuits in a processor to unregister one or more memory regions (e.g., buffers) from a parallel computing environment, where the one or more memory regions are available for polling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 causes or otherwise executes an API by one or more circuits in a processor to stop a parallel computing environment from polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 causes or otherwise executes an API by one or more circuits in a processor to stop polling for one or more errors from the one or more accelerators within a heterogeneous processor by stopping checking for error information in one or more memory regions indicated in response to the register error notification buffer API call 402 and generated by the one or more accelerators in the heterogeneous processor, as described above in connection with Figure 4As described. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause or otherwise execute an API by one or more circuits in a processor to indicate one or more memory regions that are no longer being polled for error information indicating one or more errors from one or more accelerators in a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause or otherwise execute an API by one or more circuits in a processor to indicate the one or more buffers that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor by unregistering one or more buffers in memory from a parallel computing environment (such as CUDA). In at least one embodiment, the unregister error notification buffer API call 502 is used to cause or otherwise execute an API by one or more circuits in a processor to indicate one or more buffers in the memory of a central processing unit (CPU) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause or otherwise execute an API by one or more circuits in a processor to indicate one or more buffers in the memory of one or more parallel processing units (PPUs) (such as a graphics processing unit (GPU)) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0142] In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to stop polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to stop polling for one or more errors by unregistering one or more memory regions (e.g., buffers) that are available for polling for one or more errors from one or more accelerators within a heterogeneous processor by a parallel computing environment (e.g., CUDA). In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to stop polling to obtain one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to stop polling to obtain errors from one or more accelerators within a heterogeneous processor by unregistering one or more buffers that are available for checking error information in one or more memory regions generated by one or more accelerators within a heterogeneous processor, where the one or more buffers are indicated in response to a register error notification buffer API call 402, as described above in connection with Figure 4 the foregoing. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to indicate that one or more memory regions that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to indicate that one or more buffers in memory that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to indicate that one or more buffers in the memory of a central processing unit (CPU) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification buffer API call 502 is used to cause one or more processors in the system to execute an API to indicate that one or more buffers in the memory of one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)) that are no longer being polled for error information indicating one or more errors from one or more accelerators within a heterogeneous processor.

[0143] In at least one embodiment, the unregister error notification buffer API call 502 receives, when called, one or more parameters 504, 506, 508 for indicating information about an operation to be performed. In at least one embodiment, the unregister error notification buffer API call 502 receives, when called, one or more parameters 504, 506, 508 for indicating information about an instruction to be executed. In at least one embodiment, the unregister error notification buffer API call 502 receives, when called, one or more parameters 504, 506, 508 for indicating information about a memory to be unregistered from a parallel computing environment. In at least one embodiment, the unregister error notification buffer API call 502 receives, when called, one or more parameters 504, 506, 508 for indicating information about a memory that will be unregistered from a parallel computing environment to no longer store error information that will be polled to obtain one or more errors from one or more accelerators within a heterogeneous processor.

[0144] In at least one embodiment, the unregister error notification buffer API call 502 receives as input parameters 504, 506, 508 that include a buffer array 504. In at least one embodiment, the buffer array 504 is a set of one or more indicators that include one or more buffers in memory, as described above. In at least one embodiment, the buffer array 504 is data that includes information indicating memory that will be unregistered from the parallel computing environment (such that the memory is no longer available for storing error information), such as CPU and / or GPU memory. In at least one embodiment, the buffer array 504 is data that includes information indicating memory that will be unregistered from the parallel computing environment (such that the memory is no longer available for storing error information regarding one or more accelerators within a heterogeneous processor), such as CPU and / or GPU memory. In at least one embodiment, the buffer array 504 is data that includes information indicating memory (such as CPU and / or GPU memory) that will no longer store one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 504 is data that includes information indicating one or more buffers in memory that will no longer be polled for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 504 is data that includes an array of memory locations that will no longer store error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 504 is data that includes an array of memory locations that will no longer be polled to identify one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the buffer array 504 parameter is a pointer. In at least one embodiment, the buffer array 504 parameter is a data structure that includes one or more buffers. In at least one embodiment, the buffer array 504 parameter is a pointer to a data structure, as described below. In at least one embodiment, the buffer array 504 parameter is any other data or data type that can be used to identify one or more buffers, as further described herein.

[0145] In at least one embodiment, the unregister error notification buffer API call 502 receives as input parameters 504, 506, 508 that include a buffer count 506 parameter. In at least one embodiment, the buffer count 506 parameter is data that includes information for indicating the number of elements indicated by the buffer array 504 parameter. In at least one embodiment, the buffer count 506 parameter is data that includes information for indicating the number of buffers indicated by the buffer array 504 parameter. In at least one embodiment, the buffer count 506 parameter is data that includes information for indicating the number of the buffers that will be unregistered from the parallel computing environment such that the buffers are no longer polled for error information. In at least one embodiment, the buffer count 506 parameter is data that includes information for indicating the number of the buffers that will be unregistered from the parallel computing environment such that the buffers are no longer pollable for error information generated by one or more accelerators within a heterogeneous processor.

[0146] In at least one embodiment, the unregister error notification buffer API call 502 receives as input parameters 504, 506, 508 that include additional parameter 508. In at least one embodiment, the additional parameter 508 is data that includes information for indicating any additional information available to the API in response to the unregister error notification buffer API call 502. In at least one embodiment, the additional parameter 508 includes information for indicating any additional information that can be used to perform one or more operations in response to the unregister error notification buffer API call 502 and / or is configured to perform one or more operations in response to the unregister error notification buffer API call 502.

[0147] In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 to unregister one or more memory regions (e.g., buffers) from the parallel computing environment 106 such that the one or more memory regions are no longer available for receiving and / or storing error information. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 to unregister one or more memory regions (e.g., buffers) from the parallel computing environment 106 such that the one or more memory regions are no longer available for receiving and / or storing error information that will be polled for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 to indicate a memory that is no longer being polled for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 to indicate one or more buffers in a memory that is no longer being polled for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 in the parallel computing environment 106 (e.g., Compute Unified Device Architecture (CUDA)) to stop polling for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 in the parallel computing environment 106 (e.g., CUDA) to indicate a memory that is no longer being polled for one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification buffer API call 502 is invoked, it causes the API 108 in the parallel computing environment 106 (e.g., CUDA) to indicate one or more buffers in a memory that is no longer being polled for one or more errors from one or more accelerators within a heterogeneous processor.

[0148] In at least one embodiment, in response to a de-register error notification buffer API call 502, API 108, if executed, causes one or more processors to perform a de-register error notification buffer API return 510. In at least one embodiment, the de-register error notification buffer API return 510 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the de-register error notification buffer API call 502. In at least one embodiment, the de-register error notification buffer API return 510 indicates a success identifier 512. In at least one embodiment, the success identifier 512 is data including any value used to indicate the success or successful operation of the de-register error notification buffer API call 502. In at least one embodiment, the de-register error notification buffer API return 510 indicates an error identifier 514. In at least one embodiment, the error identifier 514 is data including any value used to indicate the failure or failed operation of the de-register error notification buffer API call 502. In at least one embodiment, the error identifier 514 includes information indicating one or more specific types of errors generated as a result of or in response to the de-register error notification buffer API call 502. In at least one embodiment, the error identifier 514 includes information indicating one or more other data values generated in response to or as a result of the de-register error notification buffer API call 502.

[0149] In at least one embodiment, API 108 includes instructions that, if executed, cause one or more buffers to be de-registered from the parallel computing environment 106, where the one or more buffers are available for storing error information to be generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions that, if executed, cause one or more buffers to no longer be polled for error information generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the instructions for indicating the memory that is to be de-registered and no longer polled for error information generated by one or more accelerators within a heterogeneous processor are to be executed in response to the de-register error notification buffer API call 502 as described above. In at least one embodiment, an example software code indicating the de-register error notification buffer API call 502 in a parallel computing environment 106 (e.g., CUDA) is as follows:

[0150]

[0151] Figure 6An application programming interface (API) 602 according to at least one embodiment is shown, which is used to indicate one or more instruction sequences to be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, if the API 602 is executed, it causes the parallel computing environment 106 to register one or more callback functions to be executed to handle one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the API 602 is used to identify one or more error handlers (such as callback functions) for handling one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, a callback function or callback is a set of instructions that, if executed, causes one or more operations to be performed in response to one or more events (such as detecting an error by polling, as described above in conjunction with Figure 4 ).

[0152] In at least one embodiment, the API 602 is a set of instructions that, if executed, causes one or more processors to execute one or more functions in response to one or more API calls 602. In at least one embodiment, an API call 602 is a set of instructions that, if executed, causes one or more processors to execute the API. In at least one embodiment, an API call 602 is a function call. In at least one embodiment, an API call 602 is a software function to be called by one or more software programs. In at least one embodiment, in response to an API call 602, one or more processors are used to execute a set of instructions and then return 608. In at least one embodiment, the return 608 is a change in the control flow from the API 602 to the software program after calling the API 602. In at least one embodiment, the return 608 causes one or more data values 610, 612 to be transferred to a memory accessible by one or more software programs.

[0153] In at least one embodiment, the API and / or API call 602 is a register error notification callback API call 602. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to indicate one or more instruction sets (e.g., the instruction sets of one or more callback functions) that will be executed in response to an event (e.g., an error detected by polling or other means). In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to indicate one or more callback functions that will be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to identify one or more error handlers (e.g., callback functions) for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to indicate one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to indicate to a parallel computing environment (e.g., Compute Unified Device Architecture (CUDA)) one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to register one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor to a parallel computing environment. In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to register one or more error handlers for handling one or more errors from one or more accelerators within a heterogeneous processor to a parallel computing environment.In at least one embodiment, the register error notification callback API call 602 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to use the API to identify, to a parallel computing environment, one or more callback functions or other error handlers that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, an error handler is a set of software instructions that is executed in response to one or more errors (e.g., errors generated by one or more accelerators within a heterogeneous processor).

[0154] In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to indicate one or more instruction sets (e.g., instruction sets of one or more callback functions) that will be executed in response to an event (e.g., an error detected via polling or other means). In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to indicate one or more callback functions that will be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to identify one or more error handlers (e.g., callback functions) for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to indicate one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to register with a parallel computing environment one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to register with a parallel computing environment one or more error handlers for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, a register error notification callback API call 602, if invoked by one or more software programs, causes the API to identify for a parallel computing environment one or more callback functions or other error handlers that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0155] In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to indicate one or more instruction sets (e.g., instruction sets of one or more callback functions) that will be executed in response to an event (e.g., detecting an error by polling or otherwise). In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to indicate one or more callback functions that will be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to identify one or more error handlers (e.g., callback functions) for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to indicate one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to register one or more callback functions to a parallel computing environment, which callback functions will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to register one or more error handlers for handling one or more errors from one or more accelerators within a heterogeneous processor to a parallel computing environment. In at least one embodiment, the register error notification callback API call 602 causes or otherwise executes an API in one or more circuits in a processor to identify to a parallel computing environment one or more callback functions or other error handlers that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0156] In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to indicate one or more instruction sets (e.g., the instruction sets of one or more callback functions) that will be executed in response to an event (e.g., an error detected by polling or other means). In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to indicate one or more callback functions that will be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to identify one or more error handlers (e.g., callback functions) for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to indicate one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to register with a parallel computing environment one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to register with a parallel computing environment one or more error handlers for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the register error notification callback API call 602 causes one or more processors in the system to execute the API to identify for a parallel computing environment one or more callback functions or other error handlers that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0157] In at least one embodiment, when the register error notification callback API call 602 is called, it receives one or more parameters 604, 606 for indicating information about the operation to be performed. In at least one embodiment, when the register error notification callback API call 602 is called, it receives one or more parameters 604, 606 for indicating information about the instruction to be executed.

[0158] In at least one embodiment, the register error notification callback API call 602 receives as input arguments 604, 606 that include one or more callback functions 604. In at least one embodiment, the one or more callback functions 604 are data that include information about one or more error handlers that can be used to identify one or more errors for processing from one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more callback functions 604 are data that include information about one or more instruction sets that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more callback functions 604 are data for indicating one or more error handlers. In at least one embodiment, the one or more callback functions 604 are data for indicating one or more callback functions. In at least one embodiment, the one or more callback functions 604 are data that include one or more memory addresses. In at least one embodiment, the one or more callback functions 604 are data that include one or more pointers. In at least one embodiment, the one or more callback functions 604 are data for indicating any other information about one or more error handlers for processing one or more errors from one or more accelerators within a heterogeneous processor, such as various data structures further described herein. In at least one embodiment, the one or more callback functions 604 include one or more CUDA pointers or any other type of pointer further described herein.

[0159] In at least one embodiment, the register error notification callback API call 602 receives as input arguments 604, 606 that include additional argument 606. In at least one embodiment, the additional argument 606 is data that includes information for indicating any other information available to the API in response to the register error notification callback API call 602. In at least one embodiment, the additional argument 606 includes information for indicating one or more flags that can be used to configure one or more operations in response to the register error notification callback API call 602. In at least one embodiment, the additional argument 606 includes information for indicating any other information that can be used to perform one or more operations and / or is configured to perform one or more operations in response to the register error notification callback API call 602.

[0160] In at least one embodiment, if the register error notification callback API call 602 is invoked, it causes the API 108 to register one or more callback functions with the parallel computing environment 106 (e.g., CUDA) to be executed in response to one or more errors. In at least one embodiment, if the register error notification callback API call 602 is invoked, it causes the API 108 to register one or more callback functions with the parallel computing environment 106 (e.g., CUDA) to be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the register error notification callback API call 602 is invoked, it causes the API 108 to register one or more error handlers with the parallel computing environment 106 (e.g., CUDA) to be executed in response to one or more errors. In at least one embodiment, if the register error notification callback API call 602 is invoked, it causes the API 108 to register one or more error handlers with the parallel computing environment 106 (e.g., CUDA) to be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0161] In at least one embodiment, in response to the register error notification callback API call 602, if the API 108 is executed, it causes one or more processors to execute the register error notification callback API return 608. In at least one embodiment, the register error notification callback API return 608 is a set of instructions that, if executed, generate and / or indicate one or more data values in response to the register error notification callback API call 602. In at least one embodiment, the register error notification callback API return 608 indicates a success identifier 610. In at least one embodiment, the success identifier 610 is data including any value used to indicate the success or successful operation of the register error notification callback API call 602. In at least one embodiment, the register error notification callback API return 608 indicates an error identifier 612. In at least one embodiment, the error identifier 612 is data including any value used to indicate the failure or failed operation of the register error notification callback API call 602. In at least one embodiment, the error identifier 612 includes information indicating one or more specific types of errors generated as a result of or in response to the register error notification callback API call 602. In at least one embodiment, the error identifier 612 includes information indicating one or more other data values generated in response to or as a result of the register error notification callback API call 602.

[0162] In at least one embodiment, to specify one or more callback functions (e.g., one or more error handlers) to handle one or more errors from one or more accelerators within a heterogeneous processor, the data structures of API 108 will be used. In at least one embodiment, to specify one or more callback functions (e.g., one or more error handlers) to handle one or more errors from one or more accelerators within a heterogeneous processor, the function pointers of API 108 will be used. In at least one embodiment, to specify one or more callback functions (e.g., one or more error handlers) to handle one or more errors from one or more accelerators within a heterogeneous processor, any other type of data of API 108 that indicates a set or more sets of software instructions will be used. In at least one embodiment, an example of software code that indicates one or more callback functions (e.g., one or more error handlers) to be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor is as follows:

[0163] / **

[0164] * Callback function signature for interpreting the error buffer.

[0165] * /

[0166] typedef CUresult(*cuSocketErrorCallback)(void*data);

[0167] In at least one embodiment, API 108 includes instructions that, if executed, cause one or more operations or instructions to be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions that, if executed, register one or more callback functions that will be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions that, if executed, identify one or more error handlers for handling one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to be registered with the parallel computing environment 106 (e.g., CUDA) or to be identified to handle one or more errors from one or more accelerators within a heterogeneous processor will be executed in response to the registration error notification callback API call 602, as described above. In at least one embodiment, an example of software code that indicates the registration error notification callback API call 602 in the parallel computing environment 106 (e.g., CUDA) is as follows:

[0168]

[0169] Figure 7 An application programming interface (API) 702 according to at least one embodiment is shown, which is used to indicate one or more instruction sequences that are no longer to be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, if the API 702 is executed, it causes the parallel computing environment 106 to unregister one or more callback functions so that they are no longer to be executed to handle one or more errors from one or more accelerators within the heterogeneous processor. In at least one embodiment, the API 702 is used to cause one or more error handlers (such as callback functions) to no longer handle one or more errors from one or more accelerators within the heterogeneous processor.

[0170] In at least one embodiment, the API 702 is a set of instructions that, if executed, causes one or more processors to execute one or more functions in response to one or more API calls 702. In at least one embodiment, an API call 702 is a set of instructions that, if executed, causes one or more processors to execute the API. In at least one embodiment, an API call 702 is a function call. In at least one embodiment, an API call 702 is a software function to be called by one or more software programs. In at least one embodiment, in response to an API call 702, one or more processors are used to execute a set of instructions and then return 708. In at least one embodiment, the return 708 is a change in the control flow from the API 702 to the software program after calling the API 702. In at least one embodiment, the return 708 causes one or more data values 710, 712 to be transferred to a memory accessible by one or more software programs.

[0171] In at least one embodiment, the API and / or API call 702 is a logout error notification callback API call 702. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to indicate that one or more instruction sets (e.g., the instruction sets of one or more callback functions) that would otherwise be executed in response to an event (such as detecting an error via polling or other means) will no longer be executed. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to indicate that one or more callback functions that would otherwise be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor will no longer be executed. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to identify one or more error handlers (such as callback functions) that will no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to indicate that one or more callback functions that would otherwise be executed in response to one or more errors from one or more accelerators within a heterogeneous processor will no longer be executed. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to indicate to a parallel computing environment (such as Compute Unified Device Architecture (CUDA)) that one or more callback functions that would otherwise be executed in response to one or more errors from one or more accelerators within a heterogeneous processor will no longer be executed. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to unregister one or more callback functions from a parallel computing environment so that they will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, causes the one or more processors to use the API to unregister one or more error handlers from a parallel computing environment so that they will no longer process one or more errors from one or more accelerators within a heterogeneous processor.In at least one embodiment, the logout error notification callback API call 702 is a set of software instructions that, if executed by one or more processors, cause the one or more processors to use the API to identify, to a parallel computing environment, one or more callback functions or other error handlers that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, in combination. Figure 7 The error handlers described or otherwise referenced are a set of software instructions that will be executed in response to one or more errors (e.g., errors generated by one or more accelerators within a heterogeneous processor).

[0172] In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to indicate one or more instruction sets (e.g., instruction sets of one or more callback functions) that will no longer be executed in response to an event (e.g., an error detected via polling or otherwise). In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to indicate one or more callback functions that will no longer be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to identify one or more error handlers (e.g., callback functions) that no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to indicate one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to unregister one or more callback functions from the parallel computing environment to no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to unregister one or more error handlers from the parallel computing environment to no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702, if invoked by one or more software programs, causes the API to identify one or more callback functions or other error handlers that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0173] In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to indicate one or more instruction sets (e.g., the instruction sets of one or more callback functions) that will no longer be executed in response to an event (e.g., an error detected by polling or otherwise). In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to indicate one or more callback functions that will no longer be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to identify one or more error handlers (e.g., callback functions) that will no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to indicate one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 602 causes or otherwise executes, by one or more circuits in a processor, an API to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to unregister one or more callback functions from a parallel computing environment so that they will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to unregister one or more error handlers from a parallel computing environment so that they will no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes or otherwise executes, by one or more circuits in a processor, an API to identify one or more callback functions or other error handlers that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0174] In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to indicate one or more instruction sets (e.g., the instruction sets of one or more callback functions) that will no longer be executed in response to an event (e.g., an error detected by polling or other means). In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to indicate one or more callback functions that will no longer be executed in response to one or more errors generated by one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to identify one or more error handlers (e.g., callback functions) that will no longer handle one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to indicate one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to indicate to a parallel computing environment (e.g., CUDA) one or more callback functions that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to unregister one or more callback functions from a parallel computing environment so that they will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to unregister one or more error handlers from a parallel computing environment so that they will no longer handle one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the unregister error notification callback API call 702 causes one or more processors in the system to execute an API to identify in a parallel computing environment one or more callback functions or other error handlers that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0175] In at least one embodiment, the unregister error notification callback API call 702 receives, when called, one or more parameters 704, 706 for indicating information about an operation to be performed. In at least one embodiment, the unregister error notification callback API call 702 receives, when called, one or more parameters 704, 706 for indicating information about an instruction to be executed.

[0176] In at least one embodiment, the unregister error notification callback API call 702 receives as input a parameter 704, 706 that includes one or more callback functions 704. In at least one embodiment, the one or more callback functions 704 are data that include information about one or more error handlers that can be used to identify one or more errors that are no longer being processed from one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more callback functions 704 are data that include information about one or more instruction sets that will no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the one or more callback functions 704 are data for indicating one or more error handlers. In at least one embodiment, the one or more callback functions 704 are data for indicating one or more callback functions. In at least one embodiment, the one or more callback functions 704 are data that include one or more memory addresses. In at least one embodiment, the one or more callback functions 704 are data that include one or more pointers. In at least one embodiment, the one or more callback functions 704 are data for indicating any other information about one or more error handlers that are no longer processing one or more errors from one or more accelerators within a heterogeneous processor, such as various data structures further described herein. In at least one embodiment, the one or more callback functions 704 include one or more CUDA pointers or any other type of pointer further described herein.

[0177] In at least one embodiment, the unregister error notification callback API call 702 receives as input arguments 704, 706 that include additional arguments 706. In at least one embodiment, the additional arguments 706 are data that includes any additional information for indicating information that the API is available for in response to the unregister error notification callback API call 702. In at least one embodiment, the additional arguments 706 include information for indicating one or more flags that can be used to configure one or more operations in response to the unregister error notification callback API call 702. In at least one embodiment, the additional arguments 706 include information for indicating any additional information that can be used to perform and / or is configured to perform one or more operations in response to the unregister error notification callback API call 702.

[0178] In at least one embodiment, if the unregister error notification callback API call 702 is invoked, it causes the API 108 to unregister one or more callback functions from the parallel computing environment 106 (e.g., CUDA) such that the one or more callback functions are not executed in response to one or more errors. In at least one embodiment, if the unregister error notification callback API call 702 is invoked, it causes the API 108 to unregister one or more callback functions from the parallel computing environment 106 (e.g., CUDA) such that the one or more callback functions are not executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, if the unregister error notification callback API call 702 is invoked, it causes the API 108 to unregister one or more error handlers from the parallel computing environment 106 (e.g., CUDA) such that the one or more error handlers are not executed in response to one or more errors. In at least one embodiment, if the unregister error notification callback API call 702 is invoked, it causes the API 108 to register one or more error handlers with the parallel computing environment 106 (e.g., CUDA) such that the one or more error handlers are not executed in response to one or more errors from one or more accelerators within a heterogeneous processor.

[0179] In at least one embodiment, in response to a logout error notification callback API call 702, API 108, if executed, causes one or more processors to perform a logout error notification callback API return 708. In at least one embodiment, the logout error notification callback API return 708 is a set of instructions that, if executed, generates and / or indicates one or more data values in response to the logout error notification callback API call 702. In at least one embodiment, the logout error notification callback API return 708 indicates a success identifier 710. In at least one embodiment, the success identifier 710 is data including any value used to indicate the success or successful operation of the logout error notification callback API call 702. In at least one embodiment, the logout error notification callback API return 708 indicates an error identifier 712. In at least one embodiment, the error identifier 712 is data including any value used to indicate the failure or failed operation of the logout error notification callback API call 702. In at least one embodiment, the error identifier 712 includes information indicating one or more specific types of errors generated as a result of or in response to the logout error notification callback API call 702. In at least one embodiment, the error identifier 712 includes information indicating one or more other data values generated in response to or as a result of the logout error notification callback API call 702.

[0180] In at least one embodiment, API 108 includes instructions that, if executed, cause one or more operations or instructions to no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions that, if executed, logout one or more callback functions to no longer be executed in response to one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, API 108 includes instructions that, if executed, identify one or more error handlers that no longer process one or more errors from one or more accelerators within a heterogeneous processor. In at least one embodiment, the instructions for causing one or more operations or instructions to logout or be identified as no longer processing one or more errors from one or more accelerators within a heterogeneous processor to a parallel computing environment 106 (such as CUDA) will be executed in response to the logout error notification callback API call 702, as described above. In at least one embodiment, an example software code indicating the logout error notification callback API call 702 in a parallel computing environment 106 (such as CUDA) is as follows:

[0181]

[0182] Figure 8 Process 800 shows the execution of one or more application programming interfaces (APIs) of one or more accelerators within a heterogeneous processor by a parallel computing environment in accordance with at least one embodiment. In at least one embodiment, process 800 begins at 802 where one or more processors execute a software program 804 that includes one or more instructions that, if executed, cause the one or more processors and / or one or more other processors (e.g., a graphics processing unit (GPU) and / or one or more accelerators within one or more heterogeneous processors) to perform one or more computing operations. In at least one embodiment, the software program 804 to be executed by one or more processors includes one or more instructions that, if executed, cause one or more APIs 108 of the parallel computing environment 106 to be executed, as described above in connection with Figures 1 to 7 that which is described.

[0183] In at least one embodiment, when executing the software program 804, one or more processors execute one or more instructions to cause the one or more processors to execute one or more APIs 806 or API calls 806. In at least one embodiment, if one or more processors do not execute one or more APIs or API calls 806, then process 800 determines whether the execution 804 of one or more instructions of the software program has been completed 808, as described below. In at least one embodiment, if one or more processors execute one or more API calls 806, then the one or more processors determine whether the one or more API calls are stream API calls 812, such as the stream operation API calls described above in connection with Figure 2 that which is described.

[0184] In at least one embodiment, if one or more processors determine that one or more API calls are stream API calls 812 in response to executing one or more instructions, then the one or more processors execute one or more instructions to cause the one or more processors and / or one or more other processors (e.g., a GPU and / or an accelerator within a heterogeneous processor) to execute 814 one or more stream APIs, as described above in connection with Figure 2 that which is described. In at least one embodiment, if one or more processors determine that one or more API calls are not stream API calls 812 in response to executing one or more instructions, then the one or more processors execute one or more instructions that, if executed, cause the one or more processors to determine whether the one or more API calls are memory API calls 816, such as the memory operation API calls described above in connection with Figure 3 that which is described.

[0185] In at least one embodiment, if one or more processors determine, in response to executing one or more instructions, that one or more API calls are memory API calls 816, then the one or more processors execute one or more instructions to cause the one or more processors and / or one or more other processors (e.g., GPUs and / or accelerators within a heterogeneous processor) to execute 818 one or more memory APIs, as described above in connection with Figure 3 that. In at least one embodiment, if one or more processors determine, in response to executing one or more instructions, that one or more API calls are not memory API calls 816, then the one or more processors execute one or more instructions that, if executed, cause the one or more processors to determine whether one or more API calls are error API calls 820, e.g., such as the error notification buffer API calls as described above in connection with Figure 4 and Figure 5 that, and / or such as the error notification callback API calls as described above in connection with Figure 6 and Figure 7 that.

[0186] In at least one embodiment, if one or more processors determine, in response to executing one or more instructions, that one or more API calls are error API calls 820, then the one or more processors execute one or more instructions to cause the one or more processors and / or one or more other processors (e.g., GPUs and / or accelerators within a heterogeneous processor) to execute 822 one or more error APIs, as described above in connection with Figure 6 and Figure 7 that. In at least one embodiment, if one or more processors determine, in response to executing one or more instructions, that one or more API calls are not error API calls 820, then the one or more processors are used to execute one or more instructions that, if executed, cause the one or more processors to execute one or more other APIs 824, e.g., any API further described herein.

[0187] In at least one embodiment, during process 800 of executing one or more APIs, one or more processors executing software program 804 are configured to determine whether the execution of software program 804 is complete 808. In at least one embodiment, if one or more processors have completed 808 executing one or more instructions of software program 804, then process 800 ends 810. In at least one embodiment, if one or more processors have not completed 808 executing one or more instructions of software program 804, then the one or more processors continue to execute one or more instructions of software program 804 and / or one or more other instructions.

[0188] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one of ordinary skill in the art that the inventive concept may be practiced without one or more of these specific details.

[0189] Data center

[0190] Fig. 9 An example data center 900 according to at least one embodiment is shown. In at least one embodiment, data center 900 includes, but is not limited to, a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0191] In at least one embodiment, as Fig. 9 shown, data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources (“node C.R.”) 916(1)-916(N), where “N” represents any whole positive integer. In at least one embodiment, node C.R. 916(1)-916(N) may include, but is not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (“FPGAs”), data processing units (“DPUs”) in network devices, graphics processors, etc.), memory devices (such as dynamic random access memory), storage devices (such as solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of node C.R. 916(1)-916(N) may be servers having one or more of the above computing resources.

[0192] In at least one embodiment, the grouped computing resources 914 can include separate groupings (not shown) of node C.R.s housed within one or more racks, or numerous racks (also not shown) within data centers at various geographical locations. Separate groupings of node C.R.s within the grouped computing resources 914 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s that include CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0193] In at least one embodiment, the resource coordinator 912 can configure or otherwise control one or more node C.R.s 916(1)-916(N) and / or the grouped computing resources 914. In at least one embodiment, the resource coordinator 912 can include a software design infrastructure (“SDI”) management entity for the data center 900. In at least one embodiment, the resource coordinator 912 can include hardware, software, or some combination thereof.

[0194] In at least one embodiment, as Fig. 9As shown, the framework layer 920 includes, but is not limited to, a job scheduler 932, a configuration manager 934, a resource manager 936, and a distributed file system 938. In at least one embodiment, the framework layer 920 may include a framework that supports software 952 of the software layer 930 and / or one or more applications 942 of the application layer 940. In at least one embodiment, the software 952 or the application 942 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 920 may be, but is not limited to, a free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 938 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 932 may include a Spark driver to facilitate scheduling of the workloads supported by the various layers of the data center 900. In at least one embodiment, the configuration manager 934 may be capable of configuring different layers, such as the software layer 930 and the framework layer 920 including Spark and the distributed file system 938 for supporting large-scale data processing. In at least one embodiment, the resource manager 936 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 938 and the job scheduler 932. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 914 on the data center infrastructure layer 910. In at least one embodiment, the resource manager 936 may coordinate with the resource coordinator 912 to manage these mapped or allocated computing resources.

[0195] In at least one embodiment, the software 952 included in the software layer 930 may include software used by at least a portion of nodes C.R. 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0196] In at least one embodiment, one or more applications 942 included in the application layer 940 may include one or more types of applications used by at least a portion of nodes C.R. 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of applications may include, but are not limited to, CUDA applications.

[0197] In at least one embodiment, any one of the configuration manager 934, the resource manager 936, and the resource coordinator 912 can implement any number and type of self-modifying actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions can relieve the data center operator of the data center 900 from making potentially bad configuration decisions and can avoid underutilization and / or poorly performing parts of the data center.

[0198] Computer-based system

[0199] The following figures present, but are not limited to, exemplary computer-based systems that can be used to implement at least one embodiment.

[0200] Fig.10 A processing system 1000 according to at least one embodiment is shown. In at least one embodiment, the system 1000 includes one or more processors 1002 and one or more graphics processors 1008, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1002 or processor cores 1007. In at least one embodiment, the processing system 1000 is a processing platform integrated within a system-on-chip (SoC) integrated circuit for mobile, handheld, or embedded devices.

[0201] In at least one embodiment, the processing system 1000 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in conjunction with Figure 2-7 and / or processes described above in conjunction with Figure 8 In at least one embodiment, the processing system 1000 includes hardware that includes accelerators and / or other components within a heterogeneous processor for performing various computing operations and / or APIs and / or processes described above in conjunction with Figure 2-7 and / or processes described above in conjunction with Figure 8

[0202] In at least one embodiment, the processing system 1000 can be included in or incorporated into a server-based gaming platform, including a game console such as a game and media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, the processing system 1000 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1000 can also be coupled to or integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1000 is a television or a set-top box device having one or more processors 1002 and a graphical interface generated by one or more graphics processors 1008.​

[0203] In at least one embodiment, each of one or more processors 1002 includes one or more processor cores 1007 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 1007 is configured to process a particular instruction set 1009. In at least one embodiment, the instruction set 1009 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, multiple processor cores 1007 can each process a different instruction set 1009, which can include instructions that help to emulate other instruction sets. In at least one embodiment, the processor core 1007 can further include other processing devices, such as a digital signal processor (DSP).

[0204] In at least one embodiment, the processor 1002 includes a cache memory 1004. In at least one embodiment, the processor 1002 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of the processor 1002. In at least one embodiment, the processor 1002 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which can share the logic among the processor cores 1007 using known cache coherence techniques. In at least one embodiment, the processor 1002 further includes a register file 1006, and the processor 1002 can include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 1006 can include general-purpose registers or other registers.

[0205] In at least one embodiment, one or more processors 1002 are coupled to one or more interface buses 1010 to transfer communication signals, such as address, data, or control signals, between the processors 1002 and other components in the system 1000. In at least one embodiment, the interface bus 1010 can be a processor bus, such as a version of the Direct Media Interface (DMI) bus, in one embodiment. In at least one embodiment, the interface bus 1010 is not limited to the DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1002 includes an integrated memory controller 1016 and a Platform Controller Hub 1030. In at least one embodiment, the memory controller 1016 facilitates communication between the storage device and other components of the processing system 1000, while the Platform Controller Hub (PCH) 1030 provides connections to input / output (I / O) devices via a local I / O bus.

[0206] In at least one embodiment, the storage device 1020 can be a Dynamic Random Access Memory (DRAM) device, a Static Random Access Memory (SRAM) device, a flash memory device, a Phase Change Memory device, or have suitable performance to be used as processor memory. In at least one embodiment, the storage device 1020 can be used as the system memory of the processing system 1000 to store data 1022 and instructions 1021 for use when one or more processors 1002 execute an application or process. In at least one embodiment, the memory controller 1016 is also coupled to an optional external graphics processor 1012, which can communicate with one or more graphics processors 1008 in the processor 1002 to perform graphics and media operations. In at least one embodiment, a display device 1011 can be connected to the processor 1002. In at least one embodiment, the display device 1011 can include one or more of an internal display device, such as in a mobile electronic device or a portable computer device, or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1011 can include a Head-Mounted Display (HMD), such as a stereoscopic display device for Virtual Reality (VR) applications or Augmented Reality (AR) applications.

[0207] In at least one embodiment, the platform controller hub 1030 enables peripheral devices to be connected to the storage device 1020 and the processor 1002 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1046, a network controller 1034, a firmware interface 1028, a wireless transceiver 1026, a touch sensor 1025, and a data storage device 1024 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1024 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1025 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1026 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1028 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1034 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1010. In at least one embodiment, the audio controller 1046 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1000 includes an optional legacy I / O controller 1040 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the processing system 1000. In at least one embodiment, the platform controller hub 1030 can also be connected to one or more Universal Serial Bus (USB) controllers 1042, which connect input devices, such as a keyboard and mouse 1043 combination, a camera 1044, or other USB input devices.

[0208] In at least one embodiment, instances of the memory controller 1016 and the platform controller hub 1030 can be integrated into a discrete external graphics processor, such as the external graphics processor 1012. In at least one embodiment, the platform controller hub 1030 and / or the storage controller 1016 can be external to one or more processors 1002. For example, in at least one embodiment, the processing system 1000 can include an external storage controller 1016 and a platform controller hub 1030, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 1002.

[0209] Fig.11FIG. 1100 illustrates a computer system 1100 in accordance with at least one embodiment. In at least one embodiment, computer system 1100 may be a system having interconnected devices and components, an SOC, or some combination thereof. In at least one embodiment, computer system 1100 is formed by a processor 1102, which may include execution units for executing instructions. In at least one embodiment, computer system 1100 may include, but is not limited to, components such as processor 1102, which employs execution units including logic for executing algorithms for processing data. In at least one embodiment, computer system 1100 may include a processor such as a processor family, XeonTM, XScaleTM, and / or StrongARMTM Core TM or Nervana TM microprocessor, although other systems (including PCs, engineering workstations, set-top boxes, etc. having other microprocessors) may also be used. In at least one embodiment, computer system 1100 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0210] In at least one embodiment, computer system 1100 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, computer system 1100 includes hardware that includes accelerators and / or other components within a heterogeneous processor for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0211] ​In at least one embodiment, computer system 1100 can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular telephones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors (“DSPs”), system-on-a-chips (“SoCs”), network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system that can execute one or more instructions according to at least one embodiment.

[0212] In at least one embodiment, computer system 1100 can include, but is not limited to, a processor 1102, which can include, but is not limited to, one or more execution units 1108 that can be configured to execute Compute Unified Device Architecture (“CUDA”) (developed by NVIDIA Corporation of Santa Clara, California) programs. In at least one embodiment, a CUDA program is at least a part of a software application written in the CUDA programming language. In at least one embodiment, computer system 1100 is a single-processor desktop or server system. In at least one embodiment, computer system 1100 can be a multi-processor system. In at least one embodiment, processor 1102 can include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1102 can be coupled to a processor bus 1110 that can transfer data signals between processor 1102 and other components in computer system 1100.

[0213] In at least one embodiment, processor 1102 can include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 1104. In at least one embodiment, processor 1102 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory can reside outside of processor 1102. In at least one embodiment, processor 1102 can include a combination of internal and external caches. In at least one embodiment, register file 1106 can store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0214] In at least one embodiment, execution unit 1108, which includes, but is not limited to, logic for performing integer and floating point operations, is also located in processor 1102. Processor 1102 may also include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode for certain macroinstructions. In at least one embodiment, execution unit 1108 may include logic for processing an encapsulated instruction set 1109. In at least one embodiment, by including the encapsulated instruction set 1109 in the instruction set of general-purpose processor 1102, and the associated circuitry for the instructions to be executed, operations used by many multimedia applications can be performed using the encapsulated data in general-purpose processor 1102. In at least one embodiment, operations on the encapsulated data can be performed by using the full width of the processor's data bus to accelerate and more efficiently execute many multimedia applications, which may not require transferring smaller data units on the processor's data bus to perform one or more operations on one data element at a time.

[0215] In at least one embodiment, execution unit 1108 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1100 may include, but is not limited to, memory 1120. In at least one embodiment, memory 1120 may be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage devices. Memory 1120 may store instructions 1119 and / or data 1121 represented by data signals that may be executed by processor 1102.

[0216] In at least one embodiment, the system logic chip may be coupled to a processor bus 1110 and a memory 1120. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 1116, and the processor 1102 may communicate with the MCH 1116 via the processor bus 1110. In at least one embodiment, the MCH 1116 may provide a high-bandwidth memory path 1118 to the memory 1120 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 1116 may initiate data signals among the processor 1102, the memory 1120, and other components in the computer system 1100, and bridge data signals among the processor bus 1110, the memory 1120, and the system I / O 1122. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 1116 may be coupled to the memory 1120 via the high-bandwidth memory path 1118, and the graphics / video card 1112 may be coupled to the MCH 1116 via an Accelerated Graphics Port (“AGP”) interconnect 1114.

[0217] In at least one embodiment, the computer system 1100 may use the system I / O 1122 as a proprietary hub interface bus to couple the MCH 1116 to an I / O controller hub (“ICH”) 1130. In at least one embodiment, the ICH 1130 may provide a direct connection to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 1120, the chipset, and the processor 1102. Examples may include, but are not limited to, an audio controller 1129, a firmware hub (“Flash BIOS”) 1128, a wireless transceiver 1126, a data storage 1124, a legacy I / O controller 1123 including user input 1125 and a keyboard interface, a serial expansion port 1127 (such as USB), and a network controller 1134. The data storage 1124 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash device, or other mass storage devices.

[0218] In at least one embodiment, Fig.11 A system including interconnected hardware devices or “chips” is shown. In at least one embodiment, Fig.11 An exemplary SoC may be shown. In at least one embodiment, Fig.11The devices shown herein can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 1100 are interconnected using Compute Express Link (CXL) interconnects.

[0219] Fig.12 System 1200 is shown in accordance with at least one embodiment. In at least one embodiment, system 1200 is an electronic device utilizing processor 1210. In at least one embodiment, system 1200 can be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, an edge device communicatively coupled to one or more local or cloud service providers, a notebook computer, a desktop computer, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0220] In at least one embodiment, system 1200 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, system 1200 includes hardware that includes accelerators and / or other components within a heterogeneous processor for performing various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0221] In at least one embodiment, system 1200 can include, but is not limited to, processor 1210 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1210 is coupled using a bus or interface such as an I 2 C bus, a system management bus (“SMBus”), a low pin count (LPC) bus, a serial peripheral interface (“SPI”), a high definition audio (“HDA”) bus, a serial advanced technology attachment (“SATA”) bus, a USB (versions 1, 2, 3), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, Fig.12 A system is shown that includes interconnected hardware devices or “chips”. In at least one embodiment, Fig.12 An exemplary SoC can be shown. In at least one embodiment, Fig.12 The devices shown in Fig.12 can be interconnected with proprietary interconnect lines, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment,

[0222] In at least one embodiment,​ Fig.12 It may include a display 1224, a touch screen 1225, a touchpad 1230, a Near Field Communication unit (“NFC”) 1245, a sensor hub 1240, a thermal sensor 1246, a Fast Chipset (“EC”) 1235, a Trusted Platform Module (“TPM”) 1238, a BIOS / Firmware / Flash (“BIOS, FW Flash”) 1222, a DSP 1260, a Solid State Disk (“SSD”) or a Hard Disk Drive (“HDD”) 1220, a Wireless Local Area Network unit (“WLAN”) 1250, a Bluetooth unit 1252, a Wireless Wide Area Network unit (“WWAN”) 1256, a Global Positioning System (GPS) 1255, a camera (“USB 3.0 camera”) 1254 (such as a USB 3.0 camera), or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 1215 implemented in accordance with, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.

[0223] In at least one embodiment, other components may be communicatively coupled to the processor 1210 via the components discussed above. In at least one embodiment, an accelerometer 1241, an Ambient Light Sensor (“ALS”) 1242, a compass 1243, and a gyroscope 1244 may be communicatively coupled to the sensor hub 1240. In at least one embodiment, a thermal sensor 1239, a fan 1237, a keyboard 1236, and a touchpad 1230 may be communicatively coupled to the EC 1235. In at least one embodiment, a speaker 1263, headphones 1264, and a microphone (“mic”) 1265 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 1262, which may in turn be communicatively coupled to the DSP 1260. In at least one embodiment, the audio unit 1262 may include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a Subscriber Identity Module (“SIM”) 1257 may be communicatively coupled to the WWAN unit 1256. In at least one embodiment, components (such as the WLAN unit 1250, the Bluetooth unit 1252, and the WWAN unit 1256) may be implemented in a Next Generation Form Factor (NGFF).

[0224] Fig.13An exemplary integrated circuit 1300 according to at least one embodiment is shown. In at least one embodiment, the exemplary integrated circuit 1300 is a SoC that can be fabricated using one or more IP cores. In at least one embodiment, the integrated circuit 1300 includes one or more application processors 1305 (e.g., CPU, DPU), at least one graphics processor 1310, and may additionally include an image processor 1315 and / or a video processor 1320, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1300 includes peripheral or bus logic, which includes a USB controller 1325, a UART controller 1330, an SPI / SDIO controller 1335, and an 2 S / I 2 C controller 1340. In at least one embodiment, the integrated circuit 1300 may include a display device 1345 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1350 and a mobile industry processor interface (MIPI) display interface 1355. In at least one embodiment, storage may be provided by a flash memory subsystem 1360, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1365 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1370.

[0225] In at least one embodiment, the exemplary integrated circuit 1300 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the exemplary integrated circuit 1300 includes hardware that includes accelerators and / or other components within a heterogeneous processor for performing various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0226] Fig.14 ​FIG. 1400 illustrates a computing system 1400 in accordance with at least one embodiment. In at least one embodiment, the computing system 1400 includes a processing subsystem 1401 having one or more processors 1402 and a system memory 1404 that communicate via an interconnect path that can include a memory hub 1405. In at least one embodiment, the memory hub 1405 can be a separate component within a chipset component or can be integrated within one or more of the processors 1402. In at least one embodiment, the memory hub 1405 is coupled to an I / O subsystem 1411 via a communication link 1406. In at least one embodiment, the I / O subsystem 1411 includes an I / O hub 1407 that can enable the computing system 1400 to receive input from one or more input devices 1408. In at least one embodiment, the I / O hub 1407 can enable a display controller that is included within one or more of the processors 1402 for providing output to one or more display devices 1410A. In at least one embodiment, one or more display devices 1410A coupled to the I / O hub 1407 can include local, internal, or embedded display devices.

[0227] In at least one embodiment, the computing system 1400 is configured to perform various computing operations including one or more of the application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the computing system 1400 includes hardware that includes accelerators and / or other components within a heterogeneous processor for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or the APIs and / or the processes described above in connection with Figure 8

[0228] ​In at least one embodiment, the processing subsystem 1401 includes one or more parallel processors 1412 coupled to the memory hub 1405 via a bus or other communication link 1413. In at least one embodiment, the communication link 1413 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 1412 form a parallel or vector processing system in a computing concentration, which can include a large number of processing cores and / or processing clusters, such as a Many Integrated Core (MIC) processor. In at least one embodiment, the one or more parallel processors 1412 form a graphics processing subsystem that can output pixels to one of the one or more display devices 1410A coupled via the I / O hub 1407. In at least one embodiment, the one or more parallel processors 1412 can also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1410B.

[0229] In at least one embodiment, the system storage unit 1414 can be connected to the I / O hub 1407 to provide a storage mechanism for the computing system 1400. In at least one embodiment, the I / O switch 1416 can be used to provide an interface mechanism to enable connections between the I / O hub 1407 and other components, such as a network adapter 1418 and / or a wireless network adapter 1419 that can be integrated into the platform, and various other devices that can be added via one or more additional devices 1420. In at least one embodiment, the network adapter 1418 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1419 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices including one or more radios.

[0230] In at least one embodiment, the computing system 1400 can include other components not explicitly shown, including USB or other port connections, an optical storage drive, a video capture device, etc., which can also be connected to the I / O hub 1407. In at least one embodiment, Fig.14 the communication paths interconnecting the various components can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols (e.g., NVLink high-speed interconnect or interconnect protocol).

[0231] In at least one embodiment, one or more parallel processors 1412 include circuitry optimized for graphics and video processing (including, for example, video output circuitry) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1412 include circuitry optimized for general-purpose processing. In at least one embodiment, components of the computing system 1400 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1412, memory hub 1405, processor 1402, and I / O hub 1407 may be integrated into a system-on-a-chip (SoC) integrated circuit. In at least one embodiment, components of the computing system 1400 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 1400 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system. In at least one embodiment, the I / O subsystem 1411 and display device 1410B are omitted from the computing system 1400.

[0232] Processing system

[0233] The following figures illustrate, but are not limited to, exemplary processing systems that may be used to implement at least one embodiment.

[0234] Fig.15 An accelerated processing unit (“APU”) 1500 is shown in accordance with at least one embodiment. In at least one embodiment, the APU 1500 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the APU 1500 may be configured to execute applications, such as CUDA programs. In at least one embodiment, the APU 1500 includes, but is not limited to, a core complex 1510, a graphics complex 1540, an infrastructure 1560, an I / O interface 1570, a memory controller 1580, a display controller 1592, and a multimedia engine 1594. In at least one embodiment, the APU 1500 may include, but is not limited to, any combination of any number of core complexes 1510, any number of graphics complexes 1540, any number of display controllers 1592, and any number of multimedia engines 1594. For illustrative purposes, multiple instances of like objects are denoted by reference numerals herein, where the reference numeral identifies the object and the number in parentheses identifies the instance desired.

[0235] In at least one embodiment, the APU 1500 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or one or more application programming interfaces (APIs) described above in connection with Figure 8The described process. In at least one embodiment, the APU 1500 includes hardware and / or other components for performing the various computational operations and / or APIs described above in connection with Figure 2-7 and / or the process described above in connection with Figure 8 .

[0236] In at least one embodiment, the Core Complex 1510 is a CPU, the Graphics Complex 1540 is a GPU, and the APU 1500 is a processing unit that integrates, without limitation, 1510 and 1540 onto a single chip. In at least one embodiment, some tasks can be assigned to the Core Complex 1510 while other tasks can be assigned to the Graphics Complex 1540. In at least one embodiment, the Core Complex 1510 is configured to execute the main control software associated with the APU 1500, such as an operating system. In at least one embodiment, the Core Complex 1510 is the main processor of the APU 1500 that controls and coordinates the operation of other processors. In at least one embodiment, the Core Complex 1510 issues commands that control the operation of the Graphics Complex 1540. In at least one embodiment, the Core Complex 1510 can be configured to execute host-executable code derived from CUDA source code, and the Graphics Complex 1540 can be configured to execute device-executable code derived from CUDA source code.

[0237] In at least one embodiment, the Core Complex 1510 includes, but is not limited to, cores 1520(1)-1520(4) and the L3 cache 1530. In at least one embodiment, the Core Complex 1510 can include, but is not limited to, any number of cores 1520 and any combination of any number and type of caches. In at least one embodiment, the cores 1520 are configured to execute instructions of a specific instruction set architecture (“ISA”). In at least one embodiment, each core 1520 is a CPU core.

[0238] In at least one embodiment, each core 1520 includes, but is not limited to, a fetch / decode unit 1522, an integer execution engine 1524, a floating-point execution engine 1526, and an L2 cache 1528. In at least one embodiment, the fetch / decode unit 1522 fetches instructions, decodes these instructions, generates micro-operations, and dispatches individual micro-instructions to the integer execution engine 1524 and the floating-point execution engine 1526. In at least one embodiment, the fetch / decode unit 1522 can dispatch one micro-instruction to the integer execution engine 1524 and another micro-instruction to the floating-point execution engine 1526 simultaneously. In at least one embodiment, the integer execution engine 1524 performs operations not limited to integer and memory operations. In at least one embodiment, the floating-point engine 1526 performs operations not limited to floating-point and vector operations. In at least one embodiment, the fetch-decode unit 1522 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 1524 and the floating-point execution engine 1526.

[0239] In at least one embodiment, each core 1520(i) can access the L2 cache 1528(i) included in the core 1520(i), where i is an integer representing a particular instance of the core 1520. In at least one embodiment, each core 1520 included in the core complex 1510(j) is connected to other cores 1520 included in the core complex 1510(j) via the L3 cache 1530(j) included in the core complex 1510(j), where j is an integer representing a particular instance of the core complex 1510. In at least one embodiment, the cores 1520 included in the core complex 1510(j) can access all of the L3 caches 1530(j) included in the core complex 1510(j), where j is an integer representing a particular instance of the core complex 1510. In at least one embodiment, the L3 cache 1530 can include, but is not limited to, any number of slices.

[0240] In at least one embodiment, the graphics complex 1540 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, the graphics complex 1540 is configured to perform graphics pipeline operations such as draw commands, pixel operations, geometric calculations, and other operations associated with rendering an image to a display. In at least one embodiment, the graphics complex 1540 is configured to perform operations unrelated to graphics. In at least one embodiment, the graphics complex 1540 is configured to perform operations related to graphics and operations unrelated to graphics.

[0241] In at least one embodiment, the graphics complex 1540 includes, but is not limited to, any number of compute units 1550 and an L2 cache 1542. In at least one embodiment, the compute units 1550 share the L2 cache 1542. In at least one embodiment, the L2 cache 1542 is partitioned. In at least one embodiment, the graphics complex 1540 includes, but is not limited to, any number of compute units 1550 and any number (including zero) and type of caches. In at least one embodiment, the graphics complex 1540 includes, but is not limited to, any number of dedicated graphics hardware.

[0242] In at least one embodiment, each compute unit 1550 includes, but is not limited to, any number of SIMD units 1552 and a shared memory 1554. In at least one embodiment, each SIMD unit 1552 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 1550 can execute any number of thread blocks, but each thread block is executed on a single compute unit 1550. In at least one embodiment, a thread block includes, but is not limited to, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 1552 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predication can be used to disable one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block can synchronize together and communicate via the shared memory 1554.

[0243] In at least one embodiment, the fabric 1560 is a system interconnect that facilitates data and control transfers across the core complex 1510, graphics complex 1540, I / O interface 1570, memory controller 1580, display controller 1592, and multimedia engine 1594. In at least one embodiment, in addition to or instead of the fabric 1560, the APU 1500 may also include, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of components that are directly or indirectly linked, either internal or external to the APU 1500. In at least one embodiment, the I / O interface 1570 represents any number and type of I / O interfaces (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to the I / O interface 1570. In at least one embodiment, the peripheral devices coupled to the I / O interface 1570 may include, but are not limited to, a keyboard, a mouse, a printer, a scanner, a joystick, or other types of game controllers, media recording devices, external storage devices, network interface cards, etc.

[0244] In at least one embodiment, the display controller AMD92 displays images on one or more display devices (e.g., liquid crystal display (LCD) devices). In at least one embodiment, the multimedia engine 1594 includes, but is not limited to, any number and type of multimedia-related circuitry, such as video decoders, video encoders, image signal processors, etc. In at least one embodiment, the memory controller 1580 facilitates data transfers between the APU 1500 and the unified system memory 1590. In at least one embodiment, the core complex 1510 and the graphics complex 1540 share the unified system memory 1590.

[0245] In at least one embodiment, the APU 1500 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 1580 and memory devices (e.g., shared memory 1554) that may be dedicated to one component or shared among multiple components. In at least one embodiment, the APU 1500 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 1628, L3 cache 1530, and L2 cache 1542), each of which may be private to a component or shared among any number of components (e.g., cores 1520, core complex 1510, SIMD units 1552, compute units 1550, and graphics complex 1540).

[0246] Fig.16Shows a CPU 1600 according to at least one embodiment. In at least one embodiment, the CPU 1600 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the CPU 1600 can be configured to execute application programs. In at least one embodiment, the CPU 1600 is configured to execute main control software, such as an operating system. In at least one embodiment, the CPU 1600 issues commands to control the operation of an external GPU (not shown). In at least one embodiment, the CPU 1600 can be configured to execute host-executable code derived from CUDA source code, and the external GPU can be configured to execute device-executable code derived from such CUDA source code. In at least one embodiment, the CPU 1600 includes, but is not limited to, any number of core complexes 1610, a structure 1660, an I / O interface 1670, and a memory controller 1680.

[0247] In at least one embodiment, the CPU 1600 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the CPU 1600 includes hardware and / or other components for performing various computing operations and / or APIs described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0248] In at least one embodiment, the core complex 1610 includes, but is not limited to, cores 1620(1)-1620(4) and an L3 cache 1630. In at least one embodiment, the core complex 1610 can include, but is not limited to, any number of cores 1620 and any combination of any number and type of caches. In at least one embodiment, the cores 1620 are configured to execute instructions of a specific ISA. In at least one embodiment, each core 1620 is a CPU core.

[0249] ​In at least one embodiment, each core 1620 includes, but is not limited to, an instruction fetch / decode unit 1622, an integer execution engine 1624, a floating-point execution engine 1626, and an L2 cache 1628. In at least one embodiment, the instruction fetch / decode unit 1622 fetches instructions, decodes these instructions, generates micro-operations, and dispatches individual micro-instructions to the integer execution engine 1624 and the floating-point execution engine 1626. In at least one embodiment, the instruction fetch / decode unit 1622 can dispatch one micro-instruction to the integer execution engine 1624 and another micro-instruction to the floating-point execution engine 1626 simultaneously. In at least one embodiment, the integer execution engine 1624 performs operations not limited to integer and memory operations. In at least one embodiment, the floating-point engine 1626 performs operations not limited to floating-point and vector operations. In at least one embodiment, the instruction fetch / decode unit 1622 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 1624 and the floating-point execution engine 1626.

[0250] In at least one embodiment, each core 1620(i) can access the L2 cache 1628(i) included in the core 1620(i), where i is an integer representing a specific instance of the core 1620. In at least one embodiment, each core 1620 included in a core complex 1610(j) is connected to other cores 1620 in the core complex 1610(j) via the L3 cache 1630(j) included in the core complex 1610(j), where j is an integer representing a specific instance of the core complex 1610. In at least one embodiment, the cores 1620 included in a core complex 1610(j) can access all of the L3 caches 1630(j) included in the core complex 1610(j), where j is an integer representing a specific instance of the core complex 1610. In at least one embodiment, the L3 cache 1630 can include, but is not limited to, any number of slices.

[0251] In at least one embodiment, the structure 1660 is a system interconnect that facilitates data and control transfers across the core complexes 1610(1)-1610(N) (where N is an integer greater than zero), the I / O interfaces 1670, and the memory controller 1680. In at least one embodiment, in addition to or instead of the structure 1660, the CPU 1600 may also include, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to the CPU 1600. In at least one embodiment, the I / O interface 1670 represents any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to the I / O interface 1670. In at least one embodiment, the peripheral devices coupled to the I / O interface 1670 may include, but are not limited to, a display, a keyboard, a mouse, a printer, a scanner, a joystick or other types of game controllers, a media recording device, an external storage device, a network interface card, etc.

[0252] In at least one embodiment, the memory controller 1680 facilitates data transfers between the CPU 1600 and the system memory 1690. In at least one embodiment, the core complex 1610 and the graphics complex 1640 share the system memory 1690. In at least one embodiment, the CPU 1600 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 1680 and memory devices that may be dedicated to one component or shared among multiple components. In at least one embodiment, the CPU 1600 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 1628 and L3 cache 1630), each of which may be private to a component or shared among any number of components (e.g., cores 1620 and core complex 1610).

[0253] Fig.17An exemplary accelerator integrated slice 1790 according to at least one embodiment is shown. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, environment management, and interrupt management services on behalf of multiple graphics processing engines in multiple graphics acceleration modules. The graphics processing engines may each include a separate GPU. Optionally, the graphics processing engine may include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module may be a GPU with multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a general package, line card, or chip.

[0254] In at least one embodiment, the accelerator integrated slice 1790 is used to perform various computing operations, including the above combined Figure 2-7 One or more application programming interfaces (APIs) described above and / or in combination with Figure 8 In at least one embodiment, the accelerator integrated slice 1790 includes a method for performing the above combined Figure 2-7 Various computing operations and / or APIs described above and / or combinations thereof Figure 8 The hardware and / or other components of the processes described.

[0255] The application effective address space 1782 within the system memory 1714 stores process elements 1783. In one embodiment, the process element 1783 is stored in response to a GPU call 1781 from an application 1780 executing on the processor 1707. The process element 1783 contains the processing state of the corresponding application 1780. The work descriptor (WD) 1784 contained in the process element 1783 can be a single job requested by the application or may contain a pointer to a job queue. In at least one embodiment, the WD 1784 is a pointer to a job request queue in the application effective address space 1782.

[0256] Graphics acceleration module 1746 and / or each graphics processing engine can be shared by all or part of the processes in the system. In at least one embodiment, an infrastructure for establishing a processing state and sending WD 1784 to graphics acceleration module 1746 to start a job in a virtualized environment can be included.

[0257] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns the graphics acceleration module 1746 or an individual graphics processing engine. Since the graphics acceleration module 1746 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owning partition, and the operating system initializes the accelerator integrated circuit for the owning partition when the graphics acceleration module 1746 is allocated.

[0258] In operation, the WD fetch unit 1791 in the accelerator integrated slice 1790 fetches the next WD 1784, which includes an indication of work to be done by one or more graphics processing engines of the graphics acceleration module 1746. Data from the WD 1784 can be stored in the register 1745 and used by the memory management unit (MMU) 1739, the interrupt management circuit 1747, and / or the context management circuit 1748, as shown. For example, one embodiment of the MMU 1739 includes a segment / page walk circuit for accessing the segment / page table 1786 within the OS virtual address space 1785. The interrupt management circuit 1747 can process interrupt events (INT) 1792 received from the graphics acceleration module 1746. When performing a graphics operation, the virtual address 1793 generated by the graphics processing engine is translated to a physical address by the MMU 1739.

[0259] In one embodiment, the same register set 1745 is replicated for each graphics processing engine and / or graphics acceleration module 1746 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integrated slice 1790. Exemplary registers that can be initialized by the hypervisor are shown in Table 1.

[0260] Table 1 - Registers Initialized by the Hypervisor

[0261] 1 Slice Control Register 2 Real address (RA) plan processing area pointer 3 Permission shield override register 4 Interrupt vector table input offset 5 Interrupt vector table entry restriction 6 Status Register 7 Logical Partition ID 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Description Register

[0262] Exemplary registers that can be initialized by the operating system are shown in Table 2.

[0263] Table 2 - Registers Initialized by the Operating System

[0264] 1 Process and thread identification 2 Effective Address (EA) environment save / restore pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) stores the segment table pointer 5 Permission shielding 6 Job Descriptor

[0265] In one embodiment, each WD 1784 is specific to a particular graphics acceleration module 1746 and / or a particular graphics processing engine. It contains all the information required for the graphics processing engine to perform the work or the work to be done, or it can be a pointer to a memory location where the application has established a command queue of work to be completed.

[0266] Figures 18A-18BAn exemplary graphics processor in accordance with at least one embodiment herein is shown. In at least one embodiment, any exemplary graphics processor may be fabricated using one or more IP cores. In addition to the illustration, in at least one embodiment other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general processor cores. In at least one embodiment, the exemplary graphics processor is for use within a SoC.

[0267] Fig.18A An exemplary graphics processor 1810 of a SoC integrated circuit in accordance with at least one embodiment is shown, which may be fabricated using one or more IP cores. Fig.18B An additional exemplary graphics processor 1840 of a SoC integrated circuit in accordance with at least one embodiment is shown, which may be fabricated using one or more IP cores. In at least one embodiment, Fig.18A the graphics processor 1810 is a low-power graphics processor core. In at least one embodiment, Fig.18B the graphics processor 1840 is a higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 1810, 1840 may be Fig.13 a variant of the graphics processor 1310.

[0268] In at least one embodiment, the graphics processor 1810 is for performing various computing operations, including one or more of the application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the graphics processor 1810 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8

[0269] In at least one embodiment, the graphics processor 1810 includes a vertex processor 1805 and one or more fragment processors 1815A - 1815N (e.g., 1815A, 1815B, 1815C, 1815D to 1815N - 1, and 1815N). In at least one embodiment, the graphics processor 1810 may execute different shader programs via separate logic such that the vertex processor 1805 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1815A - 1815N perform fragment (e.g., pixel) shading operations for fragment or pixel or shader programs. In at least one embodiment, the vertex processor 1805 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1815A - 1815N use the primitives and vertex data generated by the vertex processor 1805 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 1815A - 1815N are optimized to execute fragment shader programs as provided in the OpenGL API, which may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.

[0270] In at least one embodiment, the graphics processor 1810 additionally includes one or more MMUs 1820A - 1820B, caches 1825A - 1825B, and circuit interconnects 1830A - 1830B. In at least one embodiment, the one or more MMUs 1820A - 1820B provide virtual - to - physical address mapping for the graphics processor 1810, including for the vertex processor 1805 and / or the fragment processors 1815A - 1815N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in the one or more caches 1825A - 1825B. In at least one embodiment, the one or more MMUs 1820A - 1820B may be synchronized with other MMUs within the system, including one or more MMUs associated with one or more application processors 1305, image processors 1315, and / or video processors 1320 of Fig.13 such that each processor 1305 - 1320 may participate in a shared or unified virtual memory system. In at least one embodiment, the one or more circuit interconnects 1830A - 1830B enable the graphics processor 1810 to connect to other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0271] In at least one embodiment, the graphics processor 1840 includes Fig.18AOne or more MMUs 1820A - 1820B, caches 1825A - 1825B, and circuit interconnects 1830A - 1830B of the graphics processor 1810. In at least one embodiment, the graphics processor 1840 includes one or more shader cores 1855A - 1855N (e.g., 1855A, 1855B, 1855C, 1855D, 1855E, 1855F, up to 1855N - 1 and 1855N), which provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1840 includes an inter - core task manager 1845, which acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1855A - 1855N and a tiling unit 1858 to accelerate tiling operations for tile - based rendering, where the rendering operation of the scene is subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or optimize the use of internal caches.

[0272] Fig.19A Illustrates a graphics core 1900 according to at least one embodiment. In at least one embodiment, the graphics core 1900 can be included within Fig.13 the graphics processor 1310. In at least one embodiment, the graphics core 1900 can be Fig.18B the unified shader cores 1855A - 1855N in. In at least one embodiment, the graphics core 1900 includes a shared instruction cache 1902, texture units 1918, and cache / shared memory 1920, which are shared by the execution resources within the graphics core 1900. In at least one embodiment, the graphics core 1900 can include multiple slices 1901A - 1901N or partitions per core, and the graphics processor can include multiple instances of the graphics core 1900. The slices 1901A - 1901N can include support logic, which includes local instruction caches 1904A - 1904N, thread schedulers 1906A - 1906N, thread dispatchers 1908A - 1908N, and a set of registers 1910A - 1910N. In at least one embodiment, the slices 1901A - 1901N can include a set of additional functional units (AFU) 1912A - 1912N, floating - point units (FPU) 1914A - 1914N, integer arithmetic logic units (ALU) 1916A - 1916N, address calculation units (ACU) 1913A - 1913N, double - precision floating - point units (DPFPU)

[0273] 1915A - 1915N and matrix processing units (MPU) 1917A - 1917N.

[0274] In at least one embodiment, the graphics core 1900 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the graphics core 1900 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 described above.

[0275] In one embodiment, the FPU 1914A - 1914N can perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, while the DPFPU 1915A - 1915N can perform double - precision (64 - bit) floating - point operations. In at least one embodiment, the ALU 1916A - 1916N can perform variable - precision integer operations with 8 - bit, 16 - bit, and 32 - bit precision and can be configured for mixed - precision operations. In at least one embodiment, the MPU 1917A - 1917N can also be configured for mixed - precision matrix operations, including half - precision floating - point operations and 8 - bit integer operations. In at least one embodiment, the MPU 1917A - 1917N can perform various matrix operations to accelerate CUDA programs, including enabling accelerated general matrix - to - matrix multiplication (GEMM). In at least one embodiment, the AFU 1912A - 1912N can perform additional logical operations not supported by the floating - point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0276] Fig.19BShows a general - purpose graphics processing unit (GPGPU) 1930 in at least one embodiment. In at least one embodiment, the GPGPU 1930 is highly parallel and suitable for deployment on a multi - chip module. In at least one embodiment, the GPGPU 1930 can be configured such that highly parallel computing operations can be performed by a GPU array. In at least one embodiment, the GPGPU 1930 can be directly linked to other instances of the GPGPU 1930 to create a multi - GPU cluster to improve the execution time for CUDA programs. In at least one embodiment, the GPGPU 1930 includes a host interface 1932 to enable connection to a host processor. In at least one embodiment, the host interface 1932 is a PCIe interface. In at least one embodiment, the host interface 1932 can be a vendor - specific communication interface or communication fabric. In at least one embodiment, the GPGPU 1930 receives commands from the host processor and uses a global scheduler 1934 to dispatch execution threads associated with those commands to a set of compute clusters 1936A - 1936H. In at least one embodiment, the compute clusters 1936A - 1936H share a cache memory 1938. In at least one embodiment, the cache memory 1938 can be used as a cache - of - caches for the cache memories within the compute clusters 1936A - 1936H.

[0277] In at least one embodiment, the GPGPU 1930 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the GPGPU 1930 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0278] In at least one embodiment, the GPGPU 1930 includes memories 1944A - 1944B coupled to the compute clusters 1936A - 1936H via a set of memory controllers 1942A - 1942B. In at least one embodiment, the memories 1944A - 1944B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double - data - rate (GDDR) memory.

[0279] In at least one embodiment, each of the compute clusters 1936A - 1936H includes a set of graphics cores, such as Fig.19A ​The graphics core 1900, which may include various types of integer and floating - point logic units, can perform computational operations with various precisions, including those suitable for computations related to CUDA programs. For example, in at least one embodiment, at least one subset of the floating - point units in each of the compute clusters 1936A - 1936H can be configured to perform 16 - bit or 32 - bit floating - point operations, while different subsets of floating - point units can be configured to perform 64 - bit floating - point operations.

[0280] In at least one embodiment, multiple instances of the GPGPU 1930 can be configured to operate as compute clusters. The compute clusters 1936A - 1936H can implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of the GPGPU 1930 communicate via the host interface 1932. In at least one embodiment, the GPGPU 1930 includes an I / O hub 1939 that couples the GPGPU 1930 to the GPU link 1940, enabling direct connection to other instances of the GPGPU 1930. In at least one embodiment, the GPU link 1940 is coupled to a dedicated GPU - to - GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 1930. In at least one embodiment, the GPU link 1940 is coupled to a high - speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1930 are located in separate data - processing systems and communicate via network devices accessible via the host interface 1932. In at least one embodiment, the GPU link 1940 can be configured to be able to connect to a host processor, in addition to or in place of the host interface 1932. In at least one embodiment, the GPGPU 1930 can be configured to execute CUDA programs.

[0281] Fig. 20A A parallel processor 2000 is shown in accordance with at least one embodiment. In at least one embodiment, various components of the parallel processor 2000 can be implemented using one or more integrated - circuit devices, such as programmable processors, application - specific integrated circuits (ASICs), or field - programmable gate arrays (FPGAs).

[0282] In at least one embodiment, the parallel processor 2000 is used to perform various computational operations, including one or more of the application programming interfaces (APIs) described above in connection with Figure 2-7 and / or one or more of the processes described above in connection with Figure 8 In at least one embodiment, the parallel processor 2000 includes hardware and / or other components for performing the various computational operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or one or more of the processes described above in connection with Figure 8

[0283] In at least one embodiment, the parallel processor 2000 includes a parallel processing unit 2002. In at least one embodiment, the parallel processing unit 2002 includes an I / O unit 2004 that enables communication with other devices, including other instances of the parallel processing unit 2002. In at least one embodiment, the I / O unit 2004 can be directly connected to other devices. In at least one embodiment, the I / O unit 2004 is connected to other devices by using a hub or switch interface (e.g., a memory hub 2005). In at least one embodiment, the connection between the memory hub 2005 and the I / O unit 2004 forms a communication link. In at least one embodiment, the I / O unit 2004 is connected to a host interface 2006 and a memory crossbar 2016, where the host interface 2006 receives commands for performing processing operations, and the memory crossbar 2016 receives commands for performing memory operations.

[0284] In at least one embodiment, when the host interface 2006 receives a command buffer via the I / O unit 2004, the host interface 2006 can initiate a work operation to execute those commands to a front end 2008. In at least one embodiment, the front end 2008 is coupled to a scheduler 2010 configured to allocate commands or other work items to a processing array 2012. In at least one embodiment, the scheduler 2010 ensures that the processing array 2012 is correctly configured and in an active state before tasks are assigned to the processing array 2012 in the processing array 2012. In at least one embodiment, the scheduler 2010 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2010 can be configured to perform complex scheduling and work allocation operations at both a coarse-grained and a fine-grained level, enabling fast preemption and context switching of threads executing on the processing array 2012. In at least one embodiment, host software can attest to a workload for scheduling on the processing array 2012 via one of a plurality of graphics processing doorbells. In at least one embodiment, the workload can then be automatically allocated on the processing array 2012 by scheduler 2010 logic within a microcontroller including the scheduler 2010.

[0285] In at least one embodiment, the processing array 2012 may include up to "N" processing clusters (e.g., cluster 2014A, cluster 2014B to cluster 2014N). In at least one embodiment, each of the clusters 2014A - 2014N of the processing array 2012 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 2010 may use various scheduling and / or workload assignment algorithms to assign work to the clusters 2014A - 2014N of the processing array 2012, which may vary according to the workload generated by each type of program or computation. In at least one embodiment, the scheduling may be handled dynamically by the scheduler 2010, or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by the processing array 2012. In at least one embodiment, different clusters 2014A - 2014N of the processing array 2012 may be assigned to process different types of programs or to perform different types of computations.

[0286] In at least one embodiment, the processing array 2012 may be configured to perform various types of parallel processing operations. In at least one embodiment, the processing array 2012 is configured to perform general - purpose parallel computing operations. For example, in at least one embodiment, the processing array 2012 may include logic for performing processing tasks that include filtering of video and / or audio data, performing modeling operations, including physical operations, and performing data transformation.

[0287] In at least one embodiment, the processing array 2012 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing array 2012 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. In at least one embodiment, the processing array 2012 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2002 may transfer data from the system memory via the I / O unit 2004 for processing. In at least one embodiment, during processing, the transferred data may be stored in on - chip memory (e.g., parallel processor memory 2022) during processing and then written back to the system memory.

[0288] In at least one embodiment, when the parallel processing unit 2002 is used to perform graphics processing, the scheduler 2010 can be configured to divide the processing workload into tasks of approximately equal size to better distribute the graphics processing operations to the multiple clusters 2014A - 2014N of the processing array 2012. In at least one embodiment, portions of the processing array 2012 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to generate a rendered image for display. In at least one embodiment, the intermediate data generated by one or more of the clusters 2014A - 2014N can be stored in a buffer to allow the transfer of the intermediate data between the clusters 2014A - 2014N for further processing.

[0289] In at least one embodiment, the processing array 2012 can receive processing tasks to be executed via the scheduler 2010, which receives commands defining the processing tasks from the front end 2008. In at least one embodiment, the processing tasks can include an index of the data to be processed, such as can include surface (patch) data, primitive data, vertex data, and / or pixel data, as well as status parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 2010 can be configured to obtain the index corresponding to the task, or can receive the index from the front end 2008. In at least one embodiment, the front end 2008 can be configured to ensure that the processing array 2012 is configured in a valid state before starting the workload specified by the incoming command buffer (e.g., batch - buffer, push buffer, etc.).

[0290] In at least one embodiment, each of one or more instances of the parallel processing unit 2002 can be coupled to the parallel processor memory 2022. In at least one embodiment, the parallel processor memory 2022 can be accessed via the memory crossbar 2016, which can receive memory requests from the processing array 2012 as well as the I / O unit 2004. In at least one embodiment, the memory crossbar 2016 can access the parallel processor memory 2022 via the memory interface 2018. In at least one embodiment, the memory interface 2018 can include a plurality of partitioning units (e.g., partitioning unit 2020A, partitioning unit 2020B to partitioning unit 2020N), each of which can be coupled to a portion (e.g., a memory unit) of the parallel processor memory 2022. In at least one embodiment, the plurality of partitioning units 2020A - 2020N are configured to be equal to the number of memory units such that the first partitioning unit 2020A has a corresponding first memory unit 2024A, the second partitioning unit 2020B has a corresponding memory unit 2024B, and the Nth partitioning unit 2020N has a corresponding Nth memory unit 2024N. In at least one embodiment, the number of partitioning units 2020A - 2020N may not be equal to the number of memory devices.

[0291] In at least one embodiment, the memory units 2024A - 2024N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 2024A - 2024N can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across the memory units 2024A - 2024N, allowing the partitioning units 2020A - 2020N to write portions of each rendering target in parallel to effectively utilize the available bandwidth of the parallel processor memory 2022. In at least one embodiment, a local instance of the parallel processor memory 2022 can be excluded in favor of a unified memory design that utilizes system memory in combination with local cache memory.

[0292] In at least one embodiment, any one of clusters 2014A - 2014N of processing array 2012 can process data to be written into any of memory cells 2024A - 2024N within parallel processor memory 2022. In at least one embodiment, memory crossbar 2016 can be configured to transfer the output of each cluster 2014A - 2014N to any partition unit 2020A - 2020N or another cluster 2014A - 2014N, and the cluster 2014A - 2014N can perform other processing operations on the output. In at least one embodiment, each cluster 2014A - 2014N can communicate with memory interface 2018 through memory crossbar 2016 to read from or write to various external storage devices. In at least one embodiment, memory crossbar 2016 has a connection to memory interface 2018 to communicate with I / O unit 2004, and a connection to a local instance of parallel processor memory 2022, enabling processing units within different processing clusters 2014A - 2014N to communicate with system memory or other memories that are not local to parallel processing unit 2002. In at least one embodiment, memory crossbar 2016 can use virtual channels to separate the traffic flow between clusters 2014A - 2014N and partition units 2020A - 2020N.

[0293] In at least one embodiment, multiple instances of parallel processing unit 2002 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 2002 can be configured to operate interoperably, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 2002 can include floating-point units with higher precision relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 2002 or parallel processor 2000 can be implemented in various configurations and form factors, including but not limited to desktop computers, laptop computers, or handheld personal computers, servers, workstations, gaming consoles, and / or embedded systems.

[0294] Fig. 20BA processing cluster 2094 according to at least one embodiment is shown. In at least one embodiment, the processing cluster 2094 is included within a parallel processing unit. In at least one embodiment, the processing cluster 2094 is an instance of one of the processing clusters 2014A - 2014N of FIG. 20. In at least one embodiment, the processing cluster 2094 can be configured to execute a number of threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronous threads, which uses a common instruction unit that is configured to issue instructions to a set of processing engines within each processing cluster 2094.

[0295] In at least one embodiment, the processing cluster 2094 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the processing cluster 2094 includes hardware and / or other components for performing the various computing operations and / or APIs described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8

[0296] In at least one embodiment, the operation of the processing cluster 2094 can be controlled by assigning processing tasks to the pipeline manager 2032 of the SIMT parallel processor. In at least one embodiment, the pipeline manager 2032 receives instructions from the scheduler 2010 of FIG. 20 and manages the execution of these instructions through the graphics multiprocessor 2034 and / or the texture unit 2036. In at least one embodiment, the graphics multiprocessor 2034 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures can be included within the processing cluster 2094. In at least one embodiment, one or more instances of the graphics multiprocessor 2034 can be included within the processing cluster 2094. In at least one embodiment, the graphics multiprocessor 2034 can process data, and the data crossbar 2040 can be used to distribute the processed data to one of a number of possible destinations (including other shader units). In at least one embodiment, the pipeline manager 2032 can facilitate the distribution of the processed data by specifying the destination of the processed data to be distributed via the data crossbar 2040.

[0297] ​In at least one embodiment, each graphics multiprocessor 2034 within the processing cluster 2094 may include the same set of functional execution logic (e.g., arithmetic logic unit, load store unit (LSU), etc.). In at least one embodiment, the functional execution logic may be configured in a pipeline manner, where new instructions may be issued before previous instructions are completed. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may exist.

[0298] In at least one embodiment, the instructions transmitted to the processing cluster 2094 constitute threads. In at least one embodiment, a set of threads executed across a group of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group may be assigned to a different processing engine within the graphics multiprocessor 2034. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 2034. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the cycle of processing the thread group. In at least one embodiment, the thread group may also include more threads than the number of processing engines within the graphics multiprocessor 2034. In at least one embodiment, when the thread group includes more threads than the number of processing engines within the graphics multiprocessor 2034, processing may be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups may be executed simultaneously on the graphics multiprocessor 2034.

[0299] In at least one embodiment, the graphics multiprocessor 2034 includes an internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 2034 may relinquish the internal cache and use the cache memory within the processing cluster 2094 (e.g., L1 cache 2048). In at least one embodiment, each graphics multiprocessor 2034 may also access a partitioning unit (e.g., Fig. 20AL2 caches within the partition units 2020A - 2020N), which are shared among all processing clusters 2094 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2034 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 2002 can be used as global memory. In at least one embodiment, the processing cluster 2094 includes multiple instances of the graphics multiprocessor 2034, which can share common instructions and data that can be stored in the L1 cache 2048.

[0300] In at least one embodiment, each processing cluster 2094 can include an MMU 2045 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2045 can reside within the memory interface 2018 of FIG. 20. In at least one embodiment, the MMU 2045 includes a set of page table entries (PTEs) that are used to map virtual addresses to the physical addresses of tiles (talk more about tiles) and optionally to cache line indices. In at least one embodiment, the MMU 2045 can include an address translation lookaside buffer (TLB) or a cache that can reside within the graphics multiprocessor 2034 or the L1 cache 2048 or the processing cluster 2094. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index can be used to determine whether a request for a cache line is a hit or a miss.

[0301] In at least one embodiment, the processing cluster 2094 can be configured such that each graphics multiprocessor 2034 is coupled to a texture unit 2036 to perform texture mapping operations, for example, which can involve determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read as needed from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multiprocessor 2034, and texture data is fetched from the L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2034 outputs the processed task to the data crossbar 2040 to provide the processed task to another processing cluster 2094 for further processing or store the processed task in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 2016. In at least one embodiment, the raster front operation unit (preROP) 2042 is configured to receive data from the graphics multiprocessor 2034 and direct the data to the ROP unit, which can be located together with the partitioning units described herein (e.g., the partitioning units 2020A - 2020N of FIG. 20). In at least one embodiment, the PreROP2042 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.

[0302] Fig. 20C A graphics multiprocessor 2096 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 2096 is Fig. 20B the graphics multiprocessor 2034. In at least one embodiment, the graphics multiprocessor 2096 is coupled to the pipeline manager 2032 of the processing cluster 2094. In at least one embodiment, the graphics multiprocessor 2096 has an execution pipeline that includes, but is not limited to, an instruction cache 2052, an instruction unit 2054, an address mapping unit 2056, a register file 2058, one or more GPGPU cores 2062, and one or more LSUs 2066. The GPGPU cores 2062 and LSUs 2066 are coupled to the cache memory 2072 and the shared memory 2070 via a memory and cache interconnect 2068.

[0303] In at least one embodiment, the graphics multiprocessor 2096 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the graphics multiprocessor 2096 includes means for performing the various computing operations and / or APIs described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8The hardware and / or other components of the described process.

[0304] In at least one embodiment, the instruction cache 2052 receives a stream of instructions to be executed from the pipeline manager 2032. In at least one embodiment, the instructions are cached in the instruction cache 2052 and dispatched for execution by the instruction unit 2054. In one embodiment, the instruction unit 2054 may dispatch instructions as a thread group (e.g., a warp), with each thread of the thread group assigned to a different execution unit within the GPGPU core 2062. In at least one embodiment, instructions may access any local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, the address mapping unit 2056 may be used to translate an address in the unified address space into a different memory address that can be accessed by the LSU 2066.

[0305] In at least one embodiment, the register file 2058 provides a set of registers for the functional units of the graphics multiprocessor 2096. In at least one embodiment, the register file 2058 provides temporary storage for the operands of the data paths of the functional units (e.g., the GPGPU core 2062, the LSU 2066) connected to the graphics multiprocessor 2096. In at least one embodiment, the register file 2058 is partitioned among each of the functional units such that a dedicated portion of the register file 2058 is assigned to each functional unit. In at least one embodiment, the register file 2058 is partitioned among different thread groups being executed by the graphics multiprocessor 2096.

[0306] In at least one embodiment, the GPGPU cores 2062 may each include an FPU and / or an ALU for executing the instructions of the graphics multiprocessor 2096. The GPGPU cores 2062 may be architecturally similar or may have different architectures. In at least one embodiment, a first portion of the GPGPU core 2062 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 2096 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copy rectangle or pixel blend operations. In at least one embodiment, one or more of the GPGPU cores 2062 may also include fixed or special-function logic.

[0307] In at least one embodiment, the GPGPU core 2062 includes SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, the GPGPU core 2062 can physically execute SIMD4, SIMD8, and SIMD9 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time or automatically generated when executing a program written and compiled for a single-program multiple-data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0308] In at least one embodiment, the memory and cache interconnect 2068 is an interconnect network that connects each functional unit of the graphics multiprocessor 2096 to the register file 2058 and the shared memory 2070. In at least one embodiment, the memory and cache interconnect 2068 is a crossbar interconnect that allows the LSU 2066 to perform load and store operations between the shared memory 2070 and the register file 2058. In at least one embodiment, the register file 2058 can operate at the same frequency as the GPGPU core 2062, resulting in very low latency for data transfer between the GPGPU core 2062 and the register file 2058. In at least one embodiment, the shared memory 2070 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 2096. In at least one embodiment, the cache memory 2072 can be used as, for example, a data cache to cache texture data communicated between the functional units and the texture unit 2036. In at least one embodiment, the shared memory 2070 can also be used as a program-managed cache. In at least one embodiment, in addition to the automatically cached data stored in the cache memory 2072, threads executing on the GPGPU core 2062 can also programmatically store data in the shared memory.

[0309] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated with the core on the same package or die and communicatively coupled to the core via an internal processor bus / interconnect (i.e., internal to the package or die). In at least one embodiment, regardless of how the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in the WD. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0310] Fig.21 FIG. 2100 illustrates a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processor 2100 includes a ring interconnect 2102, a pipeline front end 2104, a media engine 2137, and graphics cores 2180A-2180N. In at least one embodiment, the ring interconnect 2102 couples the graphics processor 2100 to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2100 is one of many processors integrated within a multi-core processing system.

[0311] In at least one embodiment, the graphics processor 2100 is used to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the graphics processor 2100 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8

[0312] ​In at least one embodiment, the graphics processor 2100 receives multiple batches of commands via a ring interconnect 2102. In at least one embodiment, the input commands are interpreted by a command stream converter 2103 in a pipeline front end 2104. In at least one embodiment, the graphics processor 2100 includes scalable execution logic to perform 3D geometry processing and media processing via graphics cores 2180A - 2180N. In at least one embodiment, for 3D geometry processing commands, the command stream converter 2103 provides the commands to a geometry pipeline 2136. In at least one embodiment, for at least some media processing commands, the command stream converter 2103 provides the commands to a video front end 2134, which is coupled to a media engine 2137. In at least one embodiment, the media engine 2137 includes a video quality engine (VQE) 2130 for video and image post - processing, and a multi - format encode / decode (MFX) 2133 engine for providing hardware - accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 2136 and the media engine 2137 each generate execution threads for thread execution resources provided by at least one graphics core 2180A.

[0313] In at least one embodiment, the graphics processor 2100 includes scalable thread execution resources characterized by modular graphics cores 2180A - 2180N (sometimes referred to as core slices), each modular core having multiple sub - cores 2150A - 2150N, 2160A - 2160N (sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 2100 can have any number of graphics cores 2180A through 2180N. In at least one embodiment, the graphics processor 2100 includes a graphics core 2180A having at least a first sub - core 2150A and a second sub - core 2160A. In at least one embodiment, the graphics processor 2100 is a low - power processor having a single sub - core (e.g., 2150A). In at least one embodiment, the graphics processor 2100 includes multiple graphics cores 2180A - 2180N, each graphics core including a set of first sub - cores 2150A - 2150N and a set of second sub - cores 2160A - 2160N. In at least one embodiment, each of the first sub - cores 2150A - 2150N includes at least a first set of execution units (EUs) 2152A - 2152N and media / texture samplers 2154A - 2154N. In at least one embodiment, each of the second sub - cores 2160A - 2160N includes at least a second set of execution units 2162A - 2162N and samplers 2164A - 2164N. In at least one embodiment, each sub - core 2150A - 2150N, 2160A - 2160N shares a set of shared resources 2170A - 2170N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.

[0314] Fig. 22 Illustrated is a processor 2200 according to at least one embodiment. In at least one embodiment, processor 2200 may include, but is not limited to, logic circuitry for executing instructions. In at least one embodiment, processor 2200 may execute instructions, including x86 instructions, ARM instructions, special instructions for ASICs, etc. In at least one embodiment, processor 2210 may include registers for storing packed data, such as the 64-bit wide MMXTM registers in the Intel Corporation's MMX technology-enabled microprocessors in Santa Clara, California. In at least one embodiment, the MMX registers available in integer and floating-point forms may operate with packed data elements that accompany SIMD and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, the 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or later versions (generally referred to as “SSEx” technology) may hold such packed data operands. In at least one embodiment, processor 2210 may execute instructions to accelerate CUAD programs.

[0315] In at least one embodiment, processor 2200 is used to perform various computing operations, including one or more of the application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, processor 2200 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8

[0316] In at least one embodiment, the processor 2200 includes an in-order front end (“front end”) 2201 to fetch instructions to be executed and prepare the instructions for later use in the processor pipeline. In at least one embodiment, the front end 2201 may include several units. In at least one embodiment, the instruction prefetcher 2226 fetches instructions from memory and provides the instructions to the instruction decoder 2228, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2228 decodes the received instruction into one or more operations of so-called “microinstructions” or “micro-operations” (also referred to as “uops” or “microinstructions”) for execution. In at least one embodiment, the instruction decoder 2228 parses the instruction into an opcode and corresponding data and control fields, which can be used by the microarchitecture to perform the operation. In at least one embodiment, the trace cache 2230 may assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 2234 for execution. In at least one embodiment, when the trace cache 2230 encounters a complex instruction, the microcode ROM 2232 provides the microinstructions required to complete the operation.

[0317] In at least one embodiment, some instructions can be converted into a single micro-operation, while other instructions require several micro-operations to complete the entire operation. In at least one embodiment, if more than four microinstructions are required to complete an instruction, the instruction decoder 2228 may access the microcode ROM 2232 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of microinstructions for processing at the instruction decoder 2228. In at least one embodiment, if multiple microinstructions are required to complete an operation, the instruction may be stored in the microcode ROM 2232. In at least one embodiment, the trace cache 2230 refers to an entry point programmable logic array (“PLA”) to determine the correct microinstruction pointer for reading a microcode sequence from the microcode ROM 2232 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 2232 has completed the micro-operation sequencing of the instruction, the front end 2201 of the machine may resume fetching micro-operations from the trace cache 2230.

[0318] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2203 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction stream to optimize performance as the instructions descend down the pipeline and are scheduled for execution. The out-of-order execution engine 2203 includes, but is not limited to, an allocator / register renamer 2240, a memory micro-instruction queue 2242, an integer / floating-point micro-instruction queue 2244, a memory scheduler 2246, a fast scheduler 2202, a slow / general floating-point scheduler ("slow / general FP scheduler") 2204, and a simple floating-point scheduler ("simple FP scheduler") 2206. In at least one embodiment, the fast scheduler 2202, the slow / general floating-point scheduler 2204, and the simple floating-point scheduler 2206 are also collectively referred to as "micro-instruction schedulers 2202, 2204, 2206". The allocator / register renamer 2240 allocates the machine buffers and resources required for each micro-instruction to execute in order. In at least one embodiment, the allocator / register renamer 2240 renames logical registers to entries in the register file. In at least one embodiment, the allocator / register renamer 2240 also allocates entries for each micro-instruction to one of two micro-instruction queues, the memory micro-instruction queue 2242 for memory operations and the integer / floating-point micro-instruction queue 2244 for non-memory operations, in front of the memory scheduler 2246 and the micro-instruction schedulers 2202, 2204, 2206. In at least one embodiment, the micro-instruction schedulers 2202, 2204, 2206 determine when a micro-instruction is ready for execution based on the readiness of their dependent input register operand sources and the availability of execution resources micro-instructions that need to be completed. In at least one embodiment, the fast scheduler 2202 of at least one embodiment may be scheduled on each half of the main clock cycle, while the slow / general floating-point scheduler 2204 and the simple floating-point scheduler 2206 may be scheduled once per main processor clock cycle. In at least one embodiment, the micro-instruction schedulers 2202, 2204, 2206 arbitrate the scheduling ports to schedule micro-instructions for execution.

[0319] In at least one embodiment, execution block 2211 includes, but is not limited to, integer register file / branch network 2208, floating-point register file / branch network (“FP register file / branch network”) 2210, address generation units (“AGUs”) 2212 and 2214, fast arithmetic logic units (“fast ALUs”) 2216 and 2218, slow ALU 2220, floating-point ALU (“FP”) 2222, and floating-point move unit (“FP move”) 2224. In at least one embodiment, integer register file / branch network 2208 and floating-point register file / bypass network 2210 are also referred to herein as “register files 2208, 2210”. In at least one embodiment, AGUs 2212 and 2214, fast ALUs 2216 and 2218, slow ALU 2220, floating-point ALU 2222, and floating-point move unit 2224 are also referred to herein as “execution units 2212, 2214, 2216, 2218, 2220, 2222, and 2224”. In at least one embodiment, the execution block may include, but is not limited to, any number (including zero) and type of register files, branch networks, address generation units, and execution units (in any combination).

[0320] In at least one embodiment, register files 2208, 2210 may be arranged between microinstruction schedulers 2202, 2204, 2206 and execution units 2212, 2214, 2216, 2218, 2220, 2222, and 2224. In at least one embodiment, integer register file / branch network 2208 performs integer operations. In at least one embodiment, floating-point register file / branch network 2210 performs floating-point operations. In at least one embodiment, each of register files 2208, 2210 may include, but is not limited to, a branch network that may bypass or forward a just-completed result that has not yet been written to the register file to a new dependent. In at least one embodiment, register files 2208, 2210 may communicate data with each other. In at least one embodiment, integer register file / branch network 2208 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating-point register file / branch network 2210 may include, but is not limited to, 128-bit-wide entries, since floating-point instructions typically have operands with widths of 64 to 128 bits.

[0321] In at least one embodiment, execution units 2212, 2214, 2216, 2218, 2220, 2222, 2224 may execute instructions. In at least one embodiment, register files 2208, 2210 store integer and floating-point data operand values that the microinstructions need to execute. In at least one embodiment, the processor 2200 may include, but is not limited to, any number of execution units 2212, 2214, 2216, 2218, 2220, 2222, 2224 and combinations thereof. In at least one embodiment, the floating-point ALU 2222 and the floating-point move unit 2224 may execute floating-point, MMX, SIMD, AVX, and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, the floating-point ALU 2222 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, instructions involving floating-point values may be processed with floating-point hardware. In at least one embodiment, ALU operations may be passed to the fast ALUs 2216, 2218. In at least one embodiment, the fast ALUs 2216, 2218 may execute fast operations with an effective delay of half a clock cycle. In at least one embodiment, most complex integer operations enter the slow ALU 2220 because the slow ALU 2220 may include, but is not limited to, integer execution hardware for long-delay type operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by the AGUs 2212, 2214. In at least one embodiment, the fast ALU 2216, the fast ALU 2218, and the slow ALU 2220 may perform integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 2216, the fast ALU 2218, and the slow ALU 2220 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, the floating-point ALU 2222 and the floating-point move unit 2224 may be implemented to support a certain range of operands with bits of various widths. In at least one embodiment, the floating-point ALU 2222 and the floating-point move unit 2224 may operate on 128-bit wide packed data operands in combination with SIMD and multimedia instructions.

[0322] In at least one embodiment, the microinstruction schedulers 2202, 2204, 2206 schedule dependent operations before the completion of the execution of the parent load. In at least one embodiment, since microinstructions can be scheduled and executed speculatively in the processor 2200, the processor 2200 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be dependent operations running in the pipeline, which leave the scheduler temporarily without the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that used incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0323] In at least one embodiment, the term "register" may refer to an on-board processor storage location that can be part of an instruction used to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from the programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also contains eight multimedia SIMD registers for packing data.

[0324] Fig.23 A processor 2300 is shown in accordance with at least one embodiment. In at least one embodiment, the processor 2300 includes, but is not limited to, one or more processor cores (cores)

[0325] 2302A-2302N, an integrated memory controller 2314, and an integrated graphics processor 2308. In at least one embodiment, the processor 2300 may include additional cores up to and including an additional processor core 2302N represented by the dashed box. In at least one embodiment, each processor core 2302A-2302N includes one or more internal cache units 2304A-2304N. In at least one embodiment, each processor core may also access one or more shared cache units 2306.

[0326] In at least one embodiment, the processor 2300 is used to perform various computing operations, including those described above in connection with Figure 2-7 One or more of the described application programming interfaces (APIs) and / or the processes described above in connection with Figure 8 the described processes. In at least one embodiment, the processor 2300 includes hardware and / or other components for performing the various computational operations and / or APIs described above in connection with Figure 2-7 the described processes and / or the processes described above in connection with Figure 8 the described processes.

[0327] In at least one embodiment, the internal cache units 2304A - 2304N and the shared cache unit 2306 represent the cache memory hierarchy within the processor 2300. In at least one embodiment, the cache memory units 2304A - 2304N can include at least one level of instruction and data within each processor core and one or more levels of cache in the shared mid - level cache, such as L2, L3, level 4 (L4), or other levels of cache, where the highest - level cache is classified as the LLC prior to the external memory. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 2306 and 2304A - 2304N.

[0328] In at least one embodiment, the processor 2300 may further include a set of one or more bus controller units 2316 and a system agent core 2310. In at least one embodiment, the one or more bus controller units 2316 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2310 provides management functions for the various processor components. In at least one embodiment, the system agent core 2310 includes one or more integrated memory controllers 2314 to manage access to various external memory devices (not shown).

[0329] In at least one embodiment, one or more of the processor cores 2302A - 2302N include support for multi - threading simultaneously. In at least one embodiment, the system agent core 2310 includes components for coordinating and operating the processor cores 2302A - 2302N during multi - threaded processing. In at least one embodiment, the system agent core 2310 may additionally include a power control unit (PCU) that includes logic and components to regulate one or more power states of the processor cores 2302A - 2302N and the graphics processor 2308.

[0330] In at least one embodiment, the processor 2300 further includes a graphics processor 2308 to perform graphics processing operations. In at least one embodiment, the graphics processor 2308 is coupled to a shared cache unit 2306 and a system agent core 2310 that includes one or more integrated memory controllers 2314. In at least one embodiment, the system agent core 2310 further includes a display controller 2311 for driving the output of the graphics processor to one or more coupled displays. In at least one embodiment, the display controller 2311 can also be an independent module coupled to the graphics processor 2308 via at least one interconnect, or can be integrated within the graphics processor 2308.

[0331] In at least one embodiment, a ring-based interconnect unit 2312 is used to couple the internal components of the processor 2300. In at least one embodiment, alternative interconnect units can be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 2308 is coupled to the ring interconnect 2312 via an I / O link 2313.

[0332] In at least one embodiment, the I / O link 2313 represents at least one of a variety of I / O interconnects, including a package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2318 (such as an eDRAM module). In at least one embodiment, each of the processor cores 2302A - 2302N and the graphics processor 2308 uses the embedded memory module 2318 as a shared LLC.

[0333] In at least one embodiment, the processor cores 2302A - 2302N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 2302A - 2302N are heterogeneous in terms of the ISA, where one or more of the processor cores 2302A - 2302N execute a common instruction set, while one or more other processor cores 2302A - 2302N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 2302A - 2302N are heterogeneous in terms of the microarchitecture, where one or more cores with relatively high power consumption are coupled to one or more power cores with lower power consumption. In at least one embodiment, the processor 2300 can be implemented on one or more chips or be implemented as a SoC integrated circuit.

[0334] Fig.24FIG. 2400 shows a graphics processor core according to at least one of the described embodiments. In at least one embodiment, the graphics processor core 2400 is included within a graphics core array. In at least one embodiment, the graphics processor core 2400 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2400 is an example of a graphics core slice, and the graphics processors described herein can include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 2400 can include a fixed function block 2430 coupled to a plurality of sub-cores 2401A-2401F, also referred to as sub-slices, which include modular blocks of general and fixed function logic.

[0335] In at least one embodiment, the graphics processor core 2400 is configured to perform various computing operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the graphics processor core 2400 includes hardware and / or other components for performing the various computing operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or processes described above in connection with Figure 8 In at least one embodiment, the fixed function block 2430 includes a geometry / fixed function pipeline 2436, e.g., in a lower performance and / or lower power graphics processor implementation, the geometry / fixed function pipeline 2436 can be shared by all sub-cores in the graphics processor 2400. In at least one embodiment, the geometry / fixed function pipeline 2436 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0336] In at least one embodiment, the fixed function block 2430 further includes a graphics SoC interface 2437, a graphics microcontroller 2438, and a media pipeline 2439. The graphics SoC interface 2437 provides an interface between the graphics core 2400 and other processor cores in the SoC integrated circuit system. In at least one embodiment, the graphics microcontroller 2438 is a programmable sub-processor that can be configured to manage various functions of the graphics processor 2400, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 2439 includes logic to assist in decoding, encoding, preprocessing, and / or postprocessing multimedia data including image and video data. In at least one embodiment, the media pipeline 2439 implements media operations via requests to the computing or sampling logic within the sub-cores 2401-2401F.

[0337] In at least one embodiment, the fixed function block 2430 further includes a graphics SoC interface 2437, a graphics microcontroller 2438, and a media pipeline 2439. The graphics SoC interface 2437 provides an interface between the graphics core 2400 and other processor cores in the SoC integrated circuit system. In at least one embodiment, the graphics microcontroller 2438 is a programmable sub-processor that can be configured to manage various functions of the graphics processor 2400, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 2439 includes logic to assist in decoding, encoding, preprocessing, and / or postprocessing multimedia data including image and video data. In at least one embodiment, the media pipeline 2439 implements media operations via requests to the computing or sampling logic within the sub-cores 2401-2401F.

[0338] In at least one embodiment, the SoC interface 2437 enables the graphics core 2400 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared LLC memory, system RAM, and / or embedded on-chip or package DRAM. In at least one embodiment, the SoC interface 2437 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline), and enable the use and / or implementation of global memory atoms that can be shared between the graphics core 2400 and the CPU within the SoC. In at least one embodiment, the SoC interface 2437 may also implement power management control for the graphics core 2400 and enable an interface between the clock domain of the graphics core 2400 and other clock domains within the SoC. In at least one embodiment, the SoC interface 2437 enables receipt of command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, when a media operation is to be performed, the commands and instructions may be dispatched to the media pipeline 2439, or when a graphics processing operation is to be performed, they may be assigned to a geometry and fixed-function pipeline (e.g., geometry and fixed-function pipeline 2436, geometry and fixed-function pipeline 2414).

[0339] In at least one embodiment, the graphics microcontroller 2438 may be configured to perform various scheduling and management tasks for the graphics core 2400. In at least one embodiment, the graphics microcontroller 2438 may perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 2402A-2402F, 2404A-2404F in the sub-cores 2401A-2401F. In at least one embodiment, host software executing on a CPU core of an SoC that includes the graphics core 2400 may submit a workload to one of a plurality of graphics processor doorbells, which invokes a scheduling operation on an appropriate graphics engine. In at least one embodiment, the scheduling operations include determining which workload is to be run next, submitting the workload to a command stream converter, pre-empting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 2438 may also facilitate a low-power or idle state for the graphics core 2400, thereby providing the graphics core 2400 with the ability to save and restore registers across low-power state transitions independently of the operating system and / or graphics driver software on the system.

[0340] In at least one embodiment, the graphics core 2400 may have more or fewer sub-cores than the illustrated sub-cores 2401A - 2401F, up to N modular sub-cores. For each group of N sub-cores, in at least one embodiment, the graphics core 2400 may further include shared functional logic 2410, shared and / or cache memory 2412, a geometry / fixed-function pipeline 2414, and additional fixed-function logic 2416 to accelerate various graphics and computing processing operations. In at least one embodiment, the shared functional logic 2410 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 2400. The shared and / or cache memory 2412 may be the LLC of the N sub-cores 2401A - 2401F within the graphics core 2400 and may also be used as shared memory accessible by multiple sub-cores. In at least one embodiment, a geometry / fixed-function pipeline 2414 may be included in place of the geometry / fixed-function pipeline 2436 within the fixed-function block 2430 and may include the same or similar logic units.

[0341] In at least one embodiment, the graphics core 2400 includes additional fixed-function logic 2416, which may include various fixed-function acceleration logics for use by the graphics core 2400. In at least one embodiment, the additional fixed-function logic 2416 includes an additional geometry pipeline for use only in position shading. In position shading only, there are at least two geometry pipelines, and in the full geometry pipeline and culling pipeline within the geometry / fixed-function pipelines 2416, 2436, it is the additional geometry pipeline that may be included in the additional fixed-function logic 2416. In at least one embodiment, the culling pipeline is a trimmed version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline may execute different instances of an application, each instance having a separate environment. In at least one embodiment, position shading only may hide the long culling runs of discarded triangles, thus allowing shading to be completed earlier in some cases. For example, in at least one embodiment, the culling pipeline logic in the additional fixed-function logic 2416 may execute the position shader in parallel with the main application and generally generate critical results faster than the full pipeline because the culling pipeline fetches and masks the position attributes of vertices without performing rasterization and rendering pixels to the frame buffer. In at least one embodiment, the culling pipeline may use the generated critical results to calculate visibility information for all triangles, regardless of whether these triangles are culled. In at least one embodiment, the full pipeline (which may be referred to as a replay pipeline in this case) may consume the visibility information to skip the culled triangles to only shade the visible triangles that ultimately pass to the rasterization stage.

[0342] In at least one embodiment, the additional fixed function logic 2416 may further include general-purpose target processing acceleration logic, such as fixed function matrix multiplication logic, for implementing decelerated CUAD programs.

[0343] In at least one embodiment, a set of execution resources are included within each graphics sub-core 2401A - 2401F and can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-cores 2401A - 2401F include multiple EU arrays 2402A - 2402F, 2404A - 2404F, thread dispatch and inter-thread communication (TD / IC) logic 2403A - 2403F, 3D (e.g., texture) samplers 2405A - 2405F, media samplers 2406A - 2406F, shader processors 2407A - 2407F, and shared local memory (SLM) 2408A - 2408F. Each of the EU arrays 2402A - 2402F, 2404A - 2404F contains multiple execution units, which are GUGPUs capable of servicing graphics, media, or compute operations and performing floating-point and integer / fixed-point logical operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2403A - 2403F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 2405A - 2405F can read data related to textures or other 3D graphics into memory. In at least one embodiment, the 3D samplers can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, the media samplers 2406A - 2406F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 2401A - 2401F may alternatively include a unified 3D and media sampler. In at least one embodiment, the threads executing on the execution units within each sub-core 2401A - 2401F can utilize the shared local memory 2408A - 2408F within each sub-core to enable the threads executing within a thread group to use a common pool of on-chip memory for execution.

[0344] Fig.25Shows a parallel processing unit (“PPU”) 2500 according to at least one embodiment. In at least one embodiment, the PPU 2500 is configured with machine-readable code that, if executed by the PPU 2500, causes the PPU 2500 to perform some or all of the processes and techniques described throughout this document. In at least one embodiment, the PPU 2500 is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) that are executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a set of instructions configured to be executed by the PPU 2500. In at least one embodiment, the PPU 2500 is a graphics processing unit (“GPU”), and the graphics processing unit is configured to implement a graphics rendering pipeline for processing three-dimensional (“3D”) graphics data in order to generate two-dimensional (“2D”) image data for display on a display device (such as an LCD device). In at least one embodiment, the PPU 2500 is used to perform computations such as linear algebra operations and machine learning operations. Fig.25 An example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture implemented in at least one embodiment.

[0345] In at least one embodiment, the PPU 2500 is used to perform various computational operations, including one or more application programming interfaces (APIs) described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8 In at least one embodiment, the PPU 2500 includes hardware and / or other components for performing the various computational operations and / or APIs and / or processes described above in connection with Figure 2-7 and / or the processes described above in connection with Figure 8

[0346] In at least one embodiment, one or more PPU 2500 are configured to accelerate high-performance computing (“HPC”), data center, and machine learning applications. In at least one embodiment, one or more PPU 2500 are configured to accelerate CUDA programs. In at least one embodiment, PPU 2500 includes, but is not limited to, I / O unit 2506, front-end unit 2510, scheduler unit 2512, work distribution unit 2514, hub 2516, crossbar (“Xbar”) 2520, one or more general processing clusters (“GPC”) 2518, and one or more partition units (“memory partition units”) 2522. In at least one embodiment, PPU 2500 is connected to a host processor or other PPU 2500 via one or more high-speed GPU interconnects (“GPU interconnect”) 2508. In at least one embodiment, PPU 2500 is connected to a host processor or other peripheral devices via system bus or interconnect 2502. In one embodiment, PPU 2500 is connected to local memory including one or more memory devices (“memory”) 2504. In at least one embodiment, memory device 2504 includes, but is not limited to, one or more dynamic random access memory (“DRAM”) devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high-bandwidth memory (“HBM”) subsystem, and multiple DRAM dies are stacked within each device.

[0347] In at least one embodiment, the high-speed GPU interconnect 2508 may refer to a wire-based multi-channel communication link used by the system for scaling, and includes one or more PPU 2500 in combination with one or more CPUs (“CPU”), supporting cache coherence between the PPU 2500 and the CPU and CPU master control. In at least one embodiment, the high-speed GPU interconnect 2508 transfers data and / or commands to other units of the PPU 2500 via hub 2516, such as one or more copy engines, video encoders, video decoders, power management units, and / or Fig.25 other components that may not be explicitly shown in

[0348] In at least one embodiment, I / O unit 2506 is configured to receive from the host processor via system bus 2502 ( Fig.25send and receive communications (e.g., commands, data) (not shown in the figure). In at least one embodiment, I / O unit 2506 communicates directly with the host processor via system bus 2502 or via one or more intermediate devices (such as a memory bridge). In at least one embodiment, I / O unit 2506 may communicate with one or more other processors (such as one or more PPU 2500) via system bus 2502. In at least one embodiment, I / O unit 2506 implements a PCIe interface for communication via the PCIe bus. In at least one embodiment, I / O unit 2506 implements an interface for communicating with external devices.

[0349] In at least one embodiment, I / O unit 2506 decodes the packets received via system bus 2502. In at least one embodiment, at least some of the packets represent commands configured to cause PPU2500 to perform various operations. In at least one embodiment, I / O unit 2506 sends the decoded commands to various other units of PPU 2500 as specified by the commands. In at least one embodiment, the commands are sent to the front-end unit 2510 and / or sent to the hub 2516 or other units of PPU2500, such as one or more copy engines, video encoders, video decoders, power management units, etc. ( Fig.25 not explicitly shown in the figure). In at least one embodiment, I / O unit 2506 is configured to route communications between various logical units of PPU 2500.

[0350] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides the workload to PPU 2500 for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in a memory that can be accessed (e.g., read / written) by both the host processor and PPU2500 - the host interface unit may be configured to access the buffer in the system memory connected to system bus 2502 via a memory request transmitted through I / O unit 2506 via system bus 2502. In at least one embodiment, the host processor writes the command stream to the buffer and then sends a pointer indicating the start of the command stream to PPU 2500, such that the front-end unit 2510 receives the pointer(s) to one or more command streams and manages one or more command streams, reads the commands from the command streams and forwards the commands to the respective units of PPU 2500.

[0351] In at least one embodiment, the front-end unit 2510 is coupled to a scheduler unit 2512 that configures various GPCs 2518 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 2512 is configured to track state information related to the various tasks managed by the scheduler unit 2512, where the state information can indicate which GPC 2518 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and so on. In at least one embodiment, the scheduler unit 2512 manages multiple tasks executed on one or more GPCs 2518.

[0352] In at least one embodiment, the scheduler unit 2512 is coupled to a work distribution unit 2514 that is configured to dispatch tasks for execution on the GPCs 2518. In at least one embodiment, the work distribution unit 2514 tracks the multiple scheduled tasks received from the scheduler unit 2512 and the work distribution unit 2514 manages a pool of pending tasks and a pool of active tasks for each GPC 2518. In at least one embodiment, the pool of pending tasks includes multiple time slots (e.g., 32 time slots) that contain tasks assigned to be processed by a particular GPC 2518; the pool of active tasks can include multiple time slots (e.g., 4 time slots) for tasks actively processed by the GPC 2518 such that as one of the GPCs 2518 finishes executing a task, that task is evicted from the active task pool of the GPC 2518, and one of the other tasks from the pool of pending tasks is selected and scheduled for execution on the GPC 2518. In at least one embodiment, if an active task is idle on the GPC 2518, e.g., while waiting for a data dependency to be resolved, the active task is evicted from the GPC 2518 and returned to the pool of pending tasks while another task from the pool of pending tasks is selected and scheduled for execution on the GPC 2518.

[0353] In at least one embodiment, the work distribution unit 2514 communicates with one or more GPCs 2518 via an XBar 2520. In at least one embodiment, the XBar 2520 is an interconnect network that couples many of the units of the PPU 2500 to other units of the PPU 2500 and can be configured to couple the work distribution unit 2514 to a particular GPC 2518. In at least one embodiment, one or more other units of the PPU 2500 can also be connected to the XBar 2520 via a hub 2516.

[0354] In at least one embodiment, tasks are managed by a scheduler unit 2512 and assigned to one of the GPCs 2518 by a work distribution unit 2514. The GPCs 2518 are configured to process the tasks and generate results. In at least one embodiment, the results can be consumed by other tasks in the GPC 2518, routed to a different GPC 2518 via the XBar 2520, or stored in the memory 2504. In at least one embodiment, the results can be written to the memory 2504 by a partitioning unit 2522, which implements a memory interface for writing data to or reading data from the memory 2504. In at least one embodiment, the results can be transmitted to another PPU 2500 or CPU via the high-speed GPU interconnect 2508. In at least one embodiment, the PPU 2500 includes, but is not limited to, U partitioning units 2522, which is equal to the number of separate and distinct memory devices 2504 coupled to the PPU 2500.

[0355] In at least one embodiment, a host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 2500. In one embodiment, multiple compute applications are executed simultaneously b...

Claims

1. A processor, comprising: One or more circuits for using an application programming interface (API) to instruct one or more accelerators within a heterogeneous processor to execute one or more instructions.

2. The processor according to claim 1, wherein the API is used to cause the one or more circuits to indicate a stream to which the one or more instructions are to be added, wherein the stream will be executed at least in part by the one or more accelerators within the heterogeneous processor.

3. The processor according to claim 1, wherein the API is used to cause the one or more circuits to indicate one or more streams to a parallel computing environment, and the one or more instructions will be added to the one or more streams in response to the API.

4. The processor according to claim 1, wherein the API is used to receive as input an identifier indicating a stream for storing the one or more instructions.

5. The processor according to claim 1, wherein the API is used to receive as input a list of operations to be performed by the one or more accelerators within the heterogeneous processor in response to the one or more instructions.

6. The processor according to claim 1, wherein the one or more instructions will be added to an instruction stream that will be executed at least in part by the one or more accelerators within the heterogeneous processor in response to the API.

7. The processor according to claim 1, wherein the processor is a central processing unit (CPU).

8. A system, comprising: One or more processors for using an application programming interface (API) to instruct one or more accelerators within a heterogeneous processor to execute one or more instructions.

9. The system according to claim 8, wherein the API is used to cause the one or more processors to indicate one or more first portions of a stream including the one or more instructions to be executed by the one or more accelerators within the heterogeneous processor and one or more second portions of the stream to be executed by one or more other accelerators.

10. The system according to claim 8, wherein the API is used to cause the one or more processors to indicate to a parallel computing environment a stream to which the one or more instructions are to be added.

11. The system according to claim 8, wherein the API is used to receive as input an identifier indicating a stream for storing the one or more instructions.

12. The system according to claim 8, wherein the API is used to receive as input a list of operations to be performed by the one or more accelerators within the heterogeneous processor in response to the one or more instructions.

13. The system according to claim 8, wherein the one or more instructions will be added to an instruction stream that will be executed at least in part by the one or more accelerators within the heterogeneous processor in response to the API.

14. The system according to claim 8, wherein the API is used to cause the one or more processors to indicate one or more portions of a stream to be executed by the one or more accelerators within the heterogeneous processor.

15. A method, comprising: using an application programming interface (API) to cause one or more accelerators within a heterogeneous processor to execute one or more instructions.

16. The method according to claim 15, further comprising: In response to the API, indicating one or more portions of a stream of the one or more instructions to be executed by the one or more accelerators within the heterogeneous processor.

17. The method according to claim 15 further comprises: In response to the API, indicating a stream to a parallel computing environment, wherein the stream will be executed in part by the one or more accelerators within the heterogeneous processor and in part by one or more other accelerators.

18. The method according to claim 15, wherein the API is used to receive, as input, an identifier of a stream to which the one or more instructions are to be added.

19. The method according to claim 15, wherein the API is used to receive, as input, a list of operations to be performed by the one or more accelerators within the heterogeneous processor in response to the one or more instructions.

20. The method according to claim 15, wherein the one or more instructions will be added, in response to the API, to a stream of instructions to be executed at least in part by the one or more accelerators within the heterogeneous processor.