Priority-based performance allocation

The system optimizes power allocation among multi-core CPUs by prioritizing high-priority threads, improving performance and energy efficiency through dynamic power distribution based on thread priority.

DE102025101099A1Pending Publication Date: 2025-07-31NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025101099
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2025-01-14
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Power allocation to groups of multi-core central processing units (CPUs) is inefficient, leading to suboptimal performance and energy consumption.

Method used

A system and method for asymmetric power allocation among processors based on thread priority, using an API to specify thread categories and allocate power dynamically to maximize performance within power and thermal constraints.

Benefits of technology

Enhances overall performance of processor groups by prioritizing high-priority threads with increased clock frequencies while reducing power consumption and thermal overhead, optimizing power usage across multiple processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Apparatus, systems, and techniques for allocating power to one or more processors. In at least one embodiment, processors or computer systems implement an API to allocate power to one or more processors based at least in part on indications of the priority of one or more threads to be executed by the one or more processors.
Need to check novelty before this filing date? Find Prior Art

Description

AREA

[0001] At least one embodiment relates to processing resources used to execute one or more programs. For example, at least one embodiment relates to processors or computer systems used to execute an API to allocate performance to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. BACKGROUND

[0002] Allocating power to groups of multi-core central processing units ("CPUs") can be inefficient. Techniques for allocating power to groups of multi-core CPUs can be improved. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is an example of a computing device having multiple processors according to at least one embodiment; Fig.2 is an example of a system in which power allocation is differentiated for a multi-processor computing environment according to at least one embodiment; Fig. 3 shows an example process for allocating power between multiple processors according to at least one embodiment; Fig. 4 shows an example process for performing asymmetric power allocation according to at least one embodiment; Fig. 5 shows an example process for allocating power to a processor according to at least one embodiment; Fig. 6 illustrates further aspects of an example process of allocating power to a processor in accordance with at least one embodiment; Fig. 7 shows an example process for determining thread frequency according to at least one embodiment; Fig.8 shows an exemplary data center according to at least one embodiment; Fig. 9 shows a processing system according to at least one embodiment; Fig. 10 shows a computer system according to at least one embodiment; Fig. 11 shows a system according to at least one embodiment; Fig. 12 shows an exemplary integrated circuit according to at least one embodiment; Fig. 13 shows a computer system according to at least one embodiment; Fig. 14 shows an APU according to at least one embodiment; Fig. 15 shows a CPU according to at least one embodiment; Fig. 16 shows an exemplary accelerator integration slice according to at least one embodiment; Fig. 17A-17B illustrate exemplary graphics processors according to at least one embodiment; Fig.18A shows a graphics core according to at least one embodiment; Fig. 18B shows a GPGPU according to at least one embodiment; Fig. 19A shows a parallel processor according to at least one embodiment; Fig. 19B shows a processing cluster according to at least one embodiment; Fig. 19C shows a graphics multiprocessor according to at least one embodiment; Fig. 20 shows a graphics processor according to at least one embodiment; Fig. 21 shows a processor according to at least one embodiment; Fig. 22 shows a processor according to at least one embodiment; Fig. 23 shows a graphics processor core in accordance with at least one embodiment; Fig. 24 shows a PPU according to at least one embodiment; Fig.25 shows a GPC according to at least one embodiment; Fig. 26 shows a streaming multiprocessor according to at least one embodiment; Fig. 27 shows a software stack of a programming platform according to at least one embodiment; Fig. 28 shows a CUDA implementation of a software stack of Fig. 27 according to at least one embodiment; Fig. 29 shows a ROCm implementation of a software stack of Fig. 27 in accordance with at least one embodiment; Fig. 30 shows an OpenCL implementation of a software stack from Fig. 27 in accordance with at least one embodiment; Fig. 31 illustrates software supported by a programming platform in accordance with at least one embodiment; Fig.32 shows the compilation of code for execution on the programming platforms of the Fig. 27 - 30, according to at least one embodiment; Fig. 33 shows in more detail how to compile code for execution on programming platforms of the Fig. 27-30, in accordance with at least one embodiment; Fig. 34 illustrates translating source code prior to compiling source code according to at least one embodiment; Fig. 35A shows a system configured to compile and execute CUDA source code using different types of processing units, according to at least one embodiment; Fig. Figure 35B shows a system configured to run CUDA source code from Fig. 35A compiles and executes using a CPU and a CUDA-enabled graphics processor according to at least one embodiment; Fig.Figure 35C shows a system configured to run the CUDA source code from Fig. 35A compiles and executes using a CPU and a non-CUDA-capable graphics processor according to at least one embodiment; Fig. 36 shows an example kernel created by the CUDA-to-HIP translation tool of Fig. 35C according to at least one embodiment; Fig. 37 shows a non-CUDA capable GPU from Fig. 35C in greater detail according to at least one embodiment; Fig. 38 shows how threads of an example CUDA grid are allocated to different processing units of Fig. 37, according to at least one embodiment; and Fig. 39 shows how existing CUDA code can be migrated to Data Parallel C++ code, according to at least one embodiment. DETAILED DESCRIPTION

[0003] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one skilled in the art that the inventive concepts may be practiced without one or more of these specific details.

[0004] In at least one embodiment, applications can run on a subset of threads with a range of trade-offs between increased parallelism or higher performance using fewer threads. In at least one embodiment, these benefits are enabled by power partitioning that addresses performance at the thread level. In at least one embodiment, this is done by dividing workloads into low-, medium-, and high-priority threads and adjusting the respective distribution and performance of the threads between multicore processors to improve performance in light of power consumption or thermal constraints.

[0005] In at least one embodiment, two or more processors are selectively allocated different amounts of power, each processor including multiple cores executing threads selected for power distribution among the processors. In at least one embodiment, the workloads are divided into low-priority and high-priority threads. In at least one embodiment, the threads are allocated to the two or more processors, and power is allocated to the two or more processors to maximize the performance of the high-priority threads relative to the low-priority threads.

[0006] In at least one embodiment, an API is used to characterize expected workloads. In at least one embodiment, a user of the API, such as a programmer, specifies a number of cores and threads assigned to groups of threads, as well as corresponding priorities for those groups. In at least one embodiment, the user may specify constraints regarding operating temperature, the type of workload (e.g., compute-only, memory-bound, IO-intensive, background task, etc.), or other information. In at least one embodiment, the indications are derived from profiles provided in firmware. In at least one embodiment, the indications include indications regarding expected graphics processing unit ("GPU") utilization and / or indications regarding expected central processing unit ("CPU") utilization.In at least one embodiment, a user may use the API to specify a minimum frequency for low, medium, and / or high priority threads.

[0007] In at least one embodiment, an application executes on a computing platform that includes a superchip. In at least one embodiment, the superchip is an NVIDIA Grace Super Chip. In at least one embodiment, a superchip includes two or more processors, each including multiple cores and connected to other components of the superchip via two corresponding sockets. In at least one embodiment, a socket includes any combination of circuitry and / or physical coupling to enable communication between a processor and other system components. In at least one embodiment, the processors are connected by a communications link, such as NVLink or other high-speed communications buses. In at least one embodiment, the power supplied to each processor is individually controllable.In at least one embodiment, the power is supplied via a corresponding socket.

[0008] In at least one embodiment, a programmer provides parameters to an API to control power distribution between sockets. In at least one embodiment, the parameters describe acceptable power or thermal characteristics of a workload to be executed. In at least one embodiment, firmware assigns power distribution between sockets and causes threads to execute accordingly on the processing cores of a superchip. In at least one embodiment, the API allows a user to specify categories of thread priorities (e.g., high, medium, and low) and a corresponding number of processor cores to be associated with the execution of threads in each category.

[0009] In at least one embodiment, the firmware uses information provided via the API to determine how much power to allocate to each socket and to schedule threads to run on the processor cores associated with each socket such that performance is optimized for a workload comparable to the specified workload, while the power overhead caused by the firmware remains within the specified power consumption or temperature constraints.

[0010] For example, in at least one embodiment, a user might have a power budget of 500 W for a workload to run on a Superchip. In at least one embodiment, the Superchip has two sockets, A and B, each of which is a multicore processor with 72 cores, such that the Superchip has a total of 144 cores across processors A and B. The user can use the API to specify that a workload should have a total of 144 threads, of which 30 should be high priority threads, 42 should be medium priority threads, and 72 should be priority threads, and that these threads should run at no more than 500 W.The firmware determines that this power budget can be met by running 30 high-priority threads at 3.4 GHz, 42 medium-priority threads at 3.1 GHz, and 72 low-priority threads at 2.7 GHz, and that these threads could run at these frequencies if Socket A was supplied with 180 W and Socket B with 320 W. The firmware then causes threads to run on the processor cores in a manner consistent with this allocation, for example, by running low-priority threads from the multi-core CPU connected to Socket A (using 180 W) and medium- and high-priority threads from the multi-core CPU connected to Socket B (using 320 W).

[0011] In at least one embodiment, a processor includes one or more circuits executing an API to allocate power to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. In at least one embodiment, a first processor of the one or more processors is connected to a device via a first socket, and a second processor of the one or more processors is connected to the device via a second socket, wherein power supplied to the first and second processors via the first and second sockets is to be allocated independently.In at least one embodiment, one or more circuits cause power to be allocated to a first processor of the one or more processors based at least in part on a proportion of high-priority threads to be executed by the first processor compared to low-priority threads to be executed by the first processor. In at least one embodiment, one or more circuits select one or more of the one or more threads to be executed on a first processor of the one or more processors based at least in part on the power to be allocated to the first processor.

[0012] In at least one embodiment, a first processor of the one or more processors includes a plurality of processor cores. In at least one embodiment, one or more circuits assign power allocations to the one or more processors based at least in part on utilization information provided via an application programming interface ("API"). In at least one embodiment, power allocation between at least a first and a second processor of the one or more processors is based at least in part on maintaining a power budget.

[0013] Fig.1 is an example 100 of a multi-processor computing device 102 according to at least one embodiment. In at least one embodiment, a multi-processor computing device is referred to as a superchip. In at least one embodiment, a multi-processor computing device includes two or more processors, such as two CPUs, connected by circuitry. In at least one embodiment, a multi-processor computing device includes two or more different types of processors, such as one or more CPUs and one or more graphics processing units ("GPUs").

[0014] At least in one embodiment, one or more aspects of one or more embodiments associated with Fig.1, combined with one or more aspects of one or more embodiments described herein, including at least the embodiments described in connection with Fig. 2-5. In at least one embodiment, optimizing the overall performance of a group of processors, such as processors 104 and 106, comprises selectively providing higher CPU power to selected groups of threads, in particular by varying an operating frequency of a CPU based on the priority of at least one thread executing on that particular CPU. In at least one embodiment, a default power allocation is defined among the processors. In at least one embodiment, the default allocation is changed to provide one or more of the processors with more power than other processors in order to adjust the performance of threads executing on the processors.

[0015] In at least one embodiment, optimizing the overall performance of a group of processors comprises increasing the clock frequency of processors in a group of processors that process threads with the highest priority levels in that group of processors. In at least one embodiment, the clock frequency is increased to a target frequency. In at least one embodiment, optimizing the overall performance of a group of processors comprises decreasing one or more clock frequencies of a group of processors that process threads with the lowest priority levels within that group of processors to avoid exceeding a power budget for those processors.In at least one embodiment, increasing the clock frequencies of one or more processors in a group and decreasing the clock frequencies of one or more other processors in that group enables faster parallel execution of a workload with an identical power budget for all processors in a group of processors than if no clock frequencies were increased. In at least one embodiment, increasing the clock frequencies of one or more processors in a group and decreasing the clock frequencies of one or more other processors in that group increases the overall performance of that group of processors. In at least one embodiment, a particular operating range is based on a target clock frequency for processors 104 and 106.

[0016] In at least one embodiment, a user provides at least one input to an API specifying categories of priority processor groups and a number of processors or threads associated with each group. In at least one embodiment, a user further specifies an operating temperature of the equipment. In at least one embodiment, a user further specifies a workload specification. In at least one embodiment, a workload specification is derived from profiles provided to an API by a firmware controller. In at least one embodiment, a user specifies a minimum frequency for the low-priority threads. In at least one embodiment, the selection is performed by user input using a selector or, alternatively, using a profiling tool.

[0017] In at least one embodiment, the API communicates information specified or provided by the user to the firmware controller. In at least one embodiment, the API communicates the information to other components of the multi-processor computing device 102.

[0018] Fig. 2 shows an example 200 of a computing device having multiple processors 202 in which power allocation is performed differently, according to at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described in connection with Fig. 2, combined with one or more aspects of one or more embodiments described herein, which includes embodiments described at least in connection with Fig.1 and 3-5. In at least one embodiment, optimizing the overall performance of a group of processors, such as a group including processors 204 and 206, includes selectively providing higher CPU power to certain threads, particularly by varying the power supplied to a processor based on the priority of the threads executing on that processor.

[0019] In at least one embodiment, multi-processor computing device 202 includes Boot and Power Monitoring Processor (BPMP) 214 and Baseboard Management Controller (BMC) 216 controllers, and chips 204 and 206. In at least one embodiment, one or more of BPMP 214 and BMC 216 controllers are communicatively coupled to one or more of processors 204 and 206. In at least one embodiment, BPMP 214 controls a boot flow process. In at least one embodiment, BMC 216 monitors and controls various system hardware devices, such as system sensors and other parameters.

[0020] In at least one embodiment, one or more of the processors 204 and 206 are communicatively coupled to one or more of the controllers BPMP 214 and BMC 216. In at least one embodiment, information such as telemetry data disclosed herein in connection with Fig.3-6, between one or more of the BPMP 214 and BMC 216 controllers and one or more of the processors 204 and 206. In at least one embodiment, information, such as computation instructions that manage operations of a processor, is communicated between the BPMP 214 and BMC 216 controllers and one or more of the processors 204 and 206.

[0021] In at least one embodiment, a power budget is allocated as needed for two processors A 204 and B 206 to serve threads of different priorities. As a relevant example, in at least one embodiment, a power budget of 500 W is allocated as needed for two chips A 204 and B 206 to serve high, medium, and low priority threads. In at least one embodiment, the controllers BPMP 214 and BMC 216 may determine an exemplary workload that includes 30 high priority threads 208 set to operate at 3.4 GHz, 42 medium priority threads 210 set to operate at 3.13 GHz, and 72 low priority threads 212 set to operate at 2.7 GHz.

[0022] In at least one embodiment, the firmware BPMP 214 estimates the performance of processor B 206 based on high priority threads 208. In at least one embodiment, BPMP 214 determines a maximum level of processing achievable for medium priority threads 210 while remaining within a specified power allocation for processor B. In at least one embodiment, thermal constraints are considered. In at least one embodiment, expected cooling is considered. In at least one embodiment, once BPMP 214 has a determination of processor B performance, it provides this information to BMC 216. In at least one embodiment, BMC 216 is capable of distributing a job among the processors within the multi-processor computing device according to priority.In at least one embodiment, BMC 216 controls the overall power budget of the module by adjusting the TPD of processor A 204 and the frequency according to the TDP of processor B 206, taking into account a user-specified minimum frequency requirement for low priority threads 212 executing on processor A.

[0023] In at least one embodiment, the BPMP 214 firmware then sets a maximum frequency for the processor A 204 to execute the low priority threads 212 within a power budget.

[0024] Fig. Figure 3 illustrates an example method for allocating power between multiple processors according to at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments associated with Fig.3, combined with one or more aspects of one or more embodiments described herein, which includes embodiments described at least in connection with Fig. 1-2 and 4-5 are described.

[0025] In step 310, in at least one embodiment, a user provides input to an API. In at least one embodiment, the input includes a label of a number of priority core categories and a classification of the categories with respect to a level of priority. In at least one embodiment, the levels of priority are high, medium, and low. In at least one embodiment, the input further includes a number of cores and / or threads associated with each category, an operating temperature of a machine for executing the user's workloads, and a workload term. In at least one embodiment, a workload term is derived based on profiles provided in the firmware for user selection. In at least one embodiment, a profiler tool is provided to a user for deriving this term.In at least one embodiment, the inputs include specifying a minimum frequency for low priority threads.

[0026] In at least one embodiment, a firmware controller 202 estimates at 312 the performance of processor B 206 based on high priority threads 208 and how much the medium priority threads 210 can achieve within the power allocation of processor B.

[0027] In at least one embodiment, at 314, thermal constraints are input and control outputs are determined based on how high the performance of each chip can be to maintain a certain average thermal constraint based on a cooling implementation.

[0028] In at least one embodiment, at 316, once the BPMP 202 determines the performance of processor B, it is returned to the BMC 214, and the BMC 214 dynamically adjusts the maximum power budget (TPD) of the graphics card of chip A based on the thermal design power (TDP) of processor B. Before the BMC 214 can change the power budget of processor A, the firmware BPMP 202 ensures that the TPD of processor B accommodates the minimum frequency requirements for the low-priority threads 212 of processor A.

[0029] In at least one embodiment, in step 318, the BPMP 202 firmware imposes a frequency limit on processor A based on the new power allocation and thread priorities.

[0030] Fig.4 shows an example of a process 400 for performing asymmetric power allocation in accordance with at least one embodiment. In at least one embodiment, the process 400 is performed using aspects of embodiments described with respect to Fig. 1-3 or 5. For example, in at least one embodiment, the process 400 is performed by the computing device having multiple processors configured with respect to Fig. 1 or Fig. 2 is described.

[0031] In at least one embodiment, the BMC controller 214 provides a performance estimate to the BPMP controller 202 in step 402. In at least one embodiment, the performance estimate is determined based on priority, core count, temperature, and workload. In at least one embodiment, a performance estimate depends on parameters including, but not necessarily limited to, frequency, temperature, and workload characteristics, as indicated in step 310. In at least one embodiment, an API communicates parameters to the process 400.

[0032] In at least one embodiment, in step 404, the controller BPMP estimates a processor B subset similar to BPMP 202 using a base frequency and a highest priority that a core can achieve under a thermally constrained power budget.

[0033] In at least one embodiment, in step 406, the TDP of processor B is returned to the BMC 214, which sets an appropriate TDP for processor B and processor A.

[0034] In at least one embodiment, in step 408, the maximum operating frequencies of chip A and P1 threads and P0 threads running on processor B are determined according to embodiments of algorithms described herein.

[0035] Fig.5 shows an example of a process 500 of allocating power to a processor in accordance with at least one embodiment. In at least one embodiment, BMC 502 controls an overall power budget for a multi-processor device. In at least one embodiment, the multi-processor device may include a first processor, which may be referred to as Chip1, and a second processor, which may be referred to as Chip2. In at least one embodiment, the first and second processors are packaged as chips built into the multi-processor device. In at least one embodiment, the first and second processors are circuits integrated into the device, and each communicates with other devices in the system via an interface that acts as a socket.In at least one embodiment, process 500 is a first phase of power allocation based on an estimate of the predicted power of the thread. This estimate may be performed, for example, as in step 312 shown in FIG. Fig. 3. In at least one embodiment, process 500 is followed by process 600, a second power allocation phase, as further described in connection with Fig.6. For example, in at least one embodiment, a user specifies the expected amounts of threads and thread priorities via an API. For example, in at least one embodiment, the API could be executed to indicate that an expected workload includes 20 high-priority threads, 42 medium-priority threads, and 82 low-priority threads, which may be denoted as (high:medium:low, 20:42:82). In at least one embodiment, the satellite management controller (SatMC) 504a,b determines a distribution of the workload among the processors within a multi-processor computing device.

[0036] For example, in at least one embodiment, a first processor is assigned according to the inputs, as described in connection with step 310 of Fig.3. In at least one embodiment, such an assignment is (Low, 0:0:72), which designates 72 threads to be executed at low priority. In at least one embodiment, an accompanying example assignment for a second processor of said multi-processor device could be (Low:Medium:High, 10:42:20), which designates 10 low-priority threads, 42 medium-priority threads, and 20 high-priority threads.

[0037] In at least one embodiment, BPMP1 firmware 506a and BPMP2 firmware 506b are components that are connected to the first and second processors, respectively, and receive API assignments as determined by SatMC 504a,b, but do not communicate with each other. In at least one embodiment, BPMPs 506a and 506b each calculate performance for each processor based on a number of threads of each priority and assign a frequency cap for each priority category. In at least one embodiment, the performance of the second processor is allocated at 510 according to the respective maximum clock frequencies of the cores processing high, medium, and low priority threads.

[0038] Fig.6 shows another aspect of an example process for allocating power to a processor, according to at least one embodiment. In at least one embodiment, the example process 600 is a phase related to the example process 500, as shown in Fig. 5. In at least one embodiment, sub-process 608 is similar to that associated with Fig.5, which represents a second iteration of assigning groups of threads by priority, but starts from a calculation 510 rather than an estimate 312. In at least one embodiment, a total power consumption (which may be referred to as TPD) for the first processor is assigned according to a difference between the total power consumption for the multi-processor device minus the total power consumption for the second processor. This may be referred to as TPD_Chip1 = Module TDP - TDP_Chip2. In at least one embodiment, the clock frequencies of the high / medium / low priority cores in sub-process 608 are maximized within the budget of the first processor. In at least one embodiment, sub-process 610 maintains the results of sub-process 510, as described in connection with Fig. 5 described.

[0039] Fig.7 shows an example method for determining thread frequency according to at least one embodiment. In at least one embodiment, the process determines the interaction of the frequency assigned to one or more threads. In at least one embodiment, the operating system scheduler 702 takes the priority and determines the frequency allocation to each of the threads 706a-d. In at least one embodiment, Collaborative Processor Performance Control (CPPC) 704 is a framework by which a user programs the priority and priority levels per thread for an application. In at least one embodiment, after assigning a TDP and determining the frequencies for each of the threads, the firmware BPMP 710a for a first processor and BPMP 710b for a second processor limits each of the threads 706a-d to frequency thresholds.In this example, Thread0 is running at high priority (P0) at 3.6 GHz, Thread1 is running at medium priority (P1) at 3.6 GHz, Thread2 is running at low priority (P2) at 2.2 GHz, and Thread3 is running at high priority (P0) at 3.6 GHz. In at least one embodiment, this information is published in the Advanced Configuration and Power Interface (ACPI) tables 708. In at least one embodiment, this information is available to schedulers and the operating system for further scheduling of applications. Data center

[0040] Fig. Figure 8 illustrates an exemplary data center 800 in accordance with at least one embodiment. In at least one embodiment, the data center 800 includes, without limitation, a data center infrastructure layer 810, a framework layer 820, a software layer 830, and an application layer 840.

[0041] In at least one embodiment, as in Fig.8, the data center infrastructure layer 810 may include a resource orchestrator 812, clustered compute resources 814, and node compute resources (“node CRs”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, the Node CRs 816(1)-816(N) may include any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays ("FPGAs"), data processing units ("DPUs") in network devices, graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or hard disk drives), network input / output devices ("NW I / O"), network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more Node CRs among the Node CRs 816(1)-816(N) may be a server that has one or more of the computing resources mentioned above.

[0042] In at least one embodiment, the grouped computing resources 814 may include separate groupings of node CRs housed in one or more racks (not shown), or multiple racks housed in data centers in different geographic locations (also not shown). Separate groupings of node CRs within the grouped computing resources 814 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs, including CPUs or processors, may be grouped in one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0043] In at least one embodiment, resource orchestrator 812 may configure or otherwise control one or more node CRs 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource orchestrator 812 may include a software design infrastructure ("SDI") management entity for data center 800. In at least one embodiment, resource orchestrator 812 may include hardware, software, or a combination thereof.

[0044] In at least one embodiment, as in Fig.8, the framework layer 820 includes, without limitation, a job scheduler 832, a configuration manager 834, a resource manager 836, and a distributed file system 838. In at least one embodiment, the framework layer 820 may include a framework for supporting the software 852 of the software layer 830 and / or one or more applications 842 of the application layer 840. In at least one embodiment, the software 852 or the application(s) 842 may include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be some type of free and open source software web application framework such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 838 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the job scheduler 832 may include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 800. In at least one embodiment, the configuration manager 834 may be capable of configuring different layers, such as the software layer 830 and the framework layer 820, which include Spark and the distributed file system 838 to support processing large amounts of data. In at least one embodiment, the resource manager 836 may be capable of managing clustered or grouped compute resources allocated to support the distributed file system 838 and the job scheduler 832. In at least one embodiment, the clustered or grouped compute resources may include the grouped compute resources 814 in the data center infrastructure layer 810.In at least one embodiment, the resource manager 836 may be coordinated with the resource orchestrator 812 to manage these allocated or assigned computing resources.

[0045] In at least one embodiment, the software 852 included in software layer 830 may include software used by at least portions of node CRs 816(1)-816(N), clustered computing resources 814, and / or distributed file system 838 of framework layer 820. One or more types of software may include, but are not limited to, internet website search software, email virus scanning software, database software, and streaming video content software.

[0046] In at least one embodiment, the application(s) 842 included in the application layer 840 may include one or more types of applications used by at least portions of the node CRs 816(1)-816(N), clustered computing resources 814, and / or the distributed file system 838 of the framework layer 820. At least one or more types of applications may include, without limitation, CUDA applications.

[0047] In at least one embodiment, any configuration manager 834, resource manager 836, and resource orchestrator 812 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may relieve a data center operator 800 from potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of a data center. Computer-aided systems

[0048] The following figures illustrate, without limitation, exemplary computer-based systems that may be used to implement at least one embodiment.

[0049] Fig.9 shows a processing system 900 in accordance with at least one embodiment. In at least one embodiment, the processing system 900 includes one or more processors 902 and one or more graphics processors 908, and may be a single-processor desktop system, a multiprocessor workstation system, or a server system with a large number of processors 902 or processor cores 907. In at least one embodiment, the processing system 900 is a processing platform integrated into a system-on-a-chip ("SoC") integrated circuit for use in mobile, portable, or embedded devices. In at least one embodiment, a processor core 907 is referred to as a compute unit.

[0050] In at least one embodiment, processing system 900 may include or be integrated with a server-based gaming platform, a game console, a media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, processing system 900 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 900 may also include, be coupled to, or integrated with a wearable device, such as a wearable device in the form of a smart watch, smart glasses, an augmented reality device, or a virtual reality device.In at least one embodiment, processing system 900 is a device for a television or set-top box that includes one or more processors 902 and a graphical interface generated by one or more graphics processors 908.

[0051] In at least one embodiment, one or more processors 902 each include one or more processing cores 907 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 907 is configured to process a particular instruction set 909. In at least one embodiment, the instruction set 909 may enable Complex Instruction Set Computing ("CISC"), Reduced Instruction Set Computing ("RISC"), or Very Long Instruction Word ("VLIW") computing. In at least one embodiment, the processor cores 907 may each process a different instruction set 909, which may include instructions that facilitate emulation of other instruction sets.In at least one embodiment, the processor core 907 may also include other processing devices, such as a digital signal processor ("DSP").

[0052] In at least one embodiment, processor 902 includes a cache 904. In at least one embodiment, processor 902 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache is shared by various components of processor 902. In at least one embodiment, processor 902 also uses an external cache (e.g., a Level 3 ("L3") cache or Last Level Cache ("LLC")) (not shown), which may be shared by processor cores 907 using known cache coherence techniques. In at least one embodiment, processor 902 additionally includes a register file 906, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and a register for instructions).In at least one embodiment, register file 906 may include general-purpose registers or other registers.

[0053] In at least one embodiment, one or more processors 902 are coupled to one or more interface buses 910 to communicate communication signals, such as address, data, or control signals, between the processor 902 and other components in the processing system 900. In at least one embodiment, the interface bus 910 may be a processor bus, such as a version of a Direct Media Interface ("DMI" bus). In at least one embodiment, the interface bus 910 is not limited to a DMI bus and may include one or more Peripheral Component Interconnect (e.g., "PCI," PCI Express ("PCIe")) buses, memory interconnects, or other types of interface buses. In at least one embodiment, the processor(s) 902 include an integrated memory controller 916 and a memory controller hub 930.In at least one embodiment, the memory controller 916 facilitates communication between a memory device and other components of the processing system 900, while the platform control hub ("PCH") 930 provides connections to input / output ("I / O") devices via a local I / O bus.

[0054] In at least one embodiment, memory device 920 may be a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, a phase-change memory device, or other memory device having suitable performance to serve as processor memory. In at least one embodiment, memory device 920 may operate as system memory for processing system 900 to store data 922 and instructions 921 for use when one or more processors 902 are executing an application or process. In at least one embodiment, memory controller 916 is also coupled to an optional external graphics processor 912 that can communicate with one or more graphics processors 908 in processors 902 to perform graphics and media operations.In at least one embodiment, a display device 911 may be connected to the processor(s) 902. In at least one embodiment, the display device 911 may comprise one or more internal display devices, for example, in a mobile electronic device or a laptop, or an external display device connected via an interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 911 may comprise a head-mounted display ("HMD") such as a stereoscopic display device for use in virtual reality ("VR") or augmented reality ("AR") applications.

[0055] In at least one embodiment, the storage control hub 930 enables the connection of peripherals to the storage device 920 and the processor 902 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, among others, an audio controller 946, a network interface 934, a firmware interface 928, a wireless transceiver 926, touch sensors 925, and a storage device 924 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 924 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as PCI or PCIe. In at least one embodiment, the touch sensors 925 can include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, the wireless transceiver 926 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a cellular network transceiver such as a 3G, 4G, or Long Term Evolution ("LTE") transceiver. In at least one embodiment, the firmware interface 928 enables communication with the system firmware and may, for example, be a Unified Extensible Firmware Interface ("UEFI"). In at least one embodiment, the network controller 934 may enable network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 910. In at least one embodiment, the audio controller 946 is a multi-channel high-definition audio controller.In at least one embodiment, the processing system 900 includes an optional I / O controller 940 for connecting legacy devices (e.g., Personal System 2 ("PS / 2")) to the processing system 900. In at least one embodiment, the platform control hub 930 may also be connected to one or more Universal Serial Bus ("USB") controllers 942 that connect input devices, such as keyboard and mouse combinations 943, a camera 944, or other USB input devices.

[0056] In at least one embodiment, an instance of the memory controller 916 and the memory controller hub 930 may be integrated into a discrete external graphics processor, such as the external graphics processor 912. In at least one embodiment, the platform controller hub 930 and / or the memory controller 916 may be external to one or more processors 902. For example, in at least one embodiment, the processing system 900 may include an external memory controller 916 and a platform controller hub 930, which may be configured as a memory controller hub and a peripheral controller hub within a system chipset in communication with the processor(s) 902.

[0057] Fig.10 illustrates a computer system 1000 in accordance with at least one embodiment. In at least one embodiment, computer system 1000 may be a system of interconnected devices and components, a System of Operation (SOC), or a combination thereof. In at least one embodiment, computer system 1000 is configured with a processor 1002, which may include execution units for executing an instruction. In at least one embodiment, computer system 1000 may include, without limitation, a component such as processor 1002 to employ execution units including logic for performing algorithms for processing data.In at least one embodiment, computer system 1000 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including PCs with other microprocessors, technical workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 1000 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0058] In at least one embodiment, computer system 1000 may be used in other devices such as handheld devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor (DSP), an SoC, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions.

[0059] In at least one embodiment, computer system 1000 may include, without limitation, a processor 1002, which may include, without limitation, one or more execution units 1008 that may be configured to execute a Compute Unified Device Architecture ("CUDA") program (CUDA® is developed by NVIDIA Corporation of Santa Clara, CA). In at least one embodiment, a CUDA program is at least a portion of a software application written in a CUDA programming language. In at least one embodiment, computer system 1000 is a desktop or server system having a processor. In at least one embodiment, computer system 1000 may be a multiprocessor system.In at least one embodiment, processor 1002 may comprise, without limitation, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other device, such as a digital signal processor. In at least one embodiment, processor 1002 may be connected to a processor bus 1010 that may transmit data signals between processor 1002 and other components in computer system 1000.

[0060] In at least one embodiment, processor 1002 may include, without limitation, an internal Level 1 ("L1") cache memory ("cache") 1004. In at least one embodiment, processor 1002 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may be external to processor 1002. In at least one embodiment, processor 1002 may also include a combination of both internal and external caches. In at least one embodiment, a register file 1006 may store different data types in various registers, including, without limitation, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0061] In at least one embodiment, execution unit 1008, which includes, without limitation, logic for performing integer and floating-point operations, is also located in processor 1002. Processor 1002 may also include microcode read-only memory ("ROM") ("ucode") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 1008 may include logic for handling a packed instruction set 1009. In at least one embodiment, by including a packed instruction set 1009 in an instruction set of a general-purpose processor 1002, along with associated instruction execution circuitry, operations used by many multimedia applications can be performed using packed data in a general-purpose processor 1002.In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by utilizing the full width of a processor's data bus to perform operations on packed data, thereby eliminating the need to transfer smaller units of data across a processor's data bus to perform one or more operations on one data element at a time.

[0062] In at least one embodiment, execution unit 1008 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1000 may include, without limitation, a memory 1020. In at least one embodiment, memory 1020 may be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage device. Memory 1020 may store instruction(s) 1019 and / or data 1021 represented by data signals that may be executed by processor 1002.

[0063] In at least one embodiment, a system logic chip may be connected to the processor bus 1010 and the memory 1020. In at least one embodiment, the system logic chip may include, without limitation, a memory control hub ("MCH") 1016, and the processor 1002 may communicate with the MCH 1016 via the processor bus 1010. In at least one embodiment, the MCH 1016 may provide a high-bandwidth memory path 1018 to the memory 1020 for storing instructions and data, as well as for storing graphics instructions, data, and textures. In at least one embodiment, the MCH 1016 may route data signals between the processor 1002, the memory 1020, and other components in the computer system 1000, and may bridge data signals between the processor bus 1010, the memory 1020, and a system I / O 1022.In at least one embodiment, the system logic chip may provide a graphics port for connection to a graphics controller. In at least one embodiment, the MCH 1016 may be coupled to the memory 1020 via a high-bandwidth memory path 1018, and the graphics / video card 1012 may be coupled to the MCH 1016 via an Accelerated Graphics Port ("AGP") interconnect 1014.

[0064] In at least one embodiment, computer system 1000 may use system I / O 1022, which is a proprietary hub interface, to connect MCH 1016 to I / O controller hub ("ICH") 1030. In at least one embodiment, ICH 1030 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1020, a chipset, and processor 1002. Examples may include, without limitation, an audio controller 1029, a firmware hub (“flash BIOS”) 1028, a wireless transceiver 1026, a data store 1024, a legacy I / O controller 1023 with a user input interface 1025 and a keyboard interface, a serial expansion port 1027, such as a USB, and a network controller 1034.The data storage 1024 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0065] In at least one embodiment, Fig. 10 a system comprising interconnected hardware devices or "chips." In at least one embodiment, Fig. 10 show an exemplary SoC. In at least one embodiment, the Fig. 10 may be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of system 1000 are interconnected using Compute Express Link ("CXL") interconnects.

[0066] Fig.11 shows a system 1100 in accordance with at least one embodiment. In at least one embodiment, the system 1100 is an electronic device employing a processor 1110. In at least one embodiment, the system 1100 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, an edge device communicatively coupled to one or more on-premises or cloud service providers, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0067] In at least one embodiment, system 1100 may include, without limitation, a processor 1110 communicatively coupled to any number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1110 is coupled via a bus or interface, such as an I2C bus, a System Management Bus ("SMBus"), a Low Pin Count Bus ("LPC"), a Serial Peripheral Interface ("SPI"), a High Definition Audio Bus ("HDA"), a Serial Advance Technology Attachment Bus ("SATA"), a USB bus (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter Bus ("UART"). In at least one embodiment, Fig. 11 a system comprising interconnected hardware devices or "chips." In at least one embodiment, Fig. 11 show an exemplary SoC. In at least one embodiment, the Fig.11 may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 11 interconnected using CXL connections.

[0068] In at least one embodiment, Fig.11 a display 1124, a touchscreen 1125, a touchpad 1130, a near-field communication unit (“NFC”) 1145, a sensor hub 1140, a thermal sensor 1146, an express chipset (“EC”) 1135, a trusted platform module (“TPM”) 1138, BIOS / firmware / flash memory (“BIOS, FW Flash”) 1122, a DSP 1160, a solid-state disk (“SSD”) or hard disk (“HDD”) 1120, a wireless local area network unit (“WLAN”) 1150, a Bluetooth unit 1152, a wireless wide area network unit (“WWAN”) 1156, a global positioning system (“GPS”) 1155, a camera (“USB 3.0 camera”) 1154, such as a USB 3.0 camera, or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 1115, implemented, for example, in the LPDDR3 standard. These components can be implemented in any suitable manner.

[0069] In at least one embodiment, other components may be communicatively coupled to the processor 1110 via the components described above. In at least one embodiment, an accelerometer 1141, an ambient light sensor ("ALS") 1142, a compass 1143, and a gyroscope 1144 may be communicatively coupled to the sensor hub 1140. In at least one embodiment, a thermal sensor 1139, a fan 1137, a keyboard 1136, and a touchpad 1130 may be communicatively coupled to the EC 1135. In at least one embodiment, a speaker 1163, a headset 1164, and a microphone ("mic") 1165 may be communicatively coupled to an audio unit ("audio codec and class d amp") 1162, which in turn may be communicatively coupled to the DSP 1160. In at least one embodiment, the audio unit 1162 may include, for example and without limitation, an audio encoder / decoder ("codec") and a Class D amplifier.In at least one embodiment, a SIM card ("SIM") 1157 may be communicatively coupled to the WWAN unit 1156. In at least one embodiment, components such as the WLAN unit 1150 and the Bluetooth unit 1152, as well as the WWAN unit 1156, may be implemented in a Next Generation Form Factor ("NGFF").

[0070] Fig.12 shows an example integrated circuit 1200 in accordance with at least one embodiment. In at least one embodiment, the example integrated circuit 1200 is a SoC that can be manufactured using one or more IP cores. In at least one embodiment, the integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs, DPUs), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I12S / I2C controller 1240.In at least one embodiment, integrated circuit 1200 may include a display device 1245 coupled to one or more of the following interfaces: a High-Definition Multimedia Interface ("HDMI") controller 1250 and a Mobile Industry Processor Interface ("MIPI") display interface 1255. In at least one embodiment, memory may be provided by a flash memory subsystem 1260, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1265 for accessing SDRAM or SRAM devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 1270.

[0071] Fig.13 shows a computer system 1300 according to at least one embodiment; in at least one embodiment, the computer system 1300 includes a processing subsystem 1301 having one or more processors 1302 and a system memory 1304 communicating via an interconnect path that may include a memory hub 1305. In at least one embodiment, the memory hub 1305 may be a separate component within a chipset component or integrated with one or more processors 1302. In at least one embodiment, the memory hub 1305 is coupled to an I / O subsystem 1311 via a communication link 1306. In at least one embodiment, the I / O subsystem 1311 includes an I / O hub 1307, which may enable the computer system 1300 to receive input from one or more input devices 1308.In at least one embodiment, the I / O hub 1307 may enable a display controller, which may be included in one or more processors 1302, to provide output to one or more display devices 1310A. In at least one embodiment, one or more display devices 1310A coupled to the I / O hub 1307 may comprise a local, internal, or embedded display device.

[0072] In at least one embodiment, processing subsystem 1301 includes one or more parallel processors 1312 connected to storage hub 1305 via a bus or other communication link 1313. In at least one embodiment, communication link 1313 may be any number of standards-based communication link technologies or protocols, such as, but not limited to, PCIe, or a vendor-specific communication interface or structure. In at least one embodiment, one or more parallel processors 1312 form a computationally focused parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as many integrated core processors or processing units.In at least one embodiment, one or more parallel processors 1312 form a graphics processing subsystem that can output pixels to one or more display devices 1310A coupled via I / O hub 1307. In at least one embodiment, one or more parallel processors 1312 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1310B.

[0073] In at least one embodiment, a system storage unit 1314 may be connected to the I / O hub 1307 to provide a storage mechanism for the computer system 1300. In at least one embodiment, an I / O switch 1316 may be used to provide an interface enabling connections between the I / O hub 1307 and other components, such as a network adapter 1318 and / or a wireless network adapter 1319 that may be integrated into a platform, and various other devices that may be added via one or more add-in devices 1320. In at least one embodiment, the network adapter 1318 may be an Ethernet adapter or other wired network adapter.In at least one embodiment, the wireless network adapter 1319 may include one or more Wi-Fi, Bluetooth, NFC, or other network devices that include one or more wireless radios.

[0074] In at least one embodiment, computer system 1300 may include other components not explicitly shown, including USB or other connectors, optical storage devices, video capture devices, and the like, which may also be connected to I / O hub 1307. In at least one embodiment, communication paths connecting various components in Fig. 13 interconnection may be implemented using any suitable protocols, such as PCI-based protocols (e.g., PCIe) or other bus or point-to-point communication interfaces and / or protocols, such as NVLink high-speed interconnection or interconnection protocols.

[0075] In at least one embodiment, one or more parallel processors 1312 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and form a graphics processing unit ("GPU"). In at least one embodiment, one or more parallel processors 1312 include circuitry optimized for general processing. In at least one embodiment, components of computer system 1300 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1312, memory hub 1305, processor(s) 1302, and I / O hub 1307 may be integrated into an SoC integrated circuit.In at least one embodiment, the components of computer system 1300 may be integrated into a single package to form a system-in-package ("SIP") configuration. In at least one embodiment, at least a portion of the components of computer system 1300 may be integrated into a multi-chip module ("MCM") that may be interconnected with other multi-chip modules to form a modular computer system. In at least one embodiment, I / O subsystem 1311 and display devices 1310B are not included in computer system 1300. Processing systems

[0076] The following figures illustrate, without limitation, exemplary processing systems that may be used to implement at least one embodiment.

[0077] Fig.14 shows an accelerated processing unit ("APU") 1400 in accordance with at least one embodiment. In at least one embodiment, the APU 1400 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, the APU 1400 may be configured to execute an application program, such as a CUDA program. In at least one embodiment, the APU 1400 includes, without limitation, a core complex 1410, a graphics complex 1440, a fabric 1460, I / O interfaces 1470, memory controllers 1480, a display controller 1492, and a multimedia engine 1494. In at least one embodiment, the APU 1400 may include, without limitation, any number of core complexes 1410, any number of graphics complexes 1450, any number of display controllers 1492, and any number of multimedia engines 1494 in any combination.For explanatory purposes, multiple instances of the same object are referred to here with reference numbers that identify the object and with parentheses that identify the instance where necessary.

[0078] In at least one embodiment, core complex 1410 is a CPU, graphics complex 1440 is a GPU, and APU 1400 is a processing unit that integrates, without limitation, 1410 and 1440 on a single chip. In at least one embodiment, some tasks may be assigned to core complex 1410 and other tasks to graphics complex 1440. In at least one embodiment, core complex 1410 is configured to execute main control software associated with APU 1400, such as an operating system. In at least one embodiment, core complex 1410 is the main processor of APU 1400, controlling and coordinating the operations of the other processors. In at least one embodiment, core complex 1410 issues instructions that control the operation of graphics complex 1440.In at least one embodiment, core complex 1410 may be configured to execute host executable code derived from CUDA source code, and graphics complex 1440 may be configured to execute device executable code derived from CUDA source code.

[0079] In at least one embodiment, core complex 1410 includes, without limitation, cores 1420(1)-1420(4) and an L3 cache 1430. In at least one embodiment, core complex 1410 may include, without limitation, any number of cores 1420 and any number and type of caches in any combination. In at least one embodiment, cores 1420 are configured to execute instructions of a particular instruction set architecture ("ISA"). In at least one embodiment, each core 1420 is a CPU core. In at least one embodiment, core 1420 is referred to as a compute unit.

[0080] In at least one embodiment, each core 1420 includes, without limitation, a fetch / decode unit 1422, an integer execution engine 1424, a floating-point execution engine 1426, and an L2 cache 1428. In at least one embodiment, the fetch unit 1422 fetches instructions, decodes such instructions, generates micro-operations, and sends separate micro-instructions to the integer execution engine 1424 and the floating-point execution engine 1426. In at least one embodiment, the fetch unit 1422 may concurrently send one micro-instruction to the integer execution engine 1424 and another micro-instruction to the floating-point execution engine 1426. In at least one embodiment, the integer execution engine 1424 performs, without limitation, integer and memory operations. In at least one embodiment, the floating-point engine 1426 performs, without limitation, floating-point and vector operations.In at least one embodiment, fetch unit 1422 forwards microinstructions to a single execution engine that replaces both integer execution engine 1424 and floating point execution engine 1426.

[0081] In at least one embodiment, each core 1420(i), where i is an integer representing a particular instance of core 1420, can access L2 cache 1428(i) that includes core 1420(i). In at least one embodiment, each core 1420 included in core complex 1410(j), where j is an integer representing a particular instance of core complex 1410, is connected to other cores 1420 included in core complex 1410(j) via L3 cache 1430(j) included in core complex 1410(j). In at least one embodiment, the cores 1420 included in core complex 1410(j), where j is an integer representing a particular instance of core complex 1410, may access the entire L3 cache 1430(j) included in core complex 1410(j). In at least one embodiment, L3 cache 1430 may include any number of slices, without limitation.

[0082] In at least one embodiment, graphics complex 1440 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 1440 is configured to perform graphics pipeline operations such as drawing instructions, pixel operations, geometric calculations, and other operations related to rendering an image on a display. In at least one embodiment, graphics complex 1440 is configured to perform non-graphics operations. In at least one embodiment, graphics complex 1440 is configured to perform both graphics-related and non-graphics operations.

[0083] In at least one embodiment, the graphics complex 1440 includes, without limitation, any number of compute units 1450 and an L2 cache 1442. In at least one embodiment, the compute units 1450 share the L2 cache 1442. In at least one embodiment, the L2 cache 1442 is partitioned. In at least one embodiment, the graphics complex 1440 includes, without limitation, any number of compute units 1450 and any number (including zero) and type of caches. In at least one embodiment, the graphics complex 1440 includes, without limitation, any amount of dedicated graphics hardware.

[0084] In at least one embodiment, each compute unit 1450 includes, without limitation, any number of SIMD units 1452 and a shared memory 1454. In at least one embodiment, each SIMD unit 1452 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 1450 can execute any number of thread blocks, but each thread block executes on a single compute unit 1450. In at least one embodiment, a thread block includes, without limitation, any number of threads of execution. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 1452 executes a different warp.In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process a different set of instructions based on a single set of instructions. In at least one embodiment, predication can be used to deactivate one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp thread. In at least one embodiment, different wavefronts in a thread block can synchronize with each other and communicate via a shared memory 1454.

[0085] In at least one embodiment, structure 1460 is a system interconnect that enables data and control transfers between core complex 1410, graphics complex 1440, I / O interfaces 1470, memory interfaces 1480, display controller 1492, and multimedia engine 1494. In at least one embodiment, APU 1400 may include, without limitation, any number and type of system interconnect in addition to or in place of structure 1460 that enables data and control transfers via any number and type of directly or indirectly connected components that may be internal or external to APU 1400. In at least one embodiment, I / O interfaces 1470 are representative of any number and type of I / O interfaces (e.g., PCI, PCI-Extended ("PCI-X"), PCIe, Gigabit Ethernet ("GBE"), USB, etc.).In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 1470. In at least one embodiment, peripheral devices coupled to I / O interfaces 1470 may include, without limitation, keyboards, mice, printers, scanners, joysticks or other types of gaming controllers, media recording devices, external storage devices, network interface cards, and so on.

[0086] In at least one embodiment, the display controller AMD92 displays images on one or more display devices, such as a liquid crystal display ("LCD") device. In at least one embodiment, the multimedia engine 1494 includes, without limitation, any number and type of multimedia-related circuitry, such as a video decoder, a video processor, an image signal processor, etc. In at least one embodiment, the memory controllers 1480 facilitate data transfer between the APU 1400 and a unified system memory 1490. In at least one embodiment, the core complex 1410 and the graphics complex 1440 share the unified system memory 1490.

[0087] In at least one embodiment, the APU 1400 implements a memory subsystem, including, without limitation, any number and type of memory controllers 1480 and memory devices (e.g., shared memory 1454) that may be dedicated to a component or shared among multiple components. In at least one embodiment, the APU 1400 implements a cache subsystem, including, without limitation, one or more caches (e.g., L2 caches 1428, L3 cache 1430, and L2 cache 1442), each of which may be private or shared among any number of components (e.g., cores 1420, core complex 1410, SIMD units 1452, compute units 1450, and graphics complex 1440).

[0088] Fig.15 shows a CPU 1500 in accordance with at least one embodiment. In at least one embodiment, the CPU 1500 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, the CPU 1500 may be configured to execute an application program. In at least one embodiment, the CPU 1500 is configured to execute main control software, such as an operating system. In at least one embodiment, the CPU 1500 issues instructions that control the operation of an external GPU (not shown). In at least one embodiment, the CPU 1500 may be configured to execute executable code on the host derived from CUDA source code, and an external GPU may be configured to execute executable code on the device derived from such CUDA source code.In at least one embodiment, CPU 1500 includes, without limitation, any number of core complexes 1510, fabric 1560, I / O interfaces 1570, and memory controllers 1580.

[0089] In at least one embodiment, core complex 1510 includes, without limitation, cores 1520(1)-1520(4) and an L3 cache 1530. In at least one embodiment, core complex 1510 may include, without limitation, any number of cores 1520 and any number and type of caches in any combination. In at least one embodiment, cores 1520 are configured to execute instructions of a particular ISA. In at least one embodiment, each core 1520 is a CPU core.

[0090] In at least one embodiment, each core 1520 includes, without limitation, a fetch / decode unit 1522, an integer execution engine 1524, a floating-point execution engine 1526, and an L2 cache 1528. In at least one embodiment, the fetch / decode unit 1522 fetches instructions, decodes such instructions, generates micro-operations, and sends separate micro-instructions to the integer execution engine 1524 and the floating-point execution engine 1526. In at least one embodiment, the fetch unit 1522 may concurrently send one micro-instruction to the integer execution engine 1524 and another micro-instruction to the floating-point execution engine 1526. In at least one embodiment, the integer execution engine 1524 performs, without limitation, integer and memory operations. In at least one embodiment, the floating-point engine 1526 performs, without limitation, floating-point and vector operations.In at least one embodiment, fetch unit 1522 forwards microinstructions to a single execution engine that replaces both integer execution engine 1524 and floating point execution engine 1526.

[0091] In at least one embodiment, each core 1520(i), where i is an integer representing a particular instance of core 1520, can access L2 cache 1528(i) that includes core 1520(i). In at least one embodiment, each core 1520 included in core complex 1510(j), where j is an integer representing a particular instance of core complex 1510, is connected to other cores 1520 in core complex 1510(j) via L3 cache 1530(j) included in core complex 1510(j). In at least one embodiment, the cores 1520 included in core complex 1510(j), where j is an integer representing a particular instance of core complex 1510, may access the entire L3 cache 1530(j) included in core complex 1510(j). In at least one embodiment, L3 cache 1530 may include any number of slices, without limitation.

[0092] In at least one embodiment, structure 1560 is a system interconnect that facilitates data and control transfers across core complexes 1510(1)-1510(N) (where N is an integer greater than zero), I / O interfaces 1570, and memory controllers 1580. In at least one embodiment, CPU 1500 may include, without limitation, any number and type of system interconnects in addition to or in place of structure 1560 that facilitate data and control transfers across any number and type of directly or indirectly connected components that may be internal or external to CPU 1500. In at least one embodiment, I / O interfaces 1570 are representative of any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 1570.In at least one embodiment, peripheral devices coupled to I / O interfaces 1570 may include, without limitation, displays, keyboards, mice, printers, scanners, joysticks or other types of gaming controllers, media recording devices, external storage devices, network interface cards, and so forth.

[0093] In at least one embodiment, memory controllers 1580 facilitate data transfer between CPU 1500 and system memory 1590. In at least one embodiment, core complex 1510 and graphics complex 1540 share system memory 1590. In at least one embodiment, CPU 1500 implements a memory subsystem, including, without limitation, any number and type of memory controllers 1580 and memory devices that may be dedicated to a component or shared by multiple components. In at least one embodiment, CPU 1500 implements a cache subsystem, including, without limitation, one or more caches (e.g., L2 caches 1528 and L3 caches 1530), each of which may be dedicated to or shared by any number of components (e.g., cores 1520 and core complex 1510).

[0094] Fig.16 shows an exemplary accelerator integration slice 1690 in accordance with at least one embodiment. As used herein, a "slice" comprises a particular portion of the processing resources of an accelerator integration circuit. In at least one embodiment, the accelerator integration circuit provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines included in a graphics acceleration module. The graphics processing engines may each comprise a separate GPU. Alternatively, the graphics processing engines may comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines.In at least one embodiment, the graphics acceleration module may be a GPU with multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a common package, line card, or chip.

[0095] An effective address space 1682 within system memory 1614 stores process elements 1683. In one embodiment, process elements 1683 are stored in response to GPU calls 1681 from applications 1680 executing on processor 1607. A process element 1683 contains the state of the corresponding application 1680. A work description ("WD") 1684 contained in process element 1683 may be a single job requested by an application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1684 is a pointer to a job request queue in the application's effective address space 1682.

[0096] The graphics acceleration module 1646 and / or individual graphics processing engines may be shared by all or a subset of the processes in a system. In at least one embodiment, an infrastructure for establishing process state and sending WD 1684 to the graphics acceleration module 1646 to start a job in a virtualized environment may be included.

[0097] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 1646 or an individual graphics processing engine. Because the graphics acceleration module 1646 is owned by a single process, a hypervisor initializes an accelerator integration circuit for an owning partition, and an operating system initializes the accelerator integration circuit for an owning process when the graphics acceleration module 1646 is allocated.

[0098] In operation, a WD fetch unit 1691 in accelerator integration slice 1690 fetches the next WD 1684, which includes an indication of work to be performed by one or more graphics processing engines of graphics acceleration module 1646. The data from WD 1684 may be stored in registers 1645 and used by a memory management unit ("MMU") 1639, interrupt management circuitry 1647, and / or context management circuitry 1648 (see figure). For example, one embodiment of MMU 1639 includes segment / page walkup circuitry for accessing segment / page tables 1686 in operating system virtual address space 1685. Circuitry 1647 may process interrupt events ("INT") 1692 received from graphics acceleration module 1646. When performing graphics operations, an effective address 1693 generated by a graphics processing engine is translated into a real address by the MMU 1639.

[0099] In one embodiment, the same set of registers 1645 is duplicated for each graphics processing engine and / or graphics acceleration module 1646 and may be initialized by a hypervisor or operating system. Each of these duplicated registers may comprise an accelerator integration slice 1690. Example registers that may be initialized by a hypervisor are listed in Table 1. Table 1 - Initialized hypervisor registers 1 Slice control register 2 Real Address (RA) Area pointer for scheduled processes 3 Authority mask override register 4 Interrupt vector table entry offset 5 Interrupt vector table entry boundary 6 Condition register 7 Logical partition ID 8 Real Address (RA) Pointer for the Accelerator Workload Set (Hypervisor Pointer for the Accelerator Workload Set) 9 Memory description register

[0100] Example registers that can be initialized by an operating system are listed in Table 2. Table 2 - Initialized operating system registers 1 Process and thread identification 2 Effective Address (EA) Context Store / Restore Pointer 3 Virtual Address (VA) Accelerator Workload Set Pointer (Accelerator Workload Set Pointer) 4 Virtuelle Adresse (VA) Zeiger auf die Speichersegmenttabelle 5 Autoritätsmaske 6 Arbeitsbeschreibung

[0101] In one embodiment, each WD 1684 is specific to a particular graphics acceleration module 1646 and / or a particular graphics processing engine. It contains all the information required by a graphics processing engine to perform work, or it may be a pointer to a memory location where an application has established a command queue of work to be performed.

[0102] Fig. 17A-17B illustrate example graphics processors in accordance with at least one embodiment. In at least one embodiment, each of the example graphics processors may be fabricated using one or more IP cores. In addition to the illustrated embodiments, other logic and circuitry may also be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the example graphics processors are intended for use within a SoC.

[0103] Fig. 17A shows an exemplary graphics processor 1710 of an SoC integrated circuit that may be manufactured using one or more IP cores in accordance with at least one embodiment. Fig. 17B shows another exemplary graphics processor 1740 of an integrated circuit (SoC) that may be manufactured using one or more IP cores in accordance with at least one embodiment. In at least one embodiment, the graphics processor 1710 is Fig. 17A, a low-power graphics processor core. In at least one embodiment, the graphics processor 1740 is Fig. 17B, ​​a higher performance graphics processor core. In at least one embodiment, each of the graphics processors 1710, 1740 may be a variant of the graphics processor 1210 of Fig. be 12.

[0104] In at least one embodiment, graphics processor 1710 includes a vertex processor 1705 and one or more fragment processors 1715A-1715N (e.g., 1715A, 1715B, 1715C, 1715D, through 1715N-1, and 1715N). In at least one embodiment, graphics processor 1710 may execute different shader programs via separate logic, such that vertex processor 1705 is optimized to perform operations for vertex shader programs, while one or more fragment processors 1715A-1715N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1705 performs a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data.In at least one embodiment, fragment processor(s) 1715A-1715N use the primitive and vertex data generated by vertex processor 1705 to generate a framebuffer displayed on a display device. In at least one embodiment, fragment processor(s) 1715A-1715N are optimized for executing fragment shader programs as provided in an OpenGL API, which can be used to perform similar operations as a pixel shader program as provided in a Direct 3D API.

[0105] In at least one embodiment, graphics processor 1710 additionally includes one or more MMU(s) 1720A-1720B, cache(s) 1725A-1725B, and circuit interconnect(s) 1730A-1730B. In at least one embodiment, one or more MMU(s) 1720A-1720B provide virtual-to-physical address mapping for graphics processor 1710, including vertex processor 1705 and / or fragment processor(s) 1715A-1715N, which may reference vertex or image / texture data stored in memory in addition to the vertex or image / texture data stored in one or more cache(s) 1725A-1725B. In at least one embodiment, one or more MMU(s) 1720A-1720B may be synchronized with other MMUs within a system, including one or more MMUs associated with one or more application processors 1205, image processors 1215, and / or video processors 1220 of Fig. 12, so that each processor 1205-1220 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1730A-1730B enable the graphics processor 1710 to interface with other IP cores within an SoC, either via an internal bus of the SoC or via a direct connection.

[0106] In at least one embodiment, the graphics processor 1740 includes one or more MMU(s) 1720A-1720B, caches 1725A-1725B, and circuit interconnects 1730A-1730B of the graphics processor 1710 of Fig. 17A. In at least one embodiment, graphics processor 1740 includes one or more shader cores 1755A-1755N (e.g., 1755A, 1755B, 1755C, 1755D, 1755E, 1755F through 1755N-1, and 1755N) that provide a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader code implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores may vary.In at least one embodiment, the graphics processor 1740 includes an inter-core task manager 1745 acting as a thread dispatcher to distribute execution threads to one or more shader cores 1755A-1755N and a tiling unit 1758 to accelerate tiling operations for tile-based rendering, in which rendering operations for a scene are divided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0107] Fig. 18A shows a graphics core 1800 in accordance with at least one embodiment. In at least one embodiment, the graphics core 1800 may be included in the graphics processor 1210 of Fig. 12. In at least one embodiment, the graphics core 1800 may be a unified shader core 1755A-1755N as shown in Fig. 17B. In at least one embodiment, the graphics core 1800 includes a shared instruction cache 1802, a texture unit 1818, and a cache / shared memory 1820 common to the execution resources within the graphics core 1800. In at least one embodiment, the graphics core 1800 may include multiple slices 1801A-1801N or partitions for each core, and a graphics processor may include multiple instances of the graphics core 1800. The slices 1801A-1801N may include support logic including a local instruction cache 1804A-1804N, a thread scheduler 1806A-1806N, a thread dispatcher 1808A-1808N, and a set of registers 1810A-1810N.In at least one embodiment, slices 1801A-1801N may include a set of additional functional units ("AFUs") 1812A-1812N, floating-point units ("FPUs") 1814A-1814N, integer arithmetic logic units ("ALUs") 1816-1816N, address calculation units ("ACUs") 1813A-1813N, double-precision floating-point units ("DPFPUs") 1815A-1815N, and matrix processing units ("MPUs") 1817A-1817N. In at least one embodiment, a graphics core 1800 is also referred to as a computing unit.

[0108] In at least one embodiment, the FPUs 1814A-1814N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 1815A-1815N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1816A-1816N can perform variable-precision integer operations at 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1817A-1817N can also be configured for mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, MPUs 1817-1817N may perform a variety of matrix operations to accelerate CUDA programs, including support for accelerated general purpose matrix-matrix multiplication (“GEMM”).In at least one embodiment, the AFUs 1812A-1812N may perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0109] Fig. 18B illustrates a general purpose graphics processing unit ("GPGPU") 1830 in accordance with at least one embodiment. In at least one embodiment, GPGPU 1830 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, GPGPU 1830 can be configured to allow highly parallel computational operations to be performed by an array of GPUs. In at least one embodiment, GPGPU 1830 can be directly connected to other instances of GPGPU 1830 to form a multi-GPU cluster and improve execution time for CUDA programs. In at least one embodiment, GPGPU 1830 includes a host interface 1832 to enable connection to a host processor. In at least one embodiment, host interface 1832 is a PCIe interface.In at least one embodiment, host interface 1832 may be a vendor-specific communications interface or communications structure. In at least one embodiment, GPGPU 1830 receives instructions from a host processor and uses a global scheduler 1834 to distribute the execution threads associated with those instructions among a number of compute clusters 1836A-1836H. In at least one embodiment, compute clusters 1836A-1836H share a cache 1838. In at least one embodiment, cache 1838 may serve as a higher-level cache for caches within compute clusters 1836A-1836H.

[0110] In at least one embodiment, GPGPU 1830 includes memory 1844A-1844B coupled to compute clusters 1836A-1836H via a series of memory controllers 1842A-1842B. In at least one embodiment, memory 1844A-1844B may include various types of memory devices, including DRAM or graphics random access memory, such as synchronous graphics random access memory ("SGRAM"), including graphics double data rate memory ("GDDR").

[0111] In at least one embodiment, the compute clusters 1836A-1836H each include a set of graphics cores, such as the graphics core 1800 of Fig. 18A, which may include multiple types of integer and floating-point logic units capable of performing computational operations with a range of precisions, also suitable for computations associated with CUDA programs. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 1836A-1836H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of floating-point units may be configured to perform 64-bit floating-point operations.

[0112] In at least one embodiment, multiple instances of the GPGPU 1830 can be configured to operate as a compute cluster. The compute clusters 1836A-1836H can implement any technically feasible communication techniques for synchronization and data exchange. In at least one embodiment, multiple instances of the GPGPU 1830 communicate via the host interface 1832. In at least one embodiment, the GPGPU 1830 includes an I / O hub 1839 that couples the GPGPU 1830 to a GPU interconnect 1840 that enables direct connection to other instances of the GPGPU 1830. In at least one embodiment, the GPU interconnect 1840 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 1830.In at least one embodiment, GPU link 1840 is coupled to a high-speed interconnect to send and receive data to other GPGPUs 1830 or parallel processors. In at least one embodiment, multiple GPGPU instances 1830 are located in separate computing systems and communicate via a network interface accessible via host interface 1832. In at least one embodiment, GPU link 1840 may be configured to enable connection to a processor in addition to, or alternatively to, host interface 1832. In at least one embodiment, GPGPU 1830 may be configured to execute a CUDA program.

[0113] Fig. 19A illustrates a parallel processor 1900 in accordance with at least one embodiment. In at least one embodiment, various components of the parallel processor 1900 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits ("ASICs"), or FPGAs.

[0114] In at least one embodiment, parallel processor 1900 includes a parallel processing unit 1902. In at least one embodiment, parallel processing unit 1902 includes an I / O unit 1904 that enables communication with other devices, including other instances of parallel processing unit 1902. In at least one embodiment, I / O unit 1904 can be connected directly to other devices. In at least one embodiment, I / O unit 1904 is connected to other devices via a hub or switch interface, such as storage hub 1905. In at least one embodiment, the connections between storage hub 1905 and I / O unit 1904 form a communication link.In at least one embodiment, the I / O unit 1904 is coupled to a host interface 1906 and a memory crossbar 1916, where the host interface 1906 receives commands to perform processing operations and the memory crossbar 1916 receives commands to perform memory operations.

[0115] In at least one embodiment, when host interface 1906 receives a command buffer via I / O unit 1904, host interface 1906 may direct work operations to a front end 1908 to execute those commands. In at least one embodiment, front end 1908 is coupled to a scheduler 1910 configured to dispatch commands or other work items to a processing array 1912. In at least one embodiment, scheduler 1910 ensures that processing array 1912 is properly configured and in a valid state before dispatching tasks to processing array 1912. In at least one embodiment, scheduler 1910 is implemented via firmware logic executing on a microcontroller.In at least one embodiment, the microcontroller-implemented scheduler 1910 is configurable to perform complex scheduling and work distribution operations at coarse and fine granularity, enabling fast preemption and context switching of threads executing on the processing array 1912. In at least one embodiment, host software may indicate workloads for scheduling on the processing array 1912 via one of several graphics processing doorbells. In at least one embodiment, the workloads may then be automatically distributed across the processing array 1912 by the logic of the scheduler 1910 within a microcontroller that includes the scheduler 1910.

[0116] In at least one embodiment, the processing array 1912 may include up to "N" clusters (e.g., cluster 1914A, cluster 1914B, through cluster 1914N). In at least one embodiment, each cluster 1914A-1914N of the processing array 1912 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1910 may allocate work to the clusters 1914A-1914N of the processing array 1912 using various scheduling and / or work distribution algorithms, which may vary depending on the workload incurred for each type of program or computation. In at least one embodiment, the scheduling may be performed dynamically by the scheduler 1910 or may be assisted in part by compiler logic during compilation of the program logic configured for execution by the processing array 1912.In at least one embodiment, different processing clusters 1914A-1914N of processing array 1912 may be assigned for processing different types of programs or for performing different types of computations.

[0117] In at least one embodiment, processing array 1912 may be configured to perform various types of parallel processing operations. In at least one embodiment, processing array 1912 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing array 1912 may include logic to perform processing tasks, including filtering video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.

[0118] In at least one embodiment, processing array 1912 is configured to perform parallel graphics processing operations. In at least one embodiment, processing array 1912 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing array 1912 may be configured to execute graphics processing-related shader programs, such as vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 1902 may transfer data from system memory via I / O unit 1904 for processing.In at least one embodiment, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1922) during processing and then written back to system memory.

[0119] In at least one embodiment, when the graphics processing unit 1902 is used to perform graphics processing, the scheduler 1910 may be configured to divide a processing load into approximately equal-sized tasks to enable better distribution of graphics processing operations across multiple clusters 1914A-1914N of the processing array 1912. In at least one embodiment, portions of the processing array 1912 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen operations to generate a rendered image for display.In at least one embodiment, intermediate data generated by one or more of clusters 1914A-1914N may be stored in buffers to allow intermediate data to be transferred between clusters 1914A-1914N for further processing.

[0120] In at least one embodiment, processing array 1912 may receive processing tasks to be executed via scheduler 1910, which receives commands defining processing tasks from frontend 1908. In at least one embodiment, the processing tasks may include indices of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how the data is to be processed (e.g., which program is to be executed). In at least one embodiment, scheduler 1910 may be configured to retrieve indices corresponding to the tasks or may receive indices from frontend 1908.In at least one embodiment, the front end 1908 may be configured to ensure that the processing array 1912 is configured to a valid state before initiating a workload specified by incoming command buffers (e.g., batch buffers, push buffers, etc.).

[0121] In at least one embodiment, each of one or more instances of parallel processing unit 1902 may be coupled to parallel processor memory 1922. In at least one embodiment, parallel processor memory 1922 may be accessed via memory crossbar 1916, which may receive memory requests from processing array 1912 as well as from I / O unit 1904. In at least one embodiment, memory crossbar 1916 may access parallel processor memory 1922 via a memory interface 1918. In at least one embodiment, memory interface 1918 may include a plurality of partition units (e.g., partition unit 1920A, partition unit 1920B, through partition unit 1920N), each of which may be coupled to a portion (e.g., a memory unit) of parallel processor memory 1922.In at least one embodiment, a number of partition units 1920A-1920N is configured to be equal to a number of storage units, such that a first partition unit 1920A has a corresponding first storage unit 1924A, a second partition unit 1920B has a corresponding storage unit 1924B, and an Nth partition unit 1920N has a corresponding Nth storage unit 1924N. In at least one embodiment, a number of partition units 1920A-1920N may not be equal to a number of storage devices.

[0122] In at least one embodiment, memory units 1924A-1924N may include various types of memory devices, including DRAM or graphics random access memory, such as SGRAM, including GDDR memory. In at least one embodiment, memory units 1924A-1924N may also include 3D stacks, including but not limited to high-width memories ("HBM"). In at least one embodiment, rendering targets, such as frame buffers or texture maps, may be stored across memory units 1924A-1924N so that partition units 1920A-1920N can write portions of each rendering target in parallel to efficiently utilize the available bandwidth of parallel processor memory 1922.In at least one embodiment, a local instance of parallel processor memory 1922 may be eliminated in favor of a unified memory design that utilizes system memory in conjunction with the local cache memory.

[0123] In at least one embodiment, each of the clusters 1914A-1914N of the processing array 1912 can process data written to each of the processing units 1924A-1924N in the parallel processor memory 1922. In at least one embodiment, the memory crossbar 1916 can be configured to transfer an output of each cluster 1914A-1914N to any partition unit 1920A-1920N or to another cluster 1914A-1914N that can perform additional processing operations on an output. In at least one embodiment, each cluster 1914A-1914N can communicate with the memory interface 1918 via the memory crossbar 1916 to read from or write to various external devices.In at least one embodiment, the memory crossbar 1916 includes a connection to the memory interface 1918 to communicate with the I / O unit 1904, as well as a connection to a local instance of the parallel processor memory 1922, which enables the processing units in different clusters 1914A-1914N to communicate with system memory or other memory not local to the parallel processing unit 1902. In at least one embodiment, the memory crossbar 1916 may use virtual channels to separate traffic flows between clusters 1914A-1914N and partition units 1920A-1920N.

[0124] In at least one embodiment, multiple instances of the parallel processing unit 1902 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 1902 may be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 1902 may include higher-precision floating-point units compared to other instances.In at least one embodiment, systems including one or more instances of the parallel processing unit 1902 or the parallel processor 1900 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0125] Fig. 19B shows a processing cluster 1994 in accordance with at least one embodiment. In at least one embodiment, the processing cluster 1994 is included in a parallel processing unit. In at least one embodiment, the processing cluster 1994 is one of the processing clusters 1914A-1914N of Fig. 19. In at least one embodiment, the processing cluster 1994 may be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction, multiple data (SIMD) instruction issuance techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction, multiple thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads using a common instruction unit configured to issue instructions to a number of processing engines within each processing cluster 1994.

[0126] In at least one embodiment, the operation of the processing cluster 1994 may be controlled by a pipeline manager 1932 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1932 receives instructions from the scheduler 1910 of Fig. 19 and manages the execution of these instructions via a graphics multiprocessor 1934 and / or a texture unit 1936. In at least one embodiment, the graphics multiprocessor 1934 is an exemplary example of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included in the processing cluster 1994. In at least one embodiment, the processing cluster 1994 may include one or more instances of the graphics multiprocessor 1934. In at least one embodiment, the graphics multiprocessor 1934 may process data, and a data crossbar 1940 may be used to distribute the processed data to one of several possible destinations, including other shader units.In at least one embodiment, the pipeline manager 1932 may facilitate the distribution of processed data by specifying destinations for processed data to be distributed across the data crossbar 1940.

[0127] In at least one embodiment, each graphics multiprocessor 1934 within the processing cluster 1994 may include an identical set of functional execution logic (e.g., arithmetic logic units, load / store units ("LSUs"), etc.). In at least one embodiment, the functional execution logic may be configured in a pipeline in which new instructions may be issued before previous instructions complete. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same hardware with functional units may be used to perform different operations, and any combination of functional units may be present.

[0128] In at least one embodiment, the instructions transferred to the processing cluster 1994 form a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program with different input data. In at least one embodiment, each thread within a thread group may be assigned to a different engine within the graphics multiprocessor 1934. In at least one embodiment, a thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 1934. In at least one embodiment, when a thread group includes fewer threads than a number of processing engines, one or more of the processing engines may be idle during the cycles in which that thread group is processing.In at least one embodiment, a thread group may also include more threads than the number of processing engines within the graphics multiprocessor 1934. In at least one embodiment, if a thread group includes more threads than the number of processing engines in the graphics multiprocessor 1934, processing may occur in consecutive clock cycles. In at least one embodiment, multiple thread groups may execute concurrently on the graphics multiprocessor 1934.

[0129] In at least one embodiment, the graphics multiprocessor 1934 includes an internal cache for performing load and store operations. In at least one embodiment, the graphics multiprocessor 1934 may forgo an internal cache and utilize a cache (e.g., L1 cache 1948) within the processing cluster 1994. In at least one embodiment, each graphics multiprocessor 1934 also has access to Level 2 ("L2") caches within partition units (e.g., partition units 1920A-1920N of Fig. 19A) that are shared by all processing clusters 1994 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1934 can also access off-chip global memory, which can include one or more of the parallel processor local memories and / or system memories. In at least one embodiment, any memory external to parallel processing unit 1902 can be used as global memory. In at least one embodiment, processing cluster 1994 includes multiple instances of graphics multiprocessor 1934, which can share common instructions and data, which can be stored in L1 cache 1948.

[0130] In at least one embodiment, each processing cluster 1994 may include an MMU 1945 configured to translate virtual addresses into physical addresses. In at least one embodiment, one or more instances of the MMU 1945 may reside within the memory interface 1918 of Fig. 19. In at least one embodiment, the MMU 1945 includes a set of page table entries ("PTEs") used to map a virtual address to a physical address of a tile, and optionally a cache line index. In at least one embodiment, the MMU 1945 may include address translation lookaside buffers ("TLBs") or caches, which may be located in the graphics multiprocessor 1934 or the L1 cache 1948 or the processing cluster 1994. In at least one embodiment, a physical address is processed to distribute data access locality at the surface to enable efficient interleaving of requests between the partition units. In at least one embodiment, a cache line index may be used to determine whether a request for a cache line is a hit or miss.

[0131] In at least one embodiment, processing cluster 1994 may be configured such that each graphics multiprocessor 1934 is coupled to a texture unit 1936 to perform texture mapping operations, such as determining texture pattern positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within graphics multiprocessor 1934 and fetched from an L2 cache, local parallel processor memory, or system memory as needed.In at least one embodiment, each graphics multiprocessor 1934 outputs a processed task to the data crossbar 1940 to provide the processed task to another processing cluster 1994 for further processing via the memory crossbar 1916, or to store the processed task in an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, a pre-raster operations unit ("preROP") 1942 is configured to receive data from the graphics multiprocessor 1934 and forward data to ROP units, which may be arranged with partition units as described herein (e.g., partition units 1920A-1920N of FIG. Fig. 19). In at least one embodiment, the PreROP 1942 may perform optimizations for color mixing, organizing pixel color data, and address translations.

[0132] Fig. 19C shows a graphics multiprocessor 1996 in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1996 is the graphics multiprocessor 1934 of Fig. 19B. In at least one embodiment, graphics multiprocessor 1996 is coupled to pipeline manager 1932 of processing cluster 1994. In at least one embodiment, graphics multiprocessor 1996 has an execution pipeline including, among other things, an instruction cache 1952, an instruction unit 1954, an address mapping unit 1956, a register file 1958, one or more GPGPU cores 1962, and one or more LSUs 1966. GPGPU cores 1962 and LSUs 1966 are coupled to cache memory 1972 and shared memory 1970 via a memory and cache interconnect 1968.

[0133] In at least one embodiment, instruction cache 1952 receives a stream of instructions to be executed from pipeline manager 1932. In at least one embodiment, the instructions are cached in instruction cache 1952 and forwarded for execution by instruction unit 1954. In at least one embodiment, instruction unit 1954 may dispatch instructions as thread groups (e.g., warps), with each thread of a thread group assigned to a different execution unit within GPGPU core 1962. In at least one embodiment, an instruction may access a local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, address mapping unit 1956 may be used to translate addresses in a unified address space into a unique memory address accessible by LSUs 1966.

[0134] In at least one embodiment, register file 1958 provides a set of registers for functional units of graphics multiprocessor 1996. In at least one embodiment, register file 1958 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 1962, LSUs 1966) of graphics multiprocessor 1996. In at least one embodiment, register file 1958 is partitioned among individual functional units such that each functional unit is allocated its own section of register file 1958. In at least one embodiment, register file 1958 is partitioned among different thread groups executed by graphics multiprocessor 1996.

[0135] In at least one embodiment, the GPGPU cores 1962 may each include FPUs and / or integer ALUs used to execute instructions of the graphics multiprocessor 1996. The GPGPU cores 1962 may be similar or different in architecture. In at least one embodiment, a first portion of the GPGPU cores 1962 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU cores 1962 includes a double-precision FPU. In at least one embodiment, the FPUs may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1996 may additionally include one or more fixed-function or special-function units to perform specific functions such as copying rectangles or blending pixels.In at least one embodiment, one or more of the GPGPU cores 1962 may also include fixed functional logic or special functional logic.

[0136] In at least one embodiment, GPGPU cores 1962 include SIMD logic capable of applying a single instruction to multiple data sets. In at least one embodiment, GPGPU cores 1962 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores 1962 can be generated at compile time by a shader compiler or automatically during the execution of programs written and compiled for SPMD or SIMT architectures. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logic unit.

[0137] In at least one embodiment, the memory and cache interconnect 1968 is an interconnect network that connects each functional unit of the graphics multiprocessor 1996 to the register file 1958 and the shared memory 1970. In at least one embodiment, the memory and cache interconnect 1968 is a crossbar interconnect that enables the LSU 1966 to perform load and store operations between the shared memory 1970 and the register file 1958. In at least one embodiment, the register file 1958 may operate at the same frequency as the GPGPU cores 1962, such that data transfer between the GPGPU cores 1962 and the register file 1958 has very low latency. In at least one embodiment, the shared memory 1970 may be used to enable communication between threads executing on functional units within the graphics multiprocessor 1996.For example, in at least one embodiment, cache 1972 may be used as a data cache to cache texture data transferred between functional units and texture unit 1936. In at least one embodiment, shared memory 1970 may also be used as a programmatic cache. In at least one embodiment, threads executing on GPGPU cores 1962 may programmatically store data in shared memory in addition to the automatically cached data stored in cache 1972.

[0138] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to host processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, a GPU may be communicatively coupled to the host processor cores via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, a graphics processor may be integrated on the same package or die as cores and communicate with the cores via a processor bus / interconnect located within a package or die.In at least one embodiment, regardless of how a GPU is connected, the processor cores can assign work to the GPU in the form of sequences of instructions contained in a WD. In at least one embodiment, the GPU then uses special circuitry / logic to efficiently process these instructions.

[0139] Fig. 20 shows a graphics processor 2000 according to at least one embodiment. In at least one embodiment, the graphics processor 2000 includes a ring interconnect 2002, a pipelined front end 2004, a media engine 2037, and graphics cores 2080A-2080N. In at least one embodiment, the ring interconnect 2002 couples the graphics processor 2000 to other processing units, including other graphics processors or one or more general-purpose processing cores. In at least one embodiment, the graphics processor 2000 is one of many processors integrated into a multi-core processing system.

[0140] In at least one embodiment, graphics processor 2000 receives batches of instructions via a ring interconnect 2002. In at least one embodiment, the incoming instructions are interpreted by an instruction streamer 2003 in pipeline front end 2004. In at least one embodiment, graphics processor 2000 includes scalable execution logic for performing 3D geometry processing and media processing via graphics core(s) 2080A-2080N. In at least one embodiment, instruction streamer 2003 provides instructions to geometry pipeline 2036 for 3D geometry processing instructions. In at least one embodiment, instruction streamer 2003 provides instructions to a video front end 2034 coupled to a media engine 2037 for at least some media processing instructions.In at least one embodiment, the media engine 2037 includes a video quality engine ("VQE") 2030 for post-processing videos and images and a multi-format encode / decode ("MFX") engine 2033 to enable hardware-accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2036 and the media engine 2037 each generate execution threads for threaded execution resources provided by at least one graphics core 2080A.

[0141] In at least one embodiment, graphics processor 2000 includes scalable threaded execution resources with modular graphics cores 2080A-2080N (sometimes referred to as core slices), each having a plurality of subcores 2050A-2080N, 2060A-2060N (sometimes referred to as a core slice). In at least one embodiment, graphics processor 2000 may include any number of graphics cores 2080A-2080N. In at least one embodiment, graphics processor 2000 includes a graphics core 2080A having at least a first subcore 2050A and a second subcore 2060A. In at least one embodiment, graphics processor 2000 is a low-power processor with a single subcore (e.g., subcore 2050A). In at least one embodiment, the graphics processor 2000 includes a plurality of graphics cores 2080A-2080N, each including a set of first sub-cores 2050A-2050N and a set of second sub-cores 2060A-2060N.In at least one embodiment, each subcore in the first subcores 2050A-2050N includes at least a first set of execution units ("EUs") 2052A-2052N and media / texture units 2054A-2054N. In at least one embodiment, each subcore in the second subcores 2060A-2060N includes at least a second group of execution units 2062A-2062N and samplers 2064A-2064N. In at least one embodiment, each subcore 2050A-2050N, 2060A-2060N shares a set of shared resources 2070A-2070N. In at least one embodiment, the shared resources 2070 include a shared cache and pixel operation logic.

[0142] Fig. 21 shows a processor 2100 in accordance with at least one embodiment. In at least one embodiment, the processor 2100 may include, without limitation, logic circuitry for executing instructions. In at least one embodiment, the processor 2100 may execute instructions including x86 instructions, ARM instructions, special instructions for ASICs, etc. In at least one embodiment, the processor 2110 may include registers for storing packed data, such as 64-bit wide MMX™ registers in microprocessors employing MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, which are available in both integer and floating-point form, may operate on packed data elements accompanying SIMD and Streaming SIMD Extensions ("SSE") instructions.In at least one embodiment, 128-bit XMM registers related to SSE2, SSE3, SSE4, AVX, or beyond technologies (commonly referred to as "SSEx") may accommodate such packed data operands. In at least one embodiment, processors 2110 may execute instructions to accelerate CUDA programs.

[0143] In at least one embodiment, processor 2100 includes an in-order front-end ("front-end") 2101 for fetching instructions to be executed and preparing instructions to be used later in the processor pipeline. In at least one embodiment, front-end 2101 may include multiple units. In at least one embodiment, an instruction prefetcher 2126 fetches instructions from memory and passes them to an instruction decoder 2128, which in turn decodes or interprets instructions. In at least one embodiment, instruction decoder 2128 decodes a received instruction into one or more operations, referred to as "micro-instructions" or "micro-operations" (also called "microOps" or "uOps"), for execution.In at least one embodiment, instruction decoder 2128 decomposes the instruction into opcode and corresponding data and control fields that can be used by the microarchitecture to execute operations. In at least one embodiment, a trace cache 2130 may assemble decoded uOps into program-ordered sequences or traces in a uOps queue 2134 for execution. In at least one embodiment, when trace cache 2130 encounters a complex instruction, a microcode ROM 2132 provides uOps needed to execute an operation.

[0144] In at least one embodiment, some instructions may be converted into a single micro-op, while others may require multiple micro-ops to perform a complete operation. In at least one embodiment, if more than four micro-ops are required to execute an instruction, instruction decoder 2128 may access microcode ROM 2132 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops for processing in instruction decoder 2128. In at least one embodiment, an instruction may be stored in microcode ROM 2132 if a number of micro-ops are required to perform the operation.In at least one embodiment, trace cache 2130 refers to a programmable logic array ("PLA") as an entry point to determine a correct microinstruction pointer for reading microcode sequences to complete one or more instructions from microcode ROM 2132. In at least one embodiment, after microcode ROM 2132 finishes sequencing micro-ops for an instruction, machine front-end 2101 may resume fetching micro-ops from trace cache 2130.

[0145] In at least one embodiment, the out-of-order execution engine 2103 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic includes a series of buffers to smooth and reorder the flow of instructions to optimize performance as they traverse a pipeline and are scheduled for execution. The out-of-order execution engine 2103 includes, among other things, an allocator / register renamer 2140, a memory UOps queue 2142, an integer / floating-point UOps queue 2144, a memory scheduler 2146, a fast scheduler 2102, a slow / general FP scheduler 2104, and a simple FP scheduler 2106.In at least one embodiment, the fast scheduler 2102, the slow / general floating-point scheduler 2104, and the simple floating-point scheduler 2106 are also collectively referred to herein as "uOps schedulers 2102, 2104, 2106." The allocator / register renamer 2140 allocates machine buffers and resources required by each uOps for its execution. In at least one embodiment, the allocator / register renamer 2140 renames logical registers to entries in a register file. In at least one embodiment, the allocator / register renamer 2140 also assigns each uOps an entry in one of two uOps queues, the memory uOps queue 2142 for memory operations and the integer / floating point uOps queue 2144 for non-memory operations, which precede the memory scheduler 2146 and the uOps schedulers 2102, 2104, 2106.In at least one embodiment, schedulers 2102, 2104, 2106 determine the readiness of a uOp to execute based on the readiness of its dependent input register operand sources and the availability of the execution resources required by the uOps to complete their operation. In at least one embodiment, fast scheduler 2102 may schedule one in each half of the main clock cycle, while slow / general floating-point scheduler 2104 and simple floating-point scheduler 2106 may schedule one per main clock cycle of the processor. In at least one embodiment, schedulers 2102, 2104, 2106 arbitrate for dispatch ports to schedule uOps for execution.

[0146] In at least one embodiment, execution block 2111 includes, without limitation, an integer register file / bypass network 2108, a floating-point register file / bypass network ("FP register file / bypass network") 2110, address generation units ("AGUs") 2112 and 2114, fast ALUs 2116 and 2118, a slow ALU 2120, a floating-point ALU ("FP") 2122, and a floating-point shift unit ("FP shift unit") 2124. In at least one embodiment, integer register file / bypass network 2108 and floating-point register file / bypass network 2110 are also referred to herein as "register files 2108, 2110." In at least one embodiment, the AGUSs 2112 and 2114, the fast ALUs 2116 and 2118, the slow ALU 2120, the floating-point ALU 2122, and the floating-point shift unit 2124 are also referred to herein as "execution units 2112, 2114, 2116, 2118, 2120, 2122, and 2124."In at least one embodiment, an execution block may include, without limitation, any number (including zero) and type of register files, bypass networks, address generation units, and execution units in any combination.

[0147] In at least one embodiment, register files 2108, 2110 may be located between uOps schedulers 2102, 2104, 2106 and execution units 2112, 2114, 2116, 2118, 2120, 2122, and 2124. In at least one embodiment, integer register file / bypass network 2108 performs integer operations. In at least one embodiment, floating-point register file / bypass network 2110 performs floating-point operations. In at least one embodiment, each of register files 2108, 2110 may include, without limitation, a bypass network that can bypass just-completed results that have not yet been written to the register file or forward them to new dependent uOps. In at least one embodiment, register files 2108, 2110 may exchange data with each other.In at least one embodiment, the integer register / bypass network 2108 may include, without limitation, two separate register files: one low-order data register file of 32 bits and a second high-order data register file of 32 bits. In at least one embodiment, the register file / bypass network 2110 may include, without limitation, 128-bit wide entries, since floating-point instructions typically have operands 64 to 128 bits wide.

[0148] In at least one embodiment, execution units 2112, 2114, 2116, 2118, 2120, 2122, 2124 may execute instructions. In at least one embodiment, register files 2108, 2110 store integer and floating-point data operand values ​​required for microinstruction execution. In at least one embodiment, processor 2100 may include, without limitation, any number and combination of execution units 2112, 2114, 2116, 2118, 2120, 2122, 2124. In at least one embodiment, floating-point ALU 2122 and floating-point shift unit 2124 may perform floating-point, MMX, SIMD, AVX, and SSE or other operations. In at least one embodiment, the floating-point ALU 2122 may include, without limitation, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations.In at least one embodiment, instructions involving a floating-point value may be processed using floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2116, 2118. In at least one embodiment, fast ALUs 2116, 2118 may perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 2120, as slow ALU 2120 may include, without limitation, integer execution hardware for long-latency operations, such as a multiplier, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by AGUs 2112, 2114.In at least one embodiment, the fast ALU 2116, the fast ALU 2118, and the slow ALU 2120 can perform integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 2116, the fast ALU 2118, and the slow ALU 2120 can be implemented to support a variety of data bit sizes, including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating-point ALU 2122 and the floating-point shift unit 2124 can be implemented to support a number of operands having bits of different widths. In at least one embodiment, the floating-point ALU 2122 and the floating-point shift unit 2124 can operate with 128-bit packed data operands in conjunction with SIMD and multimedia instructions.

[0149] In at least one embodiment, the uOps schedulers 2102, 2104, 2106 dispatch dependent operations before the parent load completes execution. In at least one embodiment where uOps may be speculatively scheduled and executed in the processor 2100, the processor 2100 may also include logic to handle memory misses. In at least one embodiment, when a data load fails in a data cache, there may be dependent operations in the pipeline that have left a scheduler with temporarily incorrect data. In at least one embodiment, a replay mechanism tracks instructions that use incorrect data and reexecutes them. In at least one embodiment, dependent operations may need to be replayed while independent operations are allowed to complete.In at least one embodiment, schedulers and rendering mechanisms of at least one embodiment of a processor may also be configured to intercept instruction sequences for text string comparison operations.

[0150] In at least one embodiment, the term "registers" may refer to onboard memory locations of the processor that may be used as part of instructions to identify operands. In at least one embodiment, the registers may be those that may be used from outside a processor (from a programmer's perspective). In at least one embodiment, the registers may not be limited to a particular circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein.In at least one embodiment, the registers described herein may be implemented by circuitry within a processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, 32-bit integer data is stored in integer registers. A register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.

[0151] Fig. 22 shows a processor 2200 in accordance with at least one embodiment. In at least one embodiment, the processor 2200 includes, without limitation, one or more processor cores ("cores") 2202A-2202N, an integrated memory controller 2214, and an integrated graphics processor 2208. In at least one embodiment, the processor 2200 may include additional cores, up to and including the additional processor core 2202N represented by dashed boxes. In at least one embodiment, each of the processor cores 2202A-2202N includes one or more internal cache units 2204A-2204N. In at least one embodiment, each processor core also has access to one or more shared cache units 2206. In at least one embodiment, one or more processor cores 2202A-2202N are referred to as one or more computing units.

[0152] In at least one embodiment, the internal cache units 2204A-2204N and the shared cache units 2206 represent a cache hierarchy within the processor 2200. In at least one embodiment, the cache units 2204A-2204N may include at least one level of instruction and data cache within each processor core and one or more levels of shared mid-level cache, such as L2, L3, Level 4 ("L4"), or other levels of cache, with a highest level of cache prior to external memory classified as LLC. In at least one embodiment, the cache coherence logic maintains coherence between different cache units 2206 and 2204A-2204N.

[0153] In at least one embodiment, the processor 2200 may also include a set of one or more bus control units 2216 and a system agent core 2210. In at least one embodiment, one or more bus control units 2216 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2210 provides management functions for various processor components. In at least one embodiment, the system agent core 2210 includes one or more integrated memory controllers 2214 for managing access to various external storage devices (not shown).

[0154] In at least one embodiment, one or more of the processor cores 2202A-2202N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 2210 includes components for coordinating and operating the processor cores 2202A-2202N during multi-threaded processing. In at least one embodiment, the system agent core 2210 may additionally include a power control unit ("PCU") that includes logic and components for regulating one or more power states of the processor cores 2202A-2202N and the graphics processor 2208.

[0155] In at least one embodiment, processor 2200 additionally includes graphics processor 2208 for performing graphics processing operations. In at least one embodiment, graphics processor 2208 couples to shared cache units 2206 and system agent core 2210, which includes one or more integrated memory controllers 2214. In at least one embodiment, system agent core 2210 also includes a display controller 2211 for driving the output of the graphics processor to one or more coupled displays. In at least one embodiment, display controller 2211 may also be a separate module connected to graphics processor 2208 via at least one interconnect, or it may be integrated into graphics processor 2208.

[0156] In at least one embodiment, a ring interconnect 2212 is used to couple internal components of processor 2200. In at least one embodiment, an alternative interconnect may be used, such as a point-to-point connection, a switch connection, or other techniques. In at least one embodiment, graphics processor 2208 is connected to ring interconnect 2212 via an I / O connection 2213.

[0157] In at least one embodiment, I / O interconnect 2213 represents at least one of several types of I / O interconnects, including a chassis-mounted I / O interconnect that facilitates communication between various processor components and an embedded high-performance memory module 2218, such as an eDRAM module. In at least one embodiment, each of the processor cores 2202A-2202N and the graphics processor 2208 utilize embedded memory modules 2218 as a shared LLC.

[0158] In at least one embodiment, processor cores 2202A-2202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 2202A-2202N are heterogeneous with respect to ISA, where one or more processor cores 2202A-2202N execute a common instruction set, while one or more other cores of processor cores 2202A-2202N execute a subset of a common instruction set or a different instruction set. In at least one embodiment, processor cores 2202A-2202N are heterogeneous with respect to microarchitecture, where one or more cores having relatively higher power consumption are coupled with one or more cores having lower power consumption. In at least one embodiment, processor 2200 can be implemented on one or more chips or as an integrated circuit (SoC).

[0159] Fig. 23 shows a graphics processor core 2300 in accordance with at least one of the described embodiments. In at least one embodiment, the graphics processor core 2300 is included in a graphics core array. In at least one embodiment, the graphics processor core 2300, sometimes referred to as a core slice, may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2300 is exemplary of a graphics core slice, and a graphics processor as described herein may include multiple graphics core slices based on targeted power and performance levels. In at least one embodiment, each graphics core 2300 may include a fixed functional block 2330 coupled to a plurality of sub-cores 2301A-2301F, also referred to as slices, which include modular blocks of general-purpose and fixed-function logic.

[0160] In at least one embodiment, functional block 2330 includes a geometry / fixed function pipeline 2336 that may be shared by all subcores in graphics processor 2300, for example, in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry / fixed function pipeline 2336 includes a 3D fixed function pipeline, a video front-end unit, a thread spawner and thread dispatcher, and a unified return buffer manager that manages unified return buffers.

[0161] In at least one embodiment, fixed functional block 2330 also includes a graphics SoC interface 2337, a graphics microcontroller 2338, and a media pipeline 2339. Graphics SoC interface 2337 provides an interface between graphics core 2300 and other processor cores within an SoC integrated circuit. In at least one embodiment, graphics microcontroller 2338 is a programmable subprocessor that can be configured to manage various functions of graphics processor 2300, including thread dispatch, scheduler, and preemption. In at least one embodiment, media pipeline 2339 includes logic to facilitate decoding, encoding, preprocessing, and / or postprocessing of multimedia data, including image and video data.In at least one embodiment, the media pipeline 2339 implements media operations via requests to compute or sensing logic within the subcores 2301-2301F.

[0162] In at least one embodiment, the SoC interface 2337 enables the graphics core 2300 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within an SoC, including memory hierarchy elements such as shared LLC memory, system RAM, and / or embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2337 may also enable communication with fixed devices within an SoC, such as camera imaging pipelines, and enables the use and / or implementation of global memory atomics that may be shared between graphics cores 2300 and CPUs within an SoC.In at least one embodiment, the SoC interface 2337 may also implement power management controls for the graphics core 2300 and enable an interface between a clock domain of the graphics core 2300 and other clock domains within an SoC. In at least one embodiment, the SoC interface 2337 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to deliver commands and instructions to each of one or more graphics cores within a graphics processor. In at least one embodiment, commands and instructions may be sent to the media pipeline 2339 when media operations are to be performed, or to a geometry and fixed function pipeline (e.g., geometry and fixed function pipeline 2336, geometry and fixed function pipeline 2314) when graphics processing operations are to be performed.

[0163] In at least one embodiment, graphics microcontroller 2338 may be configured to perform various scheduling and management tasks for graphics core 2300. In at least one embodiment, graphics microcontroller 2338 may perform graphics and / or compute workload scheduling on various parallel engines within execution unit (EU) arrays 2302A-2302F, 2304A-2304F within subcores 2301A-2301F. In at least one embodiment, host software executing on a CPU core of an SoC including graphics core 2300 may submit workloads to one of several graphics processor doorbells, which invokes a scheduling operation on an appropriate graphics engine.In at least one embodiment, the scheduling operations include determining the next workload to execute, submitting a workload to an instruction streamer, preempting existing workloads running on an engine, monitoring the progress of a workload, and notifying host software when a workload completes. In at least one embodiment, the graphics microcontroller 2338 may also facilitate low-power or idle states for the graphics core 2300 by providing the graphics core 2300 with the ability to save and restore registers within the graphics core 2300 across low-power state transitions independent of an operating system and / or graphics driver software on a system.

[0164] In at least one embodiment, the graphics core 2300 may include more or fewer sub-cores than the shown sub-cores 2301A-2301F, up to N modular sub-cores. For each set of N sub-cores, the graphics core 2300 may also include, in at least one embodiment, shared function logic 2310, shared and / or cache memory 2312, a geometry / fixed function pipeline 2314, and additional fixed function logic 2316 to accelerate various graphics and computational processing operations. In at least one embodiment, the shared function logic 2310 may include logical units (e.g., sampler, math, and / or inter-thread communication logic) that may be shared by all N sub-cores within the graphics core 2300.Shared and / or cache memory 2312 may be an LLC for N subcores 2301A-2301F within graphics core 2300 and may also serve as shared memory accessible by multiple subcores. In at least one embodiment, geometry / fixed function pipeline 2314 may be included in functional block 2330 instead of geometry / fixed function pipeline 2336 and may include the same or similar logic units.

[0165] In at least one embodiment, the graphics core 2300 includes additional fixed function logic 2316, which may include various fixed function accelerator logic for use by the graphics core 2300. In at least one embodiment, the additional fixed function logic 2316 includes an additional geometry pipeline for use in position-dependent shading. In position-dependent shading, there are at least two geometry pipelines: a full geometry pipeline within the geometry and fixed function pipeline 2316, 2336, and a cull pipeline, an additional geometry pipeline that may be included within the additional fixed function logic 2316. In at least one embodiment, the cull pipeline is a stripped-down version of a full geometry pipeline.In at least one embodiment, a full pipeline and a cull pipeline may execute different instances of an application, each instance having its own context. In at least one embodiment, position-dependent shading may hide long cull runs of discarded triangles, allowing shading to complete sooner in some cases. For example, in at least one embodiment, the cull pipeline logic within the additional fixed function logic 2316 may execute position shaders in parallel with a main application and generally generates critical results faster than a full pipeline because a cull pipeline retrieves and shades position attributes of vertices without performing rasterization and rendering pixels into a frame buffer.In at least one embodiment, a cull pipeline may use the generated critical results to compute visibility information for all triangles, regardless of whether those triangles are culled. In at least one embodiment, a full pipeline (which in this case may be referred to as a replay pipeline) may consume visibility information to skip rejected triangles and shade only visible triangles, which are ultimately passed to a rasterization phase.

[0166] In at least one embodiment, the additional fixed function logic 2316 may also include general processing acceleration logic, such as fixed function matrix multiplication logic to accelerate CUDA programs.

[0167] In at least one embodiment, each graphics core 2301A-2301F includes a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-cores 2301A-2301F include a plurality of EU arrays 2302A-2302F, 2304A-2304F, a thread dispatcher and inter-thread communication (“TD / IC” logic) 2303A-2303F, a 3D sampler (e.g., texture) 2305A-2305F, a media sampler 2306A-2306F, a shader processor 2307A-2307F, and a shared local memory (“SLM”) 2308A-2308F.The EU arrays 2302A-2302F, 2304A-2304F each include a plurality of execution units, which are GPGPUs capable of performing floating-point and integer / fixed-point logic operations in service of a graphics, media, or compute operation, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2303A-2303F performs local thread dispatch and thread control operations for execution units within a subcore and facilitates communication between threads executing on execution units of a subcore. In at least one embodiment, the 3D sampler 2305A-2305F can read texture or other 3D graphics data into memory. In at least one embodiment, the 3D sampler may read texture data differently based on a configured state and a texture format associated with a particular texture.In at least one embodiment, the media sampler 2306A-2306F may perform similar read operations based on a type and format associated with the media data. In at least one embodiment, each graphics core 2301A-2301F may alternately include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each of the subcores 2301A-2301F may utilize the shared local memory 2308A-2308F within each subcore to enable threads executing within a thread group to execute using a common pool of on-chip memory.

[0168] Fig. 24 shows a parallel processing unit ("PPU") 2400 in accordance with at least one embodiment. In at least one embodiment, the PPU 2400 is configured with machine-readable code that, when executed by the PPU 2400, causes the PPU 2400 to perform some or all of the processes and techniques described herein. In at least one embodiment, the PPU 2400 is a multi-threaded processor implemented on one or more integrated circuits that employs multithreading as a latency hiding technique to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) on multiple threads in parallel. In at least one embodiment, a thread refers to a thread of execution and is an instantiation of a set of instructions configured for execution by the PPU 2400.In at least one embodiment, PPU 2400 is a GPU configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device, such as an LCD device. In at least one embodiment, PPU 2400 is used to perform computations such as linear algebra operations and machine learning operations. Fig. 24 shows an example parallel processor for illustration only and should be understood as a non-limiting example of a processor architecture that may be implemented in at least one embodiment.

[0169] In at least one embodiment, one or more PPUs 2400 are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, one or more PPUs 2400 are configured to accelerate CUDA programs. In at least one embodiment, the PPU 2400 includes, without limitation, an I / O unit 2406, a front-end unit 2410, a scheduler unit 2412, a work distribution unit 2414, a hub 2416, a crossbar (“Xbar”) 2420, one or more general processing clusters (“GPCs”) 2418, and one or more partition units (“memory partition units”) 2422. In at least one embodiment, the PPU 2400 is connected to a host processor or other PPUs 2400 via one or more high-speed GPU interconnects (“GPU interconnects”) 2408.In at least one embodiment, the PPU 2400 is connected to a host processor or other peripheral devices via a system bus or interconnect 2402. In at least one embodiment, the PPU 2400 is connected to local memory, including one or more memory devices ("memory") 2404. In at least one embodiment, the memory devices 2404 include, without limitation, one or more dynamic random access memory (DRAM) devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as high-bandwidth memory ("HBM") subsystems, with multiple DRAM chips stacked within each device.

[0170] In at least one embodiment, the high-speed GPU interconnect 2408 may refer to a wired multi-lane communication link used by systems including one or more PPUs 2400 in combination with one or more CPUs and supporting cache coherency between PPUs 2400 and CPUs, as well as CPU mastering. In at least one embodiment, data and / or commands are transferred via the high-speed GPU interconnect 2408 through the hub 2416 to / from other units of the PPU 2400, such as one or more copy engines, video encoders, video decoders, power management units, and other components included in Fig. 24 may not be shown explicitly.

[0171] In at least one embodiment, the I / O unit 2406 is configured to receive communications (e.g., commands, data) from a host processor (in Fig. 24 not shown) over the system bus 2402. In at least one embodiment, the I / O unit 2406 communicates with the host processor directly over the system bus 2402 or through one or more intermediary devices, such as a memory bridge. In at least one embodiment, the I / O unit 2406 may communicate with one or more other processors, such as one or more PPUs 2400, over the system bus 2402. In at least one embodiment, the I / O unit 2406 implements a PCIe interface for communicating over a PCIe bus. In at least one embodiment, the I / O unit 2406 implements interfaces for communicating with external devices.

[0172] In at least one embodiment, the I / O unit 2406 decodes packets received over the system bus 2402. In at least one embodiment, at least some packets represent commands configured to cause the PPU 2400 to perform various operations. In at least one embodiment, the I / O unit 2406 transmits decoded commands to various other units of the PPU 2400 as specified by the commands. In at least one embodiment, commands are transmitted to the front-end unit 2410 and / or to the hub 2416 or other units of the PPU 2400, such as one or more copy engines, a video encoder, a video decoder, a power management unit, etc. (in Fig. 24). In at least one embodiment, the I / O unit 2406 is configured to direct communication between and among various logical units of the PPU 2400.

[0173] In at least one embodiment, a program executed by the host processor encodes an instruction stream in a buffer that provides workloads to the PPU 2400 for processing. In at least one embodiment, a workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) to both a processor and the PPU 2400—a host interface may be configured to access the buffer in system memory coupled to the system bus 2402 via memory requests transmitted from the I / O device 2406 over the system bus 2402.In at least one embodiment, a host processor writes an instruction stream to a buffer and then transfers a pointer to the beginning of the instruction stream to the PPU 2400, so that the front-end unit 2410 receives pointers to one or more instruction streams and manages one or more instruction streams, reading instructions from the instruction streams and forwarding instructions to various units of the PPU 2400.

[0174] In at least one embodiment, the front-end unit 2410 is coupled to the scheduler unit 2412, which configures various GPCs 2418 to process tasks defined by one or more instruction streams. In at least one embodiment, the scheduler unit 2412 is configured to track state information related to various tasks managed by the scheduler unit 2412, where the state information may indicate which of the GPCs 2418 a task is assigned to, whether the task is active or inactive, a priority level associated with the task, and so on. In at least one embodiment, the scheduler unit 2412 manages the execution of a plurality of tasks on one or more of the GPCs 2418.

[0175] In at least one embodiment, the scheduler unit 2412 is coupled to the work distribution unit 2414, which is configured to distribute tasks for execution among the GPCs 2418. In at least one embodiment, the work distribution unit 2414 tracks a number of scheduled tasks received from the scheduler unit 2412, and the work distribution unit 2414 maintains a pending task pool and an active task pool for each GPC 2418.In at least one embodiment, the pending task pool includes a number of slots (e.g., 32 slots) containing tasks assigned for processing by a particular GPC 2418; the active task pool may include a number of slots (e.g., 4 slots) for tasks actively being processed by the GPCs 2418, such that when one of the GPCs 2418 completes execution of a task, that task is removed from the active task pool for the GPC 2418, and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 2418.In at least one embodiment, when an active task on the GPC 2418 is idle, for example, while waiting for a data dependency to be resolved, the active task is removed from the GPC 2418 and returned to a pending task pool while another task in the pending task pool is selected and scheduled for execution on the GPC 2418.

[0176] In at least one embodiment, work distribution unit 2414 communicates with one or more GPCs 2418 via XBar 2420. In at least one embodiment, XBar 2420 is an interconnection network that connects many units of PPU 2400 to other units of PPU 2400 and can be configured to connect work distribution unit 2414 to a particular GPC 2418. In at least one embodiment, one or more other units of PPU 2400 can also be connected to XBar 2420 via hub 2416.

[0177] In at least one embodiment, tasks are managed by the scheduler unit 2412 and forwarded by the work distribution unit 2414 to one of the GPCs 2418. The GPC 2418 is configured to process the task and produce results. In at least one embodiment, the results may be consumed by other tasks within the GPC 2418, forwarded to another GPC 2418 via the XBar 2420, or stored in memory 2404. In at least one embodiment, results may be written to memory 2404 via partition units 2422, which implement a memory interface for reading and writing data to / from memory 2404. In at least one embodiment, the results may be transferred to another PPU 2404 or CPU via a high-speed GPU interconnect 2408.In at least one embodiment, the PPU 2400 includes, without limitation, a number U of partition units 2422 that corresponds to the number of separate and distinct storage devices 2404 connected to the PPU 2400.

[0178] In at least one embodiment, a host processor executes a driver kernel that implements an application programming interface ("API") that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 2400. In at least one embodiment, multiple applications are executed concurrently by the PPU 2400, and the PPU 2400 provides isolation, quality of service ("QoS"), and independent address spaces for multiple applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause a driver kernel to generate one or more tasks for execution by the PPU 2400, and the driver kernel issues tasks to one or more streams that are processed by the PPU 2400. In at least one embodiment, each task comprises one or more groups of related threads, which may be referred to as a warp.In at least one embodiment, a warp comprises a plurality of contiguous threads (e.g., 32 threads) that can execute in parallel. In at least one embodiment, cooperating threads may refer to a plurality of threads that comprise instructions for executing a task and that exchange data via a shared memory.

[0179] Fig. 25 shows a GPC 2500 in accordance with at least one embodiment. In at least one embodiment, the GPC 2500 is the GPC 2418 of Fig. 24. In at least one embodiment, each GPC 2500 includes, without limitation, a number of hardware units for processing tasks, and each GPC 2500 includes, without limitation, a pipeline manager 2502, a work distribution unit (“PROP”) 2504, a raster engine 2508, a work distribution crossbar (“WDX”) 2516, an MMU 2518, one or more data processing clusters (“DPCs”) 2506, and any suitable combination of parts.

[0180] In at least one embodiment, the operation of the GPC 2500 is controlled by the pipeline manager 2502. In at least one embodiment, the pipeline manager 2502 manages the configuration of one or more DPCs 2506 for processing the tasks assigned to the GPC 2500. In at least one embodiment, the pipeline manager 2502 configures at least one of the one or more DPCs 2506 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, the DPC 2506 is configured to execute a vertex shader program on a programmable streaming multiprocessor ("SM") 2514.In at least one embodiment, pipeline manager 2502 is configured to forward packets received from a work distribution unit to appropriate logical units within GPC 2500, and in at least one embodiment, some packets may be forwarded to fixed-function hardware units in PROP 2504 and / or raster engine 2508, while other packets may be forwarded to DPCs 2506 for processing by a primitive engine 2512 or SM 2514. In at least one embodiment, pipeline manager 2502 configures at least one of DPCs 2506 to implement a computational pipeline. In at least one embodiment, pipeline manager 2502 configures at least one of DPCs 2506 to execute at least a portion of a CUDA program.

[0181] In at least one embodiment, the PROP unit 2504 is configured to forward the data generated by the raster engine 2508 and the DPCs 2506 to a raster operations partition unit (“ROP”), such as the memory partition unit 2422 described above in connection with Fig. 24. In at least one embodiment, the PROP unit 2504 is configured to perform optimizations for color blending, pixel data organization, address translations, etc. In at least one embodiment, the raster engine 2508 includes, without limitation, a number of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 2508 includes, without limitation, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tiling engine, and any suitable combination thereof.In at least one embodiment, a setup engine receives transformed vertices and generates plane equations associated with a vertex-defined geometric primitive; the plane equations are passed to a coarse raster engine to generate coverage information (e.g., an x, y coverage mask for a tile) for a primitive; the output of the coarse raster engine is passed to a culling engine, where fragments associated with a primitive that fail a z-test are discarded, and to a clipping engine, where fragments that lie outside a view frustum are clipped. In at least one embodiment, fragments that survive clipping and culling are passed to a raster engine to generate attributes for pixel fragments based on plane equations generated by a setup engine.In at least one embodiment, the output of raster engine 2508 comprises fragments that are processed by any suitable unit, such as a fragment shader implemented in DPC 2506.

[0182] In at least one embodiment, each DPC 2506 included in GPC 2500 includes, without limitation, an M-Pipe Controller ("MPC") 2510; a primitive engine 2512; one or more SMs 2514; and any suitable combination thereof. In at least one embodiment, MPC 2510 controls the operation of DPC 2506 and forwards packets received from pipeline manager 2502 to the appropriate units in DPC 2506. In at least one embodiment, packets associated with a vertex are forwarded to primitive engine 2512, which is configured to retrieve vertex attributes associated with the vertex from memory; in contrast, packets associated with a shader program may be transferred to SM 2514.

[0183] In at least one embodiment, SM 2514 includes, without limitation, a programmable streaming processor configured to process tasks represented by a number of threads. In at least one embodiment, SM 2514 is multi-threaded and configured to concurrently execute a plurality of threads (e.g., 32 threads) from a given group of threads and implements a SIMD architecture in which each thread in a group of threads (e.g., a warp) is configured to process a different set of instructions based on the same set of instructions. In at least one embodiment, all threads in a group of threads execute the same instructions.In at least one embodiment, SM 2514 implements a SIMT architecture, where each thread in a group of threads is configured to process a different set of instructions based on the same instruction set, but where the individual threads in the group of threads are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, enabling concurrency between warps and serial execution within warps when threads diverge within a warp. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, enabling equal concurrency among all threads within and between warps.In at least one embodiment, execution state is maintained for each individual thread, and threads executing the same instructions may be merged and executed in parallel to improve efficiency. At least one embodiment of SM 2514 is described in connection with. Fig. 26 described in more detail.

[0184] In at least one embodiment, MMU 2518 provides an interface between GPC 2500 and a memory partition unit (e.g., partition unit 2422 of Fig. 24), and MMU 2518 provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, MMU 2518 provides one or more translation lookaside buffers (TLBs) for performing virtual address to physical address translation in memory.

[0185] Fig. Figure 26 shows a streaming multiprocessor ("SM") 2600 in accordance with at least one embodiment. In at least one embodiment, SM 2600 is SM 2514 of Fig. 25. In at least one embodiment, SM 2600 includes, without limitation, an instruction cache 2602; one or more scheduler units 2604; a register file 2608; one or more processing cores ("cores") 2610; one or more special function units ("SFUs") 2612; one or more LSUs 2614; an interconnect network 2616; a shared memory / L1 cache 2618; and any suitable combination thereof. In at least one embodiment, a work distribution unit distributes tasks for execution among GPCs of parallel processing units (PPUs), and each task is assigned to a particular data processing cluster (DPC) within a GPC, and if a task is associated with a shader program, the task is assigned to one of the SMs 2600.In at least one embodiment, the scheduler unit 2604 receives tasks from a work distribution unit and manages the scheduling of instructions for one or more thread blocks assigned to the SM 2600. In at least one embodiment, the scheduler unit 2604 schedules thread blocks for execution as warps of parallel threads, with each thread block assigned at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 2604 manages a plurality of different thread blocks by assigning warps to the different thread blocks and then dispatching instructions from a plurality of different cooperative groups to different functional units (e.g., processing cores 2610, SFUs 2612, and LSUs 2614) during each clock cycle.

[0186] In at least one embodiment, "cooperative groups" may refer to a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads communicate, thus enabling richer, more efficient parallel decompositions. In at least one embodiment, cooperative startup APIs support synchronization between thread blocks for executing parallel algorithms. In at least one embodiment, APIs of conventional programming models provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function).However, in at least one embodiment, programmers can define groups of threads with a smaller granularity than thread blocks and synchronize within the defined groups to enable higher performance, design flexibility, and software reuse in the form of common group-wide functional interfaces. In at least one embodiment, cooperative groups allow programmers to explicitly define groups of threads with sub-block and multi-block granularity and perform collective operations such as synchronization on threads in a cooperative group. In at least one embodiment, the granularity of a sub-block is as small as a single thread.In at least one embodiment, a programming model supports clean composition across software boundaries, allowing libraries and utilities to safely synchronize within their local context without requiring convergence assumptions. In at least one embodiment, cooperative group primitives enable new patterns of cooperative parallelism, including, without limitation, producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire grid of thread blocks.

[0187] In at least one embodiment, a dispatch unit 2606 is configured to transmit instructions to one or more functional units, and the scheduler unit 2604 includes, without limitation, two dispatch units 2606 that allow two different instructions from the same warp to be dispatched during each clock cycle. In at least one embodiment, each scheduler unit 2604 includes a single dispatch unit 2606 or additional dispatch units 2606.

[0188] In at least one embodiment, each SM 2600 includes, without limitation, a register file 2608 that provides a set of registers for functional units of the SM 2600. In at least one embodiment, the register file 2608 is partitioned among the individual functional units such that each functional unit is assigned a specific portion of the register file 2608. In at least one embodiment, the register file 2608 is partitioned between different warps executed by the SM 2600, and the register file 2608 provides temporary storage for operands associated with data paths of functional units. In at least one embodiment, each SM 2600 includes, without limitation, a plurality of L processing cores 2610. In at least one embodiment, the SM 2600 includes, without limitation, a large number (e.g., 128 or more) of different processing cores 2610.In at least one embodiment, each processing core 2610 includes, without limitation, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, including, without limitation, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic units implement the IEEE 754-2008 standard for floating-point computations. In at least one embodiment, the processing cores 2610 include, without limitation, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0189] In at least one embodiment, the tensor cores are configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in the processing cores 2610. In at least one embodiment, tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for training and inferring neural networks. In at least one embodiment, each tensor core operates on a 4x4 matrix and performs a matrix multiplication and accumulation operation D = A × B + C, where A, B, C, and D are 4x4 matrices.

[0190] In at least one embodiment, the inputs for matrix multiplication A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor cores operate on 16-bit floating-point input data with 32-bit floating-point accumulation. In at least one embodiment, 64 operations are used for the 16-bit floating-point multiplication, resulting in a full-precision product, which is then accumulated by 32-bit floating-point addition with other intermediate products to form a 4x4x4 matrix multiplication. In at least one embodiment, tensor cores are used to perform much larger two-dimensional or higher-dimensional matrix operations constructed from these smaller elements.In at least one embodiment, an API, such as a CUDA C++ API, provides specialized operations for loading, multiplying, accumulating, and storing matrices to efficiently utilize tensor cores in a CUDA C++ program. In at least one embodiment at the CUDA level, a warp-level interface assumes 16x16 matrices spanning all 32 threads of a warp.

[0191] In at least one embodiment, each SM 2600 includes, without limitation, M SFUs 2612 that perform specific functions (e.g., attribute evaluation, reciprocal square root, and the like). In at least one embodiment, the SFUs 2612 include, without limitation, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFUs 2612 include, without limitation, a texture unit configured to perform texture map filtering operations. In at least one embodiment, texture units are configured to load texture maps (e.g., a 2D array of texels) from memory and sample texture maps to generate sampled texture values ​​for use in shader programs executed by the SM 2600. In at least one embodiment, the texture maps are stored in shared memory / L1 cache 2618.In at least one embodiment, texture units implement texture operations such as filtering operations using mipmaps (e.g., texture maps with different levels of detail). In at least one embodiment, each SM 2600 includes, without limitation, two texture units.

[0192] In at least one embodiment, each SM 2600 includes, without limitation, N LSUs 2614 that perform load and store operations between the shared memory / L1 cache 2618 and the register file 2608. In at least one embodiment, each SM 2600 includes, without limitation, an interconnection network 2616 that connects each of the functional units to the register file 2608 and the LSU 2614 to the register file 2608 and the shared memory / L1 cache 2618. In at least one embodiment, the interconnection network 2616 is a crossbar that can be configured to connect each of the functional units to each of the registers in the register file 2608 and to connect the LSUs 2614 to the register file 2608 and the memory locations in the shared memory / L1 cache 2618.

[0193] In at least one embodiment, shared memory / L1 cache 2618 is an array of on-chip memory that enables data storage and communication between SM 2600 and a primitive engine, and between threads in SM 2600. In at least one embodiment, shared memory / L1 cache 2618 includes, without limitation, 128 KB of memory capacity and is located in a path from SM 2600 to a partition unit. In at least one embodiment, shared memory / L1 cache 2618 is used to cache reads and writes. In at least one embodiment, one or more of shared memory / L1 cache 2618, L2 cache, and memory are backing stores.

[0194] In at least one embodiment, combining data caching and shared memory functionality in a single memory block provides improved performance for both types of memory accesses. In at least one embodiment, the capacity is or can be used as cache by programs that do not use shared memory; for example, if shared memory is configured to use half the capacity, texture and load / store operations can use the remaining capacity. In at least one embodiment, integration with shared memory / L1 cache 2618 enables shared memory / L1 cache 2618 to act as a high-throughput conduit for streaming data while providing high-bandwidth, low-latency access to frequently reused data.In at least one embodiment, the configuration for general-purpose parallel computing may use a simpler configuration than that used for graphics processing. In at least one embodiment, fixed-function GPUs are bypassed, resulting in a significantly simpler programming model. In at least one embodiment, and in a configuration for general-purpose parallel computing, a work distribution unit allocates and distributes blocks of threads directly to the DPCs.In at least one embodiment, threads within a block execute the same program, using a unique thread ID in a computation to ensure that each thread produces unique results, using SM 2600 to execute a program and perform computations, shared memory / L1 cache 2618 for communication between threads, and LSU 2614 to read and write global memory via shared memory / L1 cache 2618 and a memory partition unit. In at least one embodiment, when configured for general parallel computations, SM 2600 writes instructions that scheduler unit 2604 can use to start new work on DPCs.

[0195] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, servers, supercomputers, a smartphone (e.g., a wireless wearable device), a PDA, a digital camera, a vehicle, a head-mounted display, a portable electronic device, and more. In at least one embodiment, the PPU is packaged on a single semiconductor substrate. In at least one embodiment, the PPU is included in an SoC along with one or more other devices, such as additional PPUs, memory, a RISC CPU, an MMU, a digital-to-analog converter ("DAC"), and the like.

[0196] In at least one embodiment, the PPU may be included on a graphics card that includes one or more memory devices. In at least one embodiment, a graphics card may be configured to interface with a PCIe slot on a motherboard of a desktop computer. In at least one embodiment, the PPU may be an integrated GPU ("iGPU") that includes the motherboard chipset. Software constructions for general-purpose computing

[0197] The following figures show, without limitation, exemplary software constructions for implementing at least one embodiment.

[0198] Fig. 27 shows a software stack of a programming platform in accordance with at least one embodiment. In at least one embodiment, a programming platform is a platform for utilizing hardware on a computer system to accelerate computational tasks. In at least one embodiment, a programming platform may be accessible to software developers via libraries, compiler directives, and / or programming language extensions. In at least one embodiment, a programming platform may be, but is not limited to, CUDA, Radeon Open Compute Platform ("ROCm"), OpenCL (OpenCL™ is developed by the Khronos group), SYCL, or Intel One API.

[0199] In at least one embodiment, a software stack 2700 of a programming platform provides an execution environment for an application 2701. In at least one embodiment, the application 2701 may include any computer software that can be launched on the software stack 2700. In at least one embodiment, the application 2701 may include, but is not limited to, an artificial intelligence ("AI") / machine learning ("ML") application, a high-performance computing ("HPC") application, a virtual desktop infrastructure ("VDI"), or a data center workload.

[0200] In at least one embodiment, the application 2701 and the software stack 2700 run on the hardware 2707. The hardware 2707 may, in at least one embodiment, include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of devices that support a programming platform. In at least one embodiment, such as with CUDA, the software stack 2700 may be vendor-specific and compatible only with devices from certain manufacturers. In at least one embodiment, such as with OpenCL, the software stack 2700 may be used with devices from different vendors. In at least one embodiment, the hardware 2707 includes a host connected to one or more devices that can be accessed to perform computational tasks via application programming interface (API) calls.A device within hardware 2707 may, in at least one embodiment, include a GPU, FPGA, AI engine, or other computing device (but may also include a CPU) and its memory, as opposed to a host within hardware 2707, which may, in at least one embodiment, include, but is not limited to, a CPU (but may also include a computing device) and its memory.

[0201] In at least one embodiment, the software stack 2700 of a programming platform includes, without limitation, a number of libraries 2703, a runtime 2705, and a device kernel driver 2706. In at least one embodiment, each of the libraries 2703 may include data and programming code that can be used by computer programs and utilized during software development. In at least one embodiment, the libraries 2703 may include, but are not limited to, pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. In at least one embodiment, the libraries 2703 include functions optimized for execution on one or more types of devices.In at least one embodiment, libraries 2703 may include, but are not limited to, functions for performing mathematical, deep learning, and / or other types of operations on devices. In at least one embodiment, libraries 2703 are coupled to corresponding APIs 2702, which may include one or more APIs exposing the functions implemented in libraries 2703. In at least one embodiment, a processor (e.g., CPU, GPU) executes, invokes, or otherwise uses one or more APIs to prioritize kernels. For example, a first kernel (e.g., parent kernel) may launch a second kernel (e.g., child kernel), and the second kernel may be used by a processor to launch additional kernels (e.g., grandchild kernels) independently of the first kernel.In at least one embodiment, a processor executes an API to support dynamic stream priority (e.g., updating the priority while a stream is used to perform operations). When a processor executes this API, a programmer can, for example, copy the stream priority from one stream to one or more other streams.

[0202] In at least one embodiment, software stack 2700 includes an API for supporting dynamic stream priority (e.g., updating the priority while a stream is being used to perform operations), allowing a programmer to set the priority of a stream at any time after creation. In at least one embodiment, software stack 2700 includes an API for supporting dynamic stream priority (e.g., updating the priority while the stream is being used to perform operations), allowing a programmer to obtain the current priority of a stream, where the priority is one of a plurality of attributes of a stream.In at least one embodiment, software stack 2700 includes an API for supporting dynamic stream priority (e.g., updating the priority while the stream is being used to perform operations), allowing a programmer to obtain the current priority of a stream as a single attribute. In at least one embodiment, software stack 2700 includes an API for supporting dynamic stream priority (e.g., updating the priority while the stream is being used to perform operations), allowing a programmer to launch a kernel to perform operations on a stream with a specified priority that may be different from the stream priority.

[0203] In at least one embodiment, the application 2701 is written as source code that is compiled into executable code, as described below in connection with Fig. 32-34 are discussed in more detail. The executable code of application 2701 may, in at least one embodiment, be executed at least partially in an environment provided by software stack 2700. In at least one embodiment, during execution of application 2701, code may be accessed that must be executed on a device rather than a host. In such a case, in at least one embodiment, runtime 2705 may be invoked to load and launch the required code onto the device. In at least one embodiment, runtime 2705 may comprise any technically feasible runtime system capable of supporting execution of application 2701.

[0204] In at least one embodiment, runtime 2705 is implemented as one or more runtime libraries coupled to corresponding APIs represented as API(s) 2704. One or more such runtime libraries may, in at least one embodiment, include, among other functions, memory management, execution control, device control, error handling, and / or synchronization. In at least one embodiment, the memory management functions may include, but are not limited to, functions for allocating, freeing, and copying device memory, as well as for transferring data between host memory and device memory.In at least one embodiment, execution control functions may include, but are not limited to, functions for starting a function (sometimes referred to as a "kernel" when a function is a global function that can be called from a host) on a device and for setting attribute values ​​in a buffer maintained by a runtime library for a given function to be executed on a device.

[0205] Runtime libraries and corresponding API(s) 2704 may, in at least one embodiment, be implemented in any technically feasible manner. In at least one embodiment, one (or any number of) APIs may provide a low-level set of functions for fine-grained control of a device, while another (or any number of) APIs may provide a higher-level set of such functions. In at least one embodiment, a high-level API for the runtime may be built upon a low-level API. In at least one embodiment, one or more of the runtime APIs may be language-specific APIs layered upon a language-independent runtime API.

[0206] In at least one embodiment, device kernel driver 2706 is configured to facilitate communication with an underlying device. In at least one embodiment, device kernel driver 2706 may provide low-level functionality relied upon by APIs, such as API(s) 2704, and / or other software. In at least one embodiment, device kernel driver 2706 may be configured to compile Intermediate Representation ("IR") code into binary code at runtime. With CUDA, device kernel driver 2706 may compile IR code that is not hardware-specific into binary code for a particular device at runtime (with caching of the compiled binary code), which in at least one embodiment is also referred to as "finalization code."In at least one embodiment, code finalized in this manner may be executed on a device that did not exist when the source code was originally compiled into PTX code. Alternatively, in at least one embodiment, the device source code may be compiled offline into binary code without requiring the device kernel driver 2706 to compile the IR code at runtime.

[0207] Fig. 28 shows a CUDA implementation of the software stack 2700 from Fig. 27 in accordance with at least one embodiment. In at least one embodiment, a CUDA software stack 2800 on which an application 2801 may be launched includes CUDA libraries 2803, a CUDA runtime 2805, a CUDA driver 2807, and a device kernel driver 2808. In at least one embodiment, the CUDA software stack 2800 executes on hardware 2809, which may include a graphics processor supporting CUDA and developed by NVIDIA Corporation of Santa Clara, CA.

[0208] In at least one embodiment, the application 2801, the CUDA runtime 2805, and the device kernel driver 2808 may perform similar functions to the application 2701, the runtime 2705, and the device kernel driver 2706, respectively, described above in connection with Fig. 27. In at least one embodiment, the CUDA driver 2807 includes a library (libcuda.so) that implements a CUDA driver API 2806. Similar to a CUDA runtime API 2804 implemented by a CUDA runtime library (cudart), the CUDA driver API 2806, in at least one embodiment, may provide, among other things, functions for memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability. In at least one embodiment, the CUDA driver API 2806 differs from the CUDA runtime API 2804 in that the CUDA runtime API 2804 simplifies device code management by providing implicit initialization, context management (analogous to a process), and module management (analogous to dynamically loaded libraries).In contrast to the high-level CUDA runtime API 2804, the CUDA driver API 2806 is a low-level API that, in at least one embodiment, enables finer-grained control of the device, particularly with respect to contexts and module loading. In at least one embodiment, the CUDA driver API 2806 may provide context management functionality not provided by the CUDA runtime API 2804. In at least one embodiment, the CUDA driver API 2806 is also language-independent, supporting, for example, OpenCL in addition to the CUDA runtime API 2804. Furthermore, in at least one embodiment, the development libraries, including the CUDA runtime 2805, may be considered separate from the driver components, including the user device CUDA driver 2807 and the kernel device driver 2808 (sometimes referred to as a "display" driver).

[0209] In at least one embodiment, the CUDA libraries 2803 may include, but are not limited to, mathematical libraries, deep learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries that may be accessed by parallel computing applications such as application 2801. In at least one embodiment, the CUDA libraries 2803 may include mathematical libraries such as a cuBLAS library, which is an implementation of Basic Linear Algebra Subprograms ("BLAS") for performing linear algebra operations, a cuFFT library for computing Fast Fourier Transforms ("FFTs"), and a cuRAND library for generating random numbers, among others.In at least one embodiment, the CUDA libraries 2803 may include, among others, deep learning libraries such as a cuDNN library of deep neural network primitives and a TensorRT platform for high-performance deep learning inference.

[0210] Fig. 29 shows a ROCm implementation of the software stack 2700 from Fig. 27 according to at least one embodiment. In at least one embodiment, an ROCm software stack 2900 on which an application 2901 may be launched includes a script runtime 2903, a system runtime 2905, a thunk 2907, and an ROCm kernel driver 2908. In at least one embodiment, the ROCm software stack 2900 executes on hardware 2909, which may include a GPU supporting ROCm and developed by AMD Corporation of Santa Clara, CA.

[0211] In at least one embodiment, the application 2901 may provide similar functionality as described above in connection with Fig. 27. Furthermore, the speech runtime 2903 and the system runtime 2905 may, in at least one embodiment, perform similar functions as those described above in connection with Fig. 27. In at least one embodiment, the Srpach runtime 2903 and the system runtime 2905 differ in that the system runtime 2905 is a language-independent runtime that implements a ROCr system runtime API 2904 and uses a Heterogeneous System Architecture ("HSA") runtime API. The HSA runtime API is a lightweight user-mode API that provides interfaces for accessing and interacting with an AMD graphics processor, including, in at least one embodiment, functions for memory management, execution control via architectural kernel dispatch, error handling, system and agent information, and runtime initialization and shutdown, among other things.In contrast to the system runtime 2905, the language runtime 2903, in at least one embodiment, is an implementation of a language-specific runtime API 2902 that lies on top of the ROCr system runtime API 2904. In at least one embodiment, the language runtime API may include, but is not limited to, a Heterogeneous Compute Interface for Portability ("HIP") language runtime API, a Heterogeneous Compute Compiler ("HCC") language runtime API, or an OpenCL API, among others. In particular, the HIP language is an extension of the C++ programming language with functionally similar versions of the CUDA mechanisms, and in at least one embodiment, a HIP language runtime API includes functionality similar to that of the language runtime APIs described above in connection with. Fig. 28 are similar to the CUDA runtime API 2804 discussed in section 28, such as functions for memory management, execution control, device management, error handling, and synchronization, among others.

[0212] In at least one embodiment, thunk (ROCt) 2907 is an interface 2906 that can be used to interact with the underlying ROCm driver 2908. In at least one embodiment, the ROCm driver 2908 is a ROCk driver, which is a combination of an AMD GPU driver and an HSA kernel driver (amdkfd). In at least one embodiment, the AMD GPU driver is a device kernel driver for GPUs developed by AMD that performs similar functions to the device kernel driver 2706 described above in connection with Fig. 27. In at least one embodiment, the HSA kernel driver is a driver that enables different types of processors to more effectively share system resources through hardware features.

[0213] In at least one embodiment, various libraries (not shown) may be included in the ROCm software stack 2900 above the Srpach runtime 2903 and provide similar functionality to the CUDA libraries 2803 described above in connection with Fig. 28. In at least one embodiment, various libraries may include mathematical, deep learning, and / or other libraries, such as a hipBLAS library implementing functions similar to those of CUDA cuBLAS, a rocFFT library for computing FFTs, similar to, but not limited to, CUDA cuFFT, among others.

[0214] Fig. 30 shows an OpenCL implementation of the software stack 2700 from Fig. 27 according to at least one embodiment. In at least one embodiment, an OpenCL software stack 3000 with which an application 3001 can be launched includes an OpenCL framework 3010, an OpenCL runtime 3006, and a driver 3007. In at least one embodiment, the OpenCL software stack 3000 executes on hardware 2809 that is not vendor-specific. Because OpenCL is supported by devices developed by different vendors, specific OpenCL drivers may be required to interoperate with hardware from such vendors in at least one embodiment.

[0215] In at least one embodiment, the application 3001, the OpenCL runtime 3006, the device kernel driver 3007, and the hardware 3008 may perform similar functions to the application 2701, the runtime 2705, the device kernel driver 2706, and the hardware 2707, respectively, described above in connection with Fig. 27. In at least one embodiment, the application 3001 further comprises an OpenCL kernel 3002 having code to be executed on a device.

[0216] In at least one embodiment, OpenCL defines a "platform" that enables a host to control devices connected to the host. In at least one embodiment, an OpenCL framework provides a platform layer API and a runtime API, represented as Platform API 3003 and Runtime API 3005. In at least one embodiment, Runtime API 3005 uses contexts to manage the execution of kernels on devices. In at least one embodiment, each identified device can be associated with a corresponding context, which Runtime API 3005 can use to, among other things, manage instruction queues, program objects, and kernel objects, as well as to share memory objects for that device. In at least one embodiment, Platform API 3003 provides functions to configure device contexts, among other things.to select and initialize devices, to submit work to devices via command queues, and to enable data transfer to and from devices. Furthermore, in at least one embodiment, the OpenCL framework provides various built-in functions (not shown), including mathematical functions, relational functions, and image processing functions, to name a few.

[0217] In at least one embodiment, a compiler 3004 is also included in the OpenCL framework 3010. The source code may, in at least one embodiment, be compiled offline prior to execution of an application or online during execution of an application. Unlike CUDA and ROCm, in at least one embodiment, OpenCL applications may be compiled online by compiler 3004, which is representative of any number of compilers that may be used to compile source code and / or IR code, such as Standard Portable Intermediate Representation ("SPIR-V") code, into binary code. Alternatively, in at least one embodiment, OpenCL applications may be compiled offline before such applications are executed.

[0218] Fig. 31 shows software supported by a programming platform according to at least one embodiment. In at least one embodiment, a programming platform 3104 is configured to support various programming models 3103, middleware and / or libraries 3102, and frameworks 3101 that an application 3100 may rely on. In at least one embodiment, the application 3100 may be an AI / ML application implemented, for example, using a deep learning framework such as MXNet, PyTorch, or TensorFlow, which may rely on libraries such as cuDNN, NVIDIA Collective Communications Library ("NCCL"), and / or NVIDIA Developer Data Loading Library ("DALI") CUDA libraries to provide accelerated computations on the underlying hardware.

[0219] In at least one embodiment, the programming platform 3104 may implement one of the methods described above in connection with Fig. 28, Fig. 29 respectively Fig. 30. In at least one embodiment, the programming platform 3104 supports multiple programming models 3103, which are abstractions of an underlying computer system that enable expressions of algorithms and data structures. In at least one embodiment, the programming models 3103 may expose features of the underlying hardware to improve performance. In at least one embodiment, the programming models 3103 may include, but are not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism ("C++AMP"), Open Multi-Processing ("OpenMP"), Open Accelerators ("OpenACC"), and / or Vulcan Compute.

[0220] In at least one embodiment, libraries and / or middleware 3102 provide implementations of abstractions of programming models 3104. In at least one embodiment, such libraries include data and programming code that can be used by computer programs and utilized during software development. At least in one embodiment, such middleware includes software that provides services to applications beyond those available from programming platform 3104. In at least one embodiment, libraries and / or middleware 3102 may include, but are not limited to, cuBLAS, cuFFT, cuRAND, and other CUDA libraries, or rocBLAS, rocFFT, rocRAND, and other ROCm libraries.Furthermore, in at least one embodiment, the libraries and / or middleware 3102 may include NCCL and ROCm Communication Collectives Library (“RCCL”) libraries that provide communication routines for GPUs, an MIOpen library for deep learning accelerators, and / or an Eigen library for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.

[0221] In at least one embodiment, the application frameworks 3101 depend on libraries and / or middleware 3102. In at least one embodiment, each of the application frameworks 3101 is a software framework used to implement a standard application software structure. Returning to the AI / ML example discussed above, in at least one embodiment, an AI / ML application may be implemented using a framework such as Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks.

[0222] Fig. 32 shows the compilation of code for execution on one of the programming platforms of the Fig. 27-30 according to at least one embodiment. In at least one embodiment, a compiler 3201 receives source code 3200 that includes both host code and device code. In at least one embodiment, the compiler 3201 is configured to convert the source code 3200 into host executable code 3202 for execution on a host and into device executable code 3203 for execution on a device. In at least one embodiment, the source code 3200 can be compiled either offline prior to execution of an application or online during execution of an application.

[0223] In at least one embodiment, source code 3200 may include code in any programming language supported by compiler 3201, such as C++, C, Fortran, etc. In at least one embodiment, source code 3200 may include a single-source file having a mixture of host code and device code, with the locations of the device code indicated therein. In at least one embodiment, a single-source file may be a .cu file including CUDA code or a .hip.cpp file including HIP code. Alternatively, in at least one embodiment, source code 3200 may include multiple source code files in which host code and device code are separated, instead of a single-source file.

[0224] In at least one embodiment, compiler 3201 is configured to compile source code 3200 into host executable code 3202 for execution on a host and device executable code 3203 for execution on a device. In at least one embodiment, compiler 3201 performs operations including parsing source code 3200 into an abstract system tree (AST), performing optimizations, and generating executable code. In at least one embodiment where source code 3200 comprises a single-source file, compiler 3201 may separate device code from host code in such a single-source file, combine the device code and host code into device executable code 3203 and device executable code 3203, respectively.a host executable code 3202 and link the device executable code 3203 and the host executable code 3202 into a single file, as described below with respect to . Fig. 33 is explained in more detail.

[0225] In at least one embodiment, host-executable code 3202 and device-executable code 3203 may be in any suitable format, such as binary code and / or IR code. In the case of CUDA, in at least one embodiment, host-executable code 3202 may comprise native object code, and device-executable code 3203 may comprise PTX intermediate representation code. In the case of ROCm, both host-executable code 3202 and device-executable code 3203 may comprise target binary code, in at least one embodiment.

[0226] Fig. 33 is a more detailed illustration of compiling code for execution on one of the programming platforms of the Fig. 27-30 in accordance with at least one embodiment. In at least one embodiment, a compiler 3301 is configured to receive source code 3300, compile source code 3300, and output an executable file 3310. In at least one embodiment, source code 3300 is a single-source file, such as a .cu file, a .hip.cpp file, or a file in another format that includes both host and device code. In at least one embodiment, compiler 3301 may be, but is not limited to, an NVIDIA CUDA compiler ("NVCC") for compiling CUDA code into .cu files or an HCC compiler for compiling HIP code into .hip.cpp files.

[0227] In at least one embodiment, compiler 3301 includes a compiler front-end 3302, a host compiler 3305, a device compiler 3306, and a linker 3309. In at least one embodiment, compiler front-end 3302 is configured to separate device code 3304 from host code 3303 in source code 3300. Device code 3304 is compiled by device compiler 3306 into device-executable code 3308, which, as described, may comprise binary code or IR code in at least one embodiment. Separately, host code 3303 is compiled by host compiler 3305 into executable host code 3307 in at least one embodiment.For NVCC, host compiler 3305 may be a general-purpose C / C++ compiler that outputs native object code, while device compiler 3306 may be a Low Level Virtual Machine ("LLVM")-based compiler that forks an LLVM compiler infrastructure and outputs PTX code or binary code, at least in one embodiment. For HCC, both host compiler 3305 and device compiler 3306 may be, but are not limited to, LLVM-based compilers that output target binary code, in at least one embodiment.

[0228] Following compilation of source code 3300 into host executable code 3307 and device executable code 3308, compiler 3309, in at least one embodiment, links the host executable code and device executable code 3307 and 3308 together in executable file 3310. In at least one embodiment, native object code for a host and PTX or binary code for a device may be linked together in an Executable and Linkable Format (ELF) file, which is a container format for storing object code.

[0229] Fig. 34 shows the translation of the source code prior to compilation of the source code in accordance with at least one embodiment. In at least one embodiment, the source code 3400 is passed through a translation tool 3401, which translates the source code 3400 into translated source code 3402. In at least one embodiment, a compiler 3403 is used to compile the translated source code 3402 into host-executable code 3404 and device-executable code 3405, in a process similar to the compilation of the source code 3200 by the compiler 3201 into host-executable code 3202 and device-executable code 3203, as described above in connection with Fig. 32 described.

[0230] In at least one embodiment, a translation performed by the translation tool 3401 is used to port the source code 3400 for execution in a different environment than the one in which it was originally intended to be executed. In at least one embodiment, the translation tool 3401 may include, but is not limited to, a HIP translator used to "hipify" CUDA code intended for a CUDA platform into HIP code that can be compiled and executed on a ROCm platform. In at least one embodiment, the translation of the source code 3400 may include parsing the source code 3400 and converting calls to API(s) provided by one programming model (e.g., CUDA) into corresponding calls to API(s) provided by another programming model (e.g., HIP), as described below in connection with Fig. 35A-36. Returning to the example of HIPification of CUDA code, in at least one embodiment, calls to the CUDA runtime API, the CUDA driver API, and / or the CUDA libraries may be converted into corresponding HIP API calls. In at least one embodiment, the automatic translations performed by the translation tool 3401 may sometimes be incomplete, requiring additional manual effort to fully port the source code 3400. CONFIGURING GPUS FOR GENERAL-PURPOSE COMPUTING

[0231] The following figures illustrate, without limitation, exemplary architectures for compiling and executing computational source code in accordance with at least one embodiment.

[0232] Fig. 35A shows a system 3500 configured to compile and execute CUDA source code 3510 using different types of processing units, according to at least one embodiment. In at least one embodiment, the system 3500 includes, without limitation, CUDA source code 3510, a CUDA compiler 3550, executable code on the host 3570(1), executable code on the host 3570(2), executable code for CUDA devices 3584, a CPU 3590, a CUDA-capable GPU 3594, a GPU 3592, a CUDA-to-HIP translation tool 3520, HIP source code 3530, a HIP compiler driver 3540, an HCC 3560, and executable code for HCC devices 3582.

[0233] In at least one embodiment, the CUDA source code 3510 is a collection of human-readable code in a CUDA programming language. In at least one embodiment, the CUDA code is human-readable code in a CUDA programming language. At least in one embodiment, a CUDA programming language is an extension of the C++ programming language that includes, without limitation, mechanisms for defining device code and distinguishing between device code and host code. In at least one embodiment, the device code is source code that, after compilation, is executable in parallel on a device. In at least one embodiment, a device may be a processor optimized for parallel processing of instructions, such as a CUDA-capable GPU 3590, GPU 35192, or other GPGPU, etc.In at least one embodiment, host code is source code that, after compilation, is executable on a host. In at least one embodiment, a host is a processor optimized for processing sequential instructions, such as the CPU 3590.

[0234] In at least one embodiment, the CUDA source code 3510 includes, without limitation, any number (including zero) of global functions 3512, any number (including zero) of device functions 3514, any number (including zero) of host functions 3516, and any number (including zero) of host / device functions 3518. In at least one embodiment, global functions 3512, device functions 3514, host functions 3516, and host / device functions 3518 may be intermixed in the CUDA source code 3510. In at least one embodiment, each of the global functions 3512 is executable on a device and invokable by a host. Therefore, in at least one embodiment, one or more of the global functions 3512 may serve as entry points for a device. In at least one embodiment, each of the global functions 3512 is a kernel.In at least one embodiment, and in a technique known as dynamic parallelism, one or more of the global functions 3512 define a kernel executable on a device and invokable from such a device. In at least one embodiment, a kernel is executed N (where N is any positive integer) times in parallel by N different threads on a device during execution.

[0235] In at least one embodiment, each of the device functions 3514 executes on a device and is callable only from such a device. In at least one embodiment, each of the host functions 3516 executes on a host and is callable only from such a host. In at least one embodiment, each of the host / device functions 3516 defines both a host version of a function executable on a host and callable only from such a host, and a device version of the function executable on a device and callable only from such a device.

[0236] In at least one embodiment, the CUDA source code 3510 may also include, without limitation, any number of calls to any number of functions defined via a CUDA runtime API 3502. In at least one embodiment, the CUDA runtime API 3502 may include, without limitation, any number of functions executed on a host to allocate and deallocate device memory, transfer data between host memory and device memory, manage multi-device systems, etc. In at least one embodiment, the CUDA source code 3510 may also include any number of calls to any number of functions specified in any number of other CUDA APIs. In at least one embodiment, a CUDA API may be any API intended for use by CUDA code.In at least one embodiment, the CUDA APIs include, without limitation, the CUDA Runtime API 3502, a CUDA Driver API, APIs for any number of CUDA libraries, etc. In at least one embodiment, a CUDA Driver API is a lower-level API compared to the CUDA Runtime API 3502, but enables finer-grained control of a device. In at least one embodiment, examples of CUDA libraries include, without limitation, cuBLAS, cuFFT, cuRAND, cuDNN, etc.

[0237] In at least one embodiment, the CUDA compiler 3550 compiles the input CUDA code (e.g., the CUDA source code 3510) to generate the host executable code 3570(1) and the device executable code 3584. In at least one embodiment, the CUDA compiler 3550 is NVCC. In at least one embodiment, the host executable code 3570(1) is a compiled version of the host code that includes the input source code executable on the CPU 3590. In at least one embodiment, the CPU 3590 may be any processor optimized for processing sequential instructions.

[0238] In at least one embodiment, the CUDA device executable code 3584 is a compiled version of the device code included in the input source code executable on the CUDA-capable GPU 3594. In at least one embodiment, the device executable code 3584 comprises, without limitation, binary code. In at least one embodiment, the CUDA device executable code 3584 comprises, without limitation, IR code, such as PTX code, that is further compiled at runtime by a driver into binary code for a particular target device (e.g., CUDA-capable GPU 3594). In at least one embodiment, the CUDA-capable GPU 3594 may be any processor optimized for processing parallel instructions and supporting CUDA. In at least one embodiment, the CUDA-capable graphics processor 3594 is developed by NVIDIA Corporation in Santa Clara, CA.

[0239] In at least one embodiment, the CUDA-to-HIP translation tool 3520 is configured to translate CUDA source code 3510 into functionally similar HIP source code 3530. In at least one embodiment, the HIP source code 3530 is a collection of human-readable code in a HIP programming language. In at least one embodiment, the HIP code is human-readable code in a HIP programming language. In at least one embodiment, a HIP programming language is an extension of the C++ programming language that includes, without limitation, functionally similar versions of CUDA mechanisms for defining device code and distinguishing between device code and host code. In at least one embodiment, a HIP programming language may include a subset of the functionality of a CUDA programming language.For example, in at least one embodiment, a HIP programming language includes, without limitation, mechanisms for defining global functions 3512, but such a HIP programming language may lack support for dynamic parallelism, and therefore global functions 3512 defined in HIP code may only be callable from host code.

[0240] In at least one embodiment, the HIP source code 3530 includes, without limitation, any number (including zero) of global functions 3512, any number (including zero) of device functions 3514, any number (including zero) of host functions 3516, and any number (including zero) of host / device functions 3518. In at least one embodiment, the HIP source code 3530 may also include any number of calls to any number of functions specified in a HIP runtime API 3532. In at least one embodiment, the HIP runtime API 3532 includes, without limitation, functionally similar versions of a subset of functions included in the CUDA runtime API 3502.In at least one embodiment, the HIP source code 3530 may also include any number of calls to any number of functions specified in any number of other HIP APIs. In at least one embodiment, a HIP API may be any API intended for use by HIP code and / or ROCm. In at least one embodiment, the HIP APIs include, without limitation, the HIP runtime API 3532, a HIP driver API, APIs for any number of HIP libraries, APIs for any number of ROCm libraries, etc.

[0241] In at least one embodiment, the CUDA-to-HIP translation tool 3520 converts each kernel call in the CUDA code from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the CUDA code to any number of other functionally similar HIP calls. In at least one embodiment, a CUDA call is a call to a function specified in a CUDA API, and a HIP call is a call to a function specified in a HIP API. In at least one embodiment, the CUDA-to-HIP translation tool 3520 converts any number of calls to functions specified in the CUDA runtime API 3502 to any number of calls to functions specified in the HIP runtime API 3532.

[0242] In at least one embodiment, the CUDA-to-HIP translation tool 3520 is a tool known as hipify-perl, which performs a text-based translation process. In at least one embodiment, the CUDA-to-HIP translation tool 3520 is a tool known as hipify-clang, which performs a more complex and robust translation process than hipify-perl, including parsing CUDA code using clang (a compiler front-end) and then translating the resulting symbols. In at least one embodiment, proper conversion of CUDA code to HIP code may be required in addition to the modifications (e.g., manual edits) performed by the CUDA-to-HIP translation tool 3520.

[0243] In at least one embodiment, the HIP compiler driver 3540 is a front-end that determines a target device 3546 and then configures a compiler compatible with the target device 3546 to compile the HIP source code 3530. In at least one embodiment, the target device 3546 is a processor optimized for processing parallel instructions. In at least one embodiment, the HIP compiler driver 3540 may determine the device 3546 in any technically feasible manner.

[0244] In at least one embodiment, the HIP compiler driver 3540 generates a HIP / NVCC compilation command 3542 if the target device 3546 is CUDA compatible (e.g., a CUDA-enabled GPU 3594). In at least one embodiment, and as described in connection with Fig. 35B, the HIP / NVCC compilation command 3542 configures the CUDA compiler 3550 to compile the HIP source code 3530 using, without limitation, a HIP-to-CUDA translation header and a CUDA runtime library. In at least one embodiment, and in response to the HIP / NVCC compilation command 3542, the CUDA compiler 3550 generates host-executable code 3570(1) and device-executable code 3584.

[0245] In at least one embodiment, the HIP compiler driver 3540 generates a HIP / HCC compilation command 3544 if the target device 3546 is not CUDA compatible. In at least one embodiment, and as described in connection with Fig. 35C, the HIP / HCC compile command 3544 configures HCC 3560 to compile HIP source code 3530 using, without limitation, an HCC header and a HIP / HCC runtime library. In at least one embodiment, and in response to the HIP / HCC compile command 3544, HCC 3560 generates host executable code 3570(2) and device executable code 3582. In at least one embodiment, the HCC device executable code 3582 is a compiled version of the device executable code that includes the HIP source code 3530 and is executable on the GPU 3592. In at least one embodiment, the GPU 3592 may be any processor optimized for processing parallel instructions, is non-CUDA compatible, and is HCC compatible. In at least one embodiment, the GPU 3592 is developed by AMD Corporation in Santa Clara, CA.In at least one embodiment, GPU 3592 is a non-CUDA-capable GPU 3592.

[0246] For clarification, Fig. 35A depicts three different flows that may be implemented, in at least one embodiment, to compile the CUDA source code 3510 for execution on the CPU 3590 and various devices. In at least one embodiment, a direct CUDA flow compiles the CUDA source code 3510 for execution on the CPU 3590 and the CUDA-enabled GPU 3594 without translating the CUDA source code 3510 into the HIP source code 3530. In at least one embodiment, an indirect CUDA flow translates the CUDA source code 3510 into the HIP source code 3530 and then compiles the HIP source code 3530 for execution on the CPU 3590 and the CUDA-enabled GPU 3594. In at least one embodiment, a CUDA / HCC flow translates the CUDA source code 3510 into HIP source code 3530 and then compiles the HIP source code 3530 for execution on the CPU 3590 and the GPU 3592.

[0247] A direct CUDA flow that may be implemented in at least one embodiment is depicted by dashed lines and a series of bubbles labeled A1-A3. In at least one embodiment, and as depicted with the bubble labeled A1, the CUDA compiler 3550 receives the CUDA source code 3510 and a CUDA compile command 3548 that configures the CUDA compiler 3550 to compile the CUDA source code 3510. In at least one embodiment, the CUDA source code 3510 used in a direct CUDA flow is written in a CUDA programming language based on a programming language other than C++ (e.g., C, Fortran, Python, Java, etc.).In at least one embodiment, and in response to the CUDA compilation command 3548, the CUDA compiler 3550 generates host-executable code 3570(1) and CUDA device-executable code 3584 (depicted by the bubble labeled A2). In at least one embodiment, and as depicted in the bubble labeled A3, the host-executable code 3570(1) and the device-executable code 3584 may be executed on the CPU 3590 and the CUDA-enabled GPU 3594, respectively. In at least one embodiment, the device-executable code 3584 comprises, without limitation, binary code. In at least one embodiment, the CUDA device-executable code 3584 comprises, without limitation, PTX code and is further compiled at runtime into binary code for a particular target device.

[0248] An indirect CUDA flow that may be implemented in at least one embodiment is depicted by dashed lines and a series of bubbles labeled B1-B6. In at least one embodiment, and as depicted in the bubble labeled B1, the CUDA-to-HIP translation tool 3520 receives the CUDA source code 3510. In at least one embodiment, and as depicted in the bubble labeled B2, the CUDA-to-HIP translation tool 3520 translates the CUDA source code 3510 into the HIP source code 3530. In at least one embodiment, and as depicted in bubble B3, the HIP compiler driver 3540 receives the HIP source code 3530 and determines that the target device 3546 is CUDA-capable.

[0249] In at least one embodiment, and as depicted with bubble B4, the HIP compiler driver 3540 generates the HIP / NVCC compilation instruction 3542 and transmits both the HIP / NVCC compilation instruction 3542 and the HIP source code 3530 to the CUDA compiler 3550. In at least one embodiment, and as depicted in connection with Fig. 35B, the HIP / NVCC compile command 3542 configures the CUDA compiler 3550 to compile the HIP source code 3530 using, without limitation, a HIP-to-CUDA translation header and a CUDA runtime library. In at least one embodiment, and in response to the HIP / NVCC compile command 3542, the CUDA compiler 3550 generates host-executable code 3570(1) and CUDA device-executable code 3584 (depicted with the bubble labeled B5). In at least one embodiment, and as depicted in the bubble labeled B6, the host executable code 3570(1) and the device executable code 3584 may be executed on the CPU 3590 and the CUDA-enabled GPU 3594, respectively. In at least one embodiment, the device executable code 3584 comprises, without limitation, binary code.In at least one embodiment, the CUDA code 3584 executable on the device comprises, without limitation, PTX code and is further compiled at runtime into binary code for a particular target device.

[0250] A CUDA / HCC flow that may be implemented in at least one embodiment is depicted by solid lines and a series of bubbles labeled C1-C6. In at least one embodiment, and as depicted in the bubble labeled C1, the CUDA-HIP translation tool 3520 receives the CUDA source code 3510. In at least one embodiment, and as depicted in the bubble labeled C2, the CUDA-HIP translation tool 3520 translates the CUDA source code 3510 into the HIP source code 3530. In at least one embodiment, and as depicted in the bubble labeled C3, the HIP compiler 3540 receives the HIP source code 3530 and determines that the target device 3546 is not CUDA-capable.

[0251] In at least one embodiment, the HIP compiler driver 3540 generates the HIP / HCC compile command 3544 and transmits both the HIP / HCC compile command 3544 and the HIP source code 3530 to the HCC 3560 (depicted in the balloon labeled C4). In at least one embodiment, and as described in connection with Fig. 35C, the HIP / HCC compile command 3544 configures the HCC 3560 to compile the HIP source code 3530 using, without limitation, an HCC header and a HIP / HCC runtime library. In at least one embodiment, and in response to the HIP / HCC compile command 3544, the HCC 3560 generates host executable code 3570(2) and device executable code 3582 (depicted with a bubble labeled C5). In at least one embodiment, and as depicted with the bubble labeled C6, the host executable code 3570(2) and device executable code 3582 may be executed on the CPU 3590 and the GPU 3592, respectively.

[0252] In at least one embodiment, after the CUDA source code 3510 has been translated into the HIP source code 3530, the HIP compiler driver 3540 can then be used to generate executable code for either the CUDA-capable GPU 3594 or the GPU 3592 without re-executing the CUDA-HIP translation tool 3520. In at least one embodiment, the CUDA-HIP translation tool 3520 translates the CUDA source code 3510 into the HIP source code 3530, which is then stored in memory. In at least one embodiment, the HIP compiler driver 3540 then configures the HCC 3560 to generate host executable code 3570(2) and device executable code 3582 based on the HCC. In at least one embodiment, the HIP compiler driver 3540 then configures the CUDA compiler 3550 to generate host executable code 3570(1) and CUDA device executable code 3584 based on the stored HIP source code 3530.

[0253] Fig. 35B shows a system 3504 configured to run the CUDA source code 3510 from Fig. 35A using a CPU 3590 and a CUDA-capable GPU 3594 in accordance with at least one embodiment. In at least one embodiment, the system 3504 includes, without limitation, the CUDA source code 3510, the CUDA-to-HIP translation tool 3520, the HIP source code 3530, the HIP compiler driver 3540, the CUDA compiler 3550, the host executable code 3570(1), the host executable code 3584, the CPU 3590, and the CUDA-capable GPU 3594.

[0254] In at least one embodiment, and as previously described herein in connection with Fig. 35A, the CUDA source code 3510 includes, without limitation, any number (including zero) of global functions 3512, any number (including zero) of device functions 3514, any number (including zero) of host functions 3516, and any number (including zero) of host / device functions 3518. In at least one embodiment, the CUDA source code 3510 also includes, without limitation, any number of calls to any number of functions specified in any number of CUDA APIs.

[0255] In at least one embodiment, the CUDA-to-HIP translation tool 3520 translates the CUDA source code 3510 into the HIP source code 3530. In at least one embodiment, the CUDA-to-HIP translation tool 3520 converts each kernel call in the CUDA source code 3510 from a CUDA syntax to a HIP syntax and converts any number of other CUDA calls in the CUDA source code 3510 into any number of other functionally similar HIP calls.

[0256] In at least one embodiment, the HIP compiler driver 3540 determines that the device 3546 is CUDA-capable and generates the HIP / NVCC compilation command 3542. In at least one embodiment, the HIP compiler driver 3540 then configures the CUDA compiler 3550 via the HIP / NVCC compilation command 3542 to compile HIP source code 3530. In at least one embodiment, the HIP compiler driver 3540 provides access to a HIP-to-CUDA translation header 3552 as part of the configuration of the CUDA compiler 3550. In at least one embodiment, the HIP-to-CUDA translation header 3552 translates any number of mechanisms (e.g., functions) specified in any number of HIP APIs into any number of mechanisms specified in any number of CUDA APIs.In at least one embodiment, the CUDA compiler 3550 uses the HIP-to-CUDA translation header 3552 in conjunction with a CUDA runtime library 3554 corresponding to the CUDA runtime API 3502 to generate the host executable code 3570(1) and the device executable code 3584. In at least one embodiment, the host executable code 3570(1) and the device executable code 3584 may then be executed on the CPU 3590 and the CUDA-enabled GPU 3594, respectively. In at least one embodiment, the device executable code 3584 comprises, without limitation, binary code. In at least one embodiment, the CUDA device executable code 3584 comprises, without limitation, PTX code and is further compiled at runtime into binary code for a particular target device.

[0257] Fig. 35C shows a system 3506 configured to run the CUDA source code 3510 from Fig. 35A using a CPU 3590 and a non-CUDA-capable GPU 3592 in accordance with at least one embodiment. In at least one embodiment, system 3506 includes, without limitation, CUDA source code 3510, CUDA-to-HIP translation tool 3520, HIP source code 3530, HIP compiler driver 3540, HCC 3560, host executable code 3570(2), HCC host executable code 3582, CPU 3590, and GPU 3592.

[0258] In at least one embodiment, and as previously described herein in connection with Fig. 35A, the CUDA source code 3510 includes, without limitation, any number (including zero) of global functions 3512, any number (including zero) of device functions 3514, any number (including zero) of host functions 3516, and any number (including zero) of host / device functions 3518. In at least one embodiment, the CUDA source code 3510 also includes, without limitation, any number of calls to any number of functions specified in any number of CUDA APIs.

[0259] In at least one embodiment, the CUDA-to-HIP translation tool 3520 translates the CUDA source code 3510 into the HIP source code 3530. In at least one embodiment, the CUDA-to-HIP translation tool 3520 converts each kernel call in the CUDA source code 3510 from a CUDA syntax to a HIP syntax and converts any number of other CUDA calls in the source code 3510 into any number of other functionally similar HIP calls.

[0260] In at least one embodiment, the HIP compiler driver 3540 then determines that the target device 3546 is not CUDA-capable and generates a HIP / HCC compilation command 3544. In at least one embodiment, the HIP compiler driver 3540 then configures the HCC 3560 to execute the HIP / HCC compilation command 3544 to compile the HIP source code 3530. In at least one embodiment, the HIP / HCC compiler 3544 configures the HCC 3560 to use, without limitation, a HIP / HCC runtime library 3558 and an HCC header 3556 to generate host executable code 3570(2) and HCC executable code 3582. In at least one embodiment, the HIP / HCC runtime library 3558 corresponds to the HIP runtime API 3532. In at least one embodiment, the HCC header 3556 includes, without limitation, any number and type of interoperability mechanisms for HIP and HCC.In at least one embodiment, the host executable code 3570(2) and the device executable code 3582 may be executed on the CPU 3590 and the GPU 3592, respectively.

[0261] Fig. 36 shows an example kernel generated by the CUDA-to-HIP translation tool 3520 Fig. 35C according to at least one embodiment. In at least one embodiment, the CUDA source code 3510 divides an overall problem that a particular kernel is to solve into relatively coarse subproblems that can be solved independently with thread blocks. In at least one embodiment, each thread block includes, without limitation, any number of threads. In at least one embodiment, each subproblem is divided into relatively small pieces that can be solved cooperatively and in parallel by threads within a thread block. In at least one embodiment, threads within a thread block can cooperate by sharing data via shared memory and synchronizing execution to coordinate memory accesses.

[0262] In at least one embodiment, the CUDA source code 3510 organizes thread blocks associated with a particular kernel into a one-dimensional, two-dimensional, or three-dimensional grid of thread blocks. In at least one embodiment, each thread block includes, without limitation, any number of threads, and a grid includes, without limitation, any number of thread blocks.

[0263] In at least one embodiment, a kernel is a function in device code defined using a "_global_" declaration identifier. In at least one embodiment, the dimension of a grid executing a kernel for a particular kernel invocation and the associated streams are specified using a CUDA kernel startup syntax 3610. In at least one embodiment, the CUDA kernel startup syntax 3610 is specified as "KemelName<<<RasterGröße, BlockGröße, GemeinsameSpeicherGröße, Stream> >>(KernelArguments);". In at least one embodiment, an execution configuration syntax is a "<<<...>>>" construct inserted between a kernel name ("KernelName") and a parenthesized list of kernel arguments ("KernelArguments"). In at least one embodiment, the CUDA kernel startup syntax 3610 includes, without limitation, a CUDA startup function syntax instead of an execution configuration syntax.

[0264] In at least one embodiment, GridSize is of type dim3 and specifies the dimension and size of a grid. In at least one embodiment, type dim3 is a CUDA-defined structure that includes, without limitation, the unsigned integers x, y, and z. In at least one embodiment, z defaults to one if z is not specified. In at least one embodiment, y defaults to one if y is not specified. In at least one embodiment, the number of thread blocks in a grid is equal to the product of GridSize.x, GridSize.y, and GridSize.z. In at least one embodiment, BlockSize is of type dim3 and specifies the dimension and size of each thread block. In at least one embodiment, the number of threads per thread block is equal to the product of BlockSize.x, BlockSize.y, and BlockSize.z.In at least one embodiment, each thread executing a kernel is given a unique thread ID that is accessible within the kernel via a built-in variable (e.g., "threadldx").

[0265] In at least one embodiment, and with respect to the CUDA kernel startup syntax 3610, "SharedMemorySize" is an optional argument specifying a number of bytes in shared memory that is dynamically allocated per thread block for a particular kernel invocation, in addition to the statically allocated memory. In at least one embodiment, and with respect to the CUDA kernel startup syntax 3610, the SharedMemorySize defaults to zero. In at least one embodiment, and with respect to the CUDA kernel startup syntax 3610, "Stream" is an optional argument specifying an associated stream and defaults to zero to indicate a standard stream. In at least one embodiment, a stream is a sequence of instructions (possibly from different host threads) executed in order.In at least one embodiment, different streams may execute instructions in different orders or concurrently.

[0266] In at least one embodiment, the CUDA source code 3510 includes, without limitation, a kernel definition for an exemplary kernel "MatAdd" and a main function. In at least one embodiment, the main function is host code executing on the host and including, without limitation, a kernel call that causes the MatAdd kernel to execute on a device. In at least one embodiment, and as shown, kernel MatAdd adds two matrices A and B of size NxN, where N is a positive integer, and stores the result in a matrix C. In at least one embodiment, the main function defines the variable threadsPerBlock as 16 by 16 and the variable numBlocks as N / 16 by N / 16. In at least one embodiment, the main function then returns the kernel call "MatAdd<<<numBlocks, threadsPerBlock> >(A, B, C);“In at least one embodiment, and in accordance with the CUDA kernel startup syntax 3610, the kernel MatAdd is executed using a grid of thread blocks of dimensions N / 16 by N / 16, where each thread block has a dimension of 16 by 16. In at least one embodiment, each thread block includes 256 threads, a grid is created with enough blocks to have one thread per matrix element, and each thread in such a grid executes kernel MatAdd to perform pairwise addition.

[0267] In at least one embodiment, during the translation of the CUDA source code 3510 into the HIP source code 3530, the CUDA-to-HIP translation tool 3520 translates each kernel call in the CUDA source code 3510 from the CUDA kernel startup syntax 3610 to a HIP kernel startup syntax 3620 and converts any number of other CUDA calls in the source code 3510 to any number of other functionally similar HIP calls. In at least one embodiment, the HIP kernel startup syntax 3620 is specified as "hipLaunchKerneIGGL(KemeIName,GridSize,BlockSize,SharedMemorySize,Stream,KernelArguments);". In at least one embodiment, KernelName, RasterSize, BlockSize, SharedMemorySize, Stream, and KernelArguments in the HIP kernel startup syntax 3620 have the same meaning as in the CUDA kernel startup syntax 3610 (described previously herein).In at least one embodiment, the SharedMemorySize and Stream arguments are required in the HIP kernel startup syntax 3620 and optional in the CUDA kernel startup syntax 3610.

[0268] In at least one embodiment, a portion of the Fig. 36 HIP source code 3530 is identical to a section of the Fig. 36, except for a kernel call that causes kernel MatAdd to execute on a device. In at least one embodiment, kernel MatAdd is defined in HIP source code 3530 with the same "_global_" declaration specifier as kernel MatAdd is defined in CUDA source code 3510. In at least one embodiment, a kernel call in HIP source code 3530 is "hipLaunchKerneIGGL(MatAdd, numBlocks, threadsPerBlock, 0, 0, A, B, C);", while a corresponding kernel call in CUDA source code 3510 is "MatAdd<<<numBlocks, threadsPerBlock> >(A, B, C);“

[0269] Fig. 37 shows the non-CUDA capable GPU 3592 from Fig. 35C in greater detail, in accordance with at least one embodiment. In at least one embodiment, GPU 3592 is developed by AMD Corporation of Santa Clara. In at least one embodiment, GPU 3592 may be configured to perform computational operations in a highly parallel manner. In at least one embodiment, GPU 3592 is configured to perform graphics pipeline operations such as draw instructions, pixel operations, geometric calculations, and other operations related to rendering an image on a display. In at least one embodiment, GPU 3592 is configured to perform non-graphics operations. In at least one embodiment, GPU 3592 is configured to perform both graphics-related and non-graphics operations.In at least one embodiment, the GPU 3592 may be configured to execute the device code included in the HIP source code 3530.

[0270] In at least one embodiment, the GPU 3592 includes, without limitation, any number of programmable processing units 3720, an instruction processor 3710, an L2 cache 3722, memory controllers 3770, DMA engines 3780(1), system memory controllers 3782, DMA engines 3780(2), and GPU controllers 3784. In at least one embodiment, each programmable processing unit 3720 includes, without limitation, a workload manager 3730 and any number of compute units 3740. In at least one embodiment, the instruction processor 3710 reads instructions from one or more instruction queues (not shown) and dispatches the instructions to the workload managers 3730. In at least one embodiment, for each programmable processing unit 3720, the associated workload manager 3730 dispatches work to the processors 3740 located in the programmable processing unit 3720 contained computing units 3740.In at least one embodiment, each compute unit 3740 may execute any number of thread blocks, but each thread block executes on a single compute unit 3740. In at least one embodiment, a workgroup is a thread block.

[0271] In at least one embodiment, each compute unit 3740 includes, without limitation, any number of SIMD units 3750 and a shared memory 3760. In at least one embodiment, each SIMD unit 3750 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, each SIMD unit 3750 includes, without limitation, a vector ALU 3752 and a processor register file 3754. In at least one embodiment, each SIMD unit 3750 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process a different set of instructions based on a single set of instructions.In at least one embodiment, predication can be used to deactivate one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp thread. In at least one embodiment, different wavefronts in a thread block can synchronize with each other and communicate via a shared memory 3760.

[0272] In at least one embodiment, the programmable processing units 3720 are referred to as "shader engines." In at least one embodiment, each programmable processing unit 3720 includes, without limitation, any amount of dedicated graphics hardware in addition to the compute units 3740. In at least one embodiment, each programmable processing unit 3720 includes, without limitation, any number (at least zero) of geometry processors, any number (at least zero) of rasterizers, any number (at least zero) of render backends, a workload manager 3730, and any number of compute units 3740.

[0273] In at least one embodiment, compute units 3740 share L2 cache 3722. In at least one embodiment, L2 cache 3722 is partitioned. In at least one embodiment, GPU memory 3790 is accessible to all compute units 3740 in GPU 3592. In at least one embodiment, memory controllers 3770 and system memory controllers 3782 facilitate data transfers between GPU 3592 and a host, and DMA engines 3780(1) enable asynchronous memory transfers between GPU 3592 and such a host. In at least one embodiment, memory controllers 3770 and GPU controllers 3784 facilitate data transfers between GPU 3592 and other GPUs 3592, and DMA engines 3780(2) enable asynchronous memory transfers between GPU 3592 and other GPUs 3592.

[0274] In at least one embodiment, the graphics processor 3592 includes, without limitation, any number and type of system interconnects that enable data and control transfers across any number and type of directly or indirectly connected components that may be internal or external to the graphics processor 3592. In at least one embodiment, the GPU 3592 includes, without limitation, any number and type of I / O interfaces (e.g., PCIe) connected to any number and type of peripheral devices. In at least one embodiment, the GPU 3592 may include, without limitation, any number (including zero) of display engines and any number (including zero) of multimedia engines.In at least one embodiment, the GPU 3592 implements a memory subsystem, including, without limitation, any number and type of memory controllers (e.g., memory controllers 3770 and system memory controllers 3782) and storage devices (e.g., shared memories 3760) that may be dedicated to a component or shared among multiple components. In at least one embodiment, the GPU 3592 implements a cache subsystem, including, without limitation, one or more cache memories (e.g., L2 cache 3722), each of which may be dedicated to or shared among any number of components (e.g., SIMD units 3750, compute units 3740, and programmable processing units 3720).

[0275] Fig. 38 shows how threads of an exemplary CUDA raster 3820 are mapped to different compute units 3740 of Fig. 37. In at least one embodiment, and for illustrative purposes only, grid 3820 has a grid size of BX times BY times 1 and a block size of TX times TY times 1. Therefore, in at least one embodiment, grid 3820 has, without limitation, (BX * BY) thread blocks 3830, and each thread block 3830 includes, without limitation, (TX * TY) threads 3840. Threads 3840 are in Fig. 38 depicted as ornate arrows.

[0276] In at least one embodiment, grid 3820 is mapped to programmable processing unit 3720(1), which includes, without limitation, compute units 3740(1)-3740(C). In at least one embodiment, and as illustrated, (BJ * BY) thread blocks 3830 are assigned to compute unit 3740(1), and the remaining thread blocks 3830 are assigned to compute unit 3740(2). In at least one embodiment, each thread block 3830 may include, without limitation, any number of warps, and each warp is mapped to a different SIMD unit 3750 of Fig. 37 shown.

[0277] In at least one embodiment, the warps in a given thread block 3830 may synchronize with each other and communicate via a shared memory 3760 that includes an associated compute unit 3740. For example, and in at least one embodiment, warps in thread block 3830(BJ,1) may synchronize with each other and communicate via shared memory 3760(1). For example, and in at least one embodiment, warps in thread block 3830(BJ+1,1) may synchronize with each other and communicate via shared memory 3760(2).

[0278] Fig. Figure 39 shows how existing CUDA code can be migrated to Data Parallel C++ code in accordance with at least one embodiment. Data Parallel C++ (DPC++) may refer to an open, standards-based, single-architecture alternative to proprietary languages ​​that allows developers to reuse code for different hardware targets (CPUs and accelerators such as GPUs and FPGAs) and also to perform custom tuning for a specific accelerator. DPC++ uses similar and / or identical C and C++ constructs in accordance with ISO C++, which developers should be familiar with. DPC++ incorporates the SYCL standard from The Khronos Group to support data parallelism and heterogeneous programming.SYCL is a cross-platform abstraction layer that builds on the underlying concepts, portability, and efficiency of OpenCL, allowing code for heterogeneous processors to be written in a single-source style using standard C++. SYCL can enable single-source development, where C++ template functions can contain both host and device code to construct complex algorithms that leverage OpenCL acceleration, and then reuse them throughout their source code for different data types.

[0279] In at least one embodiment, a DPC++ compiler is used to compile DPC++ source code that can be deployed on various hardware targets. In at least one embodiment, a DPC++ compiler is used to generate DPC++ applications that can be deployed on various hardware targets, and a DPC++ compatibility tool can be used to migrate CUDA applications into a multiplatform program in DPC++. In at least one embodiment, a DPC++ base toolkit includes a DPC++ compiler for deploying applications on various hardware targets; a DPC++ library for increasing productivity and performance on CPUs, GPUs, and FPGAs; a DPC++ compatibility tool for migrating CUDA applications into multiplatform applications; and any suitable combination thereof.

[0280] In at least one embodiment, a DPC++ programming model is used to simplify one or more aspects of programming CPUs and accelerators by using modern C++ features to express parallelism with a programming language called Data Parallel C++. The DPC++ programming language can be used to reuse code for hosts (e.g., a CPU) and accelerators (e.g., a GPU or FPGA) using a single source language, with execution and memory dependencies clearly communicated. Mappings within the DPC++ code can be used to run an application on a hardware or set of hardware devices that best accelerate a workload. A host can be available to simplify the development and debugging of device code, even on platforms that do not have an accelerator.

[0281] In at least one embodiment, the CUDA source code 3900 is provided as input to a DPC++ compatibility tool 3902 to generate human-readable DPC++ 3904. In at least one embodiment, the human-readable DPC++ 3904 includes inline comments generated by the DPC++ compatibility tool 3902 that guide a developer on how and / or where to modify the DPC++ code to complete the coding and tuning for desired performance 3906, thereby generating DPC++ source code 3908.

[0282] In at least one embodiment, the CUDA source code 3900 is or includes a collection of human-readable source code in a CUDA programming language. In at least one embodiment, the CUDA source code 3900 is human-readable source code in a CUDA programming language. In at least one embodiment, a CUDA programming language is an extension of the C++ programming language that includes, without limitation, mechanisms for defining device code and distinguishing between device code and host code. In at least one embodiment, the device code is source code that, once compiled, is executable on a device (e.g., a GPU or an FPGA) and may include one or more parallelizable workflows that may be executed on one or more processor cores of a device.In at least one embodiment, a device may be a processor optimized for parallel processing of instructions, such as a CUDA-capable GPU, GPU or other GPGPU, etc. In at least one embodiment, host code is source code that, after compilation, is executable on a host. In at least one embodiment, some or all of the host code and device code may execute in parallel on a CPU and a GPU / FPGA. In at least one embodiment, a host is a processor optimized for processing sequential instructions, such as a CPU. The method described in connection with . Fig. 39 may be consistent with the CUDA source code described elsewhere in this document.

[0283] In at least one embodiment, the DPC++ compatibility tool 3902 refers to an executable tool, program, application, or other suitable type of tool used to facilitate the migration from CUDA source code 3900 to DPC++ source code 3908. In at least one embodiment, the DPC++ compatibility tool 3902 is a command-line-based code migration tool available as part of a DPC++ toolkit and used to port existing CUDA source code to DPC++. In at least one embodiment, the DPC++ compatibility tool 3902 converts some or all of a CUDA application's source code from CUDA to DPC++ and generates a resulting file written at least partially in DPC++, referred to as human-readable DPC++ 3904.In at least one embodiment, the human-readable DPC++ 3904 includes comments generated by the DPC++ compatibility tool 3902 to indicate where user intervention might be required. In at least one embodiment, user intervention is required when the CUDA source code 3900 calls a CUDA API that does not have an analogous DPC++ API; other examples where user intervention is required are discussed in more detail later.

[0284] In at least one embodiment, a workflow for migrating CUDA source code 3900 (e.g., an application or a portion thereof) includes creating one or more compilation database files; migrating from CUDA to DPC++ using a DPC++ compatibility tool 3902; completing the migration and verifying correctness, thereby producing DPC++ source code 3908; and compiling DPC++ source code 3908 with a DPC++ compiler to produce a DPC++ application. In at least one embodiment, a compatibility tool provides a utility that intercepts instructions used in Makefile execution and stores them in a compilation database file. In at least one embodiment, a file is stored in JSON format. In at least one embodiment, an intercepted instruction converts the Makefile instruction into a DPC compatibility instruction.

[0285] In at least one embodiment, intercept-build is a helper script that intercepts a compilation process to capture compilation options, macro definitions, and include paths, and writes this data to a compilation database file. In at least one embodiment, a compilation database file is a JSON file. In at least one embodiment, the DPC++ Compatibility Tool 3902 analyzes a compilation database and applies options when migrating input sources. In at least one embodiment, the use of intercept-build is optional but is strongly recommended for Make or CMake-based environments. In at least one embodiment, a migration database includes commands, directories, and files: the command may include necessary compilation flags; the directory may include paths to header files; the file may include paths to CUDA files.

[0286] In at least one embodiment, the DPC++ compatibility tool 3902 migrates CUDA code (e.g., applications) written in CUDA to DPC++ by generating DPC++ wherever possible. In at least one embodiment, the DPC++ compatibility tool 3902 is available as part of a toolkit. In at least one embodiment, a DPC++ toolkit includes an intercept build tool. In at least one embodiment, an intercept built tool creates a compilation database that captures compilation instructions for migrating CUDA files. In at least one embodiment, a compilation database generated by an intercept built tool is used by the DPC++ compatibility tool 3902 to migrate CUDA code to DPC++. In at least one embodiment, non-CUDA C++ code and files are migrated unchanged.In at least one embodiment, the DPC++ Compatibility Tool 3902 generates human-readable DPC++ 3904, which may be DPC++ code that, as generated by the DPC++ Compatibility Tool 3902, cannot be compiled by the DPC++ compiler and requires additional digging to verify portions of the code that were not migrated correctly and may require manual intervention, for example, by a developer. In at least one embodiment, the DPC++ Compatibility Tool 3902 provides hints or tools embedded in the code to assist developers in manually migrating additional code that could not be automatically migrated. In at least one embodiment, migration is a one-time operation for a source file, project, or application.

[0287] In at least one embodiment, the DPC++ compatibility tool 39002 is capable of successfully migrating all sections of the CUDA code to DPC++, and there may be only an optional step of manually reviewing and tuning the performance of the generated DPC++ source code. In at least one embodiment, the DPC++ compatibility tool 3902 directly generates DPC++ source code 3908 that is compiled by a DPC++ compiler, without requiring or utilizing human intervention to modify the DPC++ code generated by the DPC++ compatibility tool 3902. In at least one embodiment, the DPC++ compatibility tool generates compilable DPC++ code that can optionally be tuned by a developer for performance, readability, maintainability, other various considerations, or a combination thereof.

[0288] In at least one embodiment, one or more CUDA source files are at least partially migrated to DPC++ source files using the DPC++ Compatibility Tool 3902. In at least one embodiment, the CUDA source code includes one or more header files, which may also include CUDA header files. In at least one embodiment, a CUDA source file includes a<cuda.h> header file and a<stdio.h> -Header file that can be used to print text. In at least one embodiment, a portion of a CUDA source file for a vector addition kernel can be written as or with reference to: #include<cuda.h> #include<stdio.h> #define VECTOR_SIZE 256 [] global_void VectorAddKernel(float* A, float* B, float* C) A[threadldx.x] = threadldx.x + 1.0f; B[threadldx.x] = threadldx.x + 1.0f; C[threadldx.x] = A[threadldx.x] + B[threadldx.x]; int main() { float *d_A, *d_B, *d_C; cudaMalloc(&d_A, VECTOR_SIZE*sizeof(float)); cudaMalloc(&d_B, VECTOR_SIZE*sizeof(float)); cudaMalloc(&d_C, VECTOR_SlZE*sizeof(float)); VectorAddKernel<<<1, VECTOR_SlZE>>>(d_A, d_B, d_C); float Result[VECTOR_SIZE] = {}; cudaMemcpy(Result, d_C, VECTOR_SIZE*sizeof(float), cudaMemcpyDeviceToHost); cudaFree(d_A); cudaFree(d_B); cudaFree(d_C); for (int i=0; i <VECTOR_SIZE; i++ { if (i % 16 == 0) { printf("\n");} printf("%f ", Result[i]); return 0;.

[0289] In at least one embodiment, and in conjunction with the CUDA source code file presented above, the DPC++ compatibility tool 3902 analyzes the CUDA source code and replaces the header files with appropriate DPC++ and SYCL header files. In at least one embodiment, the DPC++ header files include auxiliary declarations. In CUDA, there is the concept of a thread ID, and accordingly, in DPC++ or SYCL, there is a local identifier for each element.

[0290] In at least one embodiment, and in connection with the CUDA source file presented above, there are two vectors A and B that are initialized, and a vector addition result is placed into vector C as part of VectorAddKernel(). In at least one embodiment, the DPC++ compatibility tool 3902 converts CUDA thread IDs used to index work items to standard SYCL addressing for work items via a local ID as part of the migration from CUDA code to DPC++ code. In at least one embodiment, the DPC++ code generated by the DPC++ compatibility tool 3902 may be optimized, for example, by reducing the dimensionality of an nd_item, thereby increasing memory and / or processor utilization.

[0291] In at least one embodiment, and in conjunction with the CUDA source file presented above, memory allocation is migrated. In at least one embodiment, cudaMalloc() is migrated to a unified shared SYCL call, malloc_device(), passed a device and a context, using SYCL concepts such as platform, device, context, and queue. In at least one embodiment, a SYCL platform may include multiple devices (e.g., host and GPU devices); a device may have multiple queues to which jobs may be submitted; each device may have a context; and a context may include multiple devices and manage shared memory objects.

[0292] In at least one embodiment, and in conjunction with the CUDA source file presented above, a main() function calls VectorAddKernel() to add two vectors A and B and store the result in vector C. In at least one embodiment, the CUDA code for calling VectorAddKernel() is replaced with DPC++ code to submit a kernel to a command queue for execution. In at least one embodiment, a command group handler cgh submits data, synchronization, and computations submitted to the queue. Parallel_for is called for a number of global elements and a number of work items in that work group where VectorAddKernel() is called.

[0293] In at least one embodiment, and in conjunction with the CUDA source file presented above, the CUDA calls for copying device memory and then freeing memory for vectors A, B, and C are migrated into corresponding DPC++ calls. In at least one embodiment, the C++ code (e.g., the standard ISO C++ code for printing a vector of floating-point variables) is migrated unchanged, without being modified by the DPC++ compatibility tool 3902. In at least one embodiment, the DPC++ compatibility tool 3902 modifies the CUDA APIs for setting up the memory and / or the host calls to execute the kernel on the device for acceleration. In at least one embodiment, and in conjunction with the CUDA source file presented above, a corresponding human-readable DPC++ 3904 (which can be compiled, for example) is written as or with reference to: #include <CL / sycl.hpp> #include <dpct / dpct.hpp> #define VECTOR_SIZE 256 void VectorAddKernel(float* A, float* B, float* C, sycl::nd_item<3> item_ct1) { A[item_ct1.get_local_id(2)] = item_ct1.get_local_id(2) + 1.0f; B[item_ct1.get_local_id(2)] = item_ct1.get_local_id(2) + 1.0f; C[item_ct1.get_local_id(2)] = A[item_ct1.get_local_id(2)] + B[item_ct1.get_local_id(2)]; int main() { float *d_A, *d_B, *d_C; d_A = (float *)sycl::malloc_device(VECTOR_SIZE * sizeof(float), dpct:get current deviceQ, dpct::get_default_context()); d_B = (float *)sycl::malloc_device(VECTOR_SIZE * sizeof(float), dpct:get current deviceQ, dpct::get_default_context()); d_C = (float *)sycl::malloc_device(VECTOR_SIZE * sizeof(float), dpct::get_current_device(), dpct::get_default_context()); dpct::get_default_queue_wait().submit([&](sycl::handler &cgh) { cgh.parallel_for( sycl::nd_range<3>(sycl::range<3>(1, 1, 1) * sycl::range<3>(1, 1, VECTOR_SIZE) * sycl::range<3>(1, 1, VECTOR_SIZE)), [=](sycl::nd_items<3> item_ct1) { VectorAddKernel(d_A, d_B, d_C, item_ct1);});}); float Result[VECTOR_SlZE] = {}; dpct::get_default_queue_wait() .memcpy(Result, d_C, VECTOR_SIZE * sizeof(float)) .wait(); sycl::free(d_A, dpct:get default contextQ); sycl::free(d_B, dpct::get_default_context()); sycl::free(d_C, dpct::get_default_context()); for (int i=0; i<VECTOR_SIZE; i++ { if (i % 16 == 0) { printf("\n"); printf("%f ", Result[i]); return 0; .

[0294] In at least one embodiment, the human-readable DPC++ 3904 refers to the output generated by the DPC++ Compatibility Tool 3902 and may be optimized in one way or another. In at least one embodiment, the human-readable DPC++ 3904 generated by the DPC++ Compatibility Tool 3902 may be manually edited by a developer after migration to make it more maintainable, improve performance, or address other considerations. In at least one embodiment, the DPC++ code generated by the DPC++ Compatibility Tool 39002, such as disclosed by DPC++, may be optimized by removing the repeated calls to get_current_device() and / or get_default_context() for each malloc_device() call.In at least one embodiment, the DPC++ code generated above uses a three-dimensional nd_range that can be refactored to use only a single dimension, thereby reducing memory usage. In at least one embodiment, a developer can manually edit the DPC++ code generated by the DPC++ Compatibility Tool 3902 and replace the use of uniform shared memory with accessors. In at least one embodiment, the DPC++ Compatibility Tool 3902 includes an option to change the way CUDA code is migrated to DPC++ code. In at least one embodiment, the DPC++ Compatibility Tool 3902 is very verbose because it uses a general template for migrating CUDA code to DPC++ code that works for a large number of cases.

[0295] In at least one embodiment, a workflow for migrating from CUDA to DPC++ includes the following steps: preparing for migration using the Intercept build script; migrating CUDA projects to DPC++ using the DPC++ Compatibility Tool 3902; manually reviewing and editing the migrated source files for completeness and correctness; and compiling the final DPC++ code to produce a DPC++ application.In at least one embodiment, manual review of the DPC++ source code may be required in one or more scenarios, including, but not limited to: migrated API does not return an error code (CUDA code may return an error code that can then be consumed by the application, but SYCL uses exceptions to report errors and therefore does not use error codes to expose bugs); CUDA Compute Capability-dependent logic is not supported by DPC++; instruction could not be removed.In at least one embodiment, scenarios where DPC++ code requires manual intervention may include, without limitation, error code logic being replaced with (*,0) code or commented out; equivalent DPC++ API not available; CUDA Compute Capability dependent logic; hardware dependent API (clock()); missing features, unsupported API; execution time measurement logic; handling built-in vector type conflicts; cuBLAS API migration; and more.

[0296] In at least one embodiment, one or more of the techniques described herein use a oneAPI programming model. In at least one embodiment, a oneAPI programming model refers to a programming model for interacting with various compute accelerator architectures. In at least one embodiment, oneAPI refers to an application programming interface (API) designed to interact with various compute accelerator architectures. In at least one embodiment, a oneAPI programming model uses a DPC++ programming language. In at least one embodiment, the DPC++ programming language is a high-level language for data-parallel programming productivity. In at least one embodiment, a DPC++ programming language is based at least in part on the C and / or C++ programming languages.In at least one embodiment, a oneAPI programming model is a programming model developed by Intel Corporation of Santa Clara, CA.

[0297] In at least one embodiment, oneAPl and / or the oneAPl programming model is used to interact with various accelerator, GPU, processor, and / or variants thereof architectures. In at least one embodiment, oneAPl comprises a set of libraries implementing various functionalities. In at least one embodiment, oneAPl comprises at least one oneAPI DPC++ library, oneAPI math kernel library, oneAPl data analytics library, oneAPI deep neural network library, oneAPl collective communication library, oneAPl threading building block library, oneAPl video processing library, and / or variations thereof.

[0298] In at least one embodiment, a oneAPI DPC++ library, also referred to as oneDPL, is a library that implements algorithms and functions for accelerating DPC++ kernel programming. In at least one embodiment, oneDPL implements one or more Standard Template Library (STL) functions. In at least one embodiment, oneDPL implements one or more parallel STL functions. In at least one embodiment, oneDPL provides a set of library classes and functions, such as parallel algorithms, iterators, function object classes, range-based API, and / or variations thereof. In at least one embodiment, oneDPL implements one or more C++ Standard Library classes and / or functions. In at least one embodiment, oneDPL implements one or more random number generator functions.

[0299] In at least one embodiment, a oneAPI math kernel library, also referred to as oneMKL, is a library that implements various optimized and parallelized routines for various mathematical functions and / or operations. In at least one embodiment, oneMKL implements one or more Basic Linear Algebra Subprograms (BLAS) and / or Linear Algebra Package (LAPACK) dense linear algebra routines. In at least one embodiment, oneMKL implements one or more sparse BLAS routines for linear algebra. At least in one embodiment, oneMKL implements one or more random number generators (RNGs). In at least one embodiment, oneMKL implements one or more vector math (VM) routines for mathematical operations on vectors. In at least one embodiment, oneMKL implements one or more fast Fourier transform (FFT) functions.

[0300] In at least one embodiment, a oneAp1 data analysis library, also referred to as oneDAL, is a library that implements various data analysis applications and distributed computing. In at least one embodiment, oneDAL implements various algorithms for preprocessing, transformation, analysis, modeling, validation, and decision-making for data analysis in batch, online, and distributed processing modes of computation. In at least one embodiment, oneDAL implements various C++ and / or Java APIs and various connectors to one or more data sources. In at least one embodiment, oneDAL implements DPC++ API extensions to a traditional C++ interface and enables GPU utilization for various algorithms.

[0301] In at least one embodiment, a oneAPl deep neural network library, also referred to as oneDNN, is a library that implements various deep learning functions. In at least one embodiment, oneDNN implements various functions, algorithms, and / or variations of neural networks, machine learning, and deep learning.

[0302] In at least one embodiment, a collective oneAPl communication library, also referred to as oneCCL, is a library that implements various applications for deep learning and machine learning workloads. In at least one embodiment, oneCCL builds on lower-level communication middleware, such as Message Passing Interface (MPI) and libfabrics. In at least one embodiment, oneCCL enables a number of deep learning-specific optimizations, such as prioritization, persistent operations, out-of-order execution, and / or variations thereof. In at least one embodiment, oneCCL implements various CPU and GPU functions.

[0303] In at least one embodiment, a oneAPI Threading Building Blocks Library, also referred to as oneTBB, is a library that implements various parallelized processes for various applications. In at least one embodiment, oneTBB is used for task-based, collaborative parallel programming on a host. In at least one embodiment, oneTBB implements generic parallel algorithms. In at least one embodiment, oneTBB implements concurrent containers. In at least one embodiment, oneTBB implements a scalable memory allocator. In at least one embodiment, oneTBB implements a work-stealing task scheduler. In at least one embodiment, oneTBB implements low-level synchronization primitives. In at least one embodiment, oneTBB is compiler-independent and can be deployed on various processors, such as GPUs, PPUs, CPUs, and / or variations thereof.

[0304] In at least one embodiment, a oneAPl video processing library, also referred to as oneVPL, is a library used to accelerate video processing in one or more applications. In at least one embodiment, oneVPL implements various video decoding, encoding, and processing functions. In at least one embodiment, oneVPL implements various functions for media pipelines on CPUs, GPUs, and other accelerators. In at least one embodiment, oneVPL implements device detection and selection in media-centric and video analytics workloads. In at least one embodiment, oneVPL implements API primitives for sharing zero-copy buffers.

[0305] In at least one embodiment, a oneAPI programming model uses a DPC++ programming language. In at least one embodiment, a DPC++ programming language is a programming language that includes, without limitation, functionally similar versions of CUDA mechanisms to define device code and to distinguish between device code and host code. In at least one embodiment, a DPC++ programming language may include a subset of the functionality of a CUDA programming language. In at least one embodiment, one or more CUDA programming model operations are performed using a oneAPI programming model with a DPC++ programming language.

[0306] It should be noted that while the embodiments described herein may refer to a CUDA programming model, the techniques described herein may be used with any suitable programming model, such as HIP, oneAPI (e.g., using oneAPI-based programming to perform or implement a method disclosed herein), and / or variations thereof.

[0307] In at least one embodiment, one or more components of the systems and / or processors disclosed above may communicate with one or more CPUs, ASICs, GPUs, FPGAs, or other hardware, circuit, or integrated circuit components that may include, for example, an upscaler or upsampler for upscaling an image, an image blender or image blender component for blending, mixing, or stitching images, a sampler for sampling an image (e.g., as part of a DSP), a neural network circuit configured to execute an upscaler to upscale an image (e.g., from a low-resolution image to a high-resolution image), or other hardware to modify or generate an image, frame, or video to adjust its resolution, size, or pixels;one or more components of systems and / or processors disclosed above may use components described in this disclosure to perform methods, operations, or instructions that generate or modify an image.;

[0308] At least one embodiment of the disclosure may be described in terms of the following clauses: 1. A processor comprises: one or more circuits for implementing an API to allocate power to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. 2. The processor of clause 1, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power delivered to the first and second processors via the first and second sockets is to be allocated independently. 3. The processor of clause 1 or 2, wherein the one or more circuits are to cause power allocation to a first processor of the one or more processors based at least in part on a ratio of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. 4. The processor of any one of clauses 1 to 3, wherein the one or more circuits are to select one or more of the one or more threads to execute on a first processor of the one or more processors based at least in part on power to be allocated to the first processor. 5. The processor of any of clauses 1-4, wherein a first processor of the one or more processors comprises a plurality of processor cores. 6. The processor of any of clauses 1-5, wherein the one or more circuits are to allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface (“API”). 7. The processor of any of clauses 1-6, wherein power allocation between at least a first and a second processor of the one or more processors is based at least in part on maintaining a power budget. 8. A system comprising: one or more circuits for implementing an API to allocate power to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. 9. The system of clause 8, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power supplied to the first and second processors via the first and second sockets is to be allocated independently. 10. The system of clause 8 or 9, wherein the one or more circuits are to cause power allocation to a first processor of the one or more processors based at least in part on a proportion of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. 11. The system of any of clauses 8-10, wherein the one or more circuits select one or more of the one or more threads to execute on a first processor of the one or more processors based at least in part on power to be allocated to the first processor. 12. A system as described in any of clauses 8-11, wherein a first processor of the one or more processors comprises a plurality of processor cores. 13. The system of any of clauses 8-12, wherein the one or more circuits are to allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface (“API”). 14. The system of any of clauses 8-13, wherein power allocation between at least a first and a second processor of the one or more processors is based at least in part on maintaining a power budget. 15. A method comprising: using a processor including one or more circuits to execute an API to allocate power to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. 16. The method of any one of clauses 15, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power supplied to the first and second processors via the first and second sockets is to be allocated independently. 17. The method of clause 15 or 16, wherein the one or more circuits are to cause power allocation to a first processor of the one or more processors based at least in part on a proportion of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. 18. The method as described in any of clauses 15-17, wherein the one or more circuits select one or more of the one or more threads to execute on a first processor of the one or more processors based at least in part on power to be allocated to the first processor. 19. The method as described in any one of clauses 15-18, wherein a first processor of the one or more processors comprises a plurality of processor cores. 20. The method of any of clauses 15-19, wherein the one or more circuits allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface (“API”).

[0309] Other variations are within the spirit of the present disclosure. Thus, while disclosed techniques are susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. It should be understood, however, that the disclosure is not intended to limit the disclosure to any particular form or forms, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure as defined by the appended claims.

[0310] The use of the terms "a" and "an" and "the" and similar designations in connection with the description of disclosed embodiments (particularly in connection with the following claims) is to be construed to include both the singular and the plural, unless otherwise stated herein or clearly contradicted by context, and not as a definition of any term. The terms "comprising," "having," and "including," unless otherwise stated, are to be understood as open-ended terms (i.e., "including, but not limited to"). The term "connected," when left unchanged and referring to physical connections, is to be understood as partially or wholly contained in, attached to, or connected to a component, even if something in between.Reference to ranges of values ​​is intended merely as a shorthand method for referring individually to each value falling within the range, unless otherwise noted here, and each value is included in the specification as if listed individually herein. Use of the term "set" (e.g., "a set of items") or "subset" is to be understood, unless otherwise noted or contradicted by context, as a non-empty collection comprising one or more elements. Furthermore, unless otherwise noted or contradicted by context, the term "subset" of a corresponding set is not necessarily to be understood as a proper subset of the corresponding set; rather, subset and corresponding set may be the same.

[0311] Subjunctive expressions, such as sentences of the form "at least one of A, B, and C" or "at least one of A, B, and C," are, unless explicitly stated otherwise or clearly contradicted by the context, understood with the context as they are generally used to represent that an element, term, etc., can be either A or B or C or any non-empty subset of the set of A and B and C. In the illustrative example of a set having three members, the subjunctive expressions "at least one of A, B, and C" and "at least one of A, B, and C" refer to one of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such subjunctive formulations are not intended to generally imply that, in certain embodiments, at least one of A, at least one of B, and at least one of C must be present.Unless otherwise noted or contradicted by the context, the term "plurality" indicates a state of plurality (for example, "a plurality of items" indicates multiple items). The number of items in a plurality is at least two, but may be more if indicated either explicitly or by the context. Furthermore, unless otherwise noted or otherwise clear from the context, "based on" means "at least partly based on" and not "exclusively based on."

[0312] The operations of the processes described herein may be performed in any suitable order, unless otherwise specified herein or apparent from the context. In at least one embodiment, a process such as the processes described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are collectively executed on one or more processors, by hardware, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electrical or electromagnetic transmission) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within transient signal transceivers. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media on which executable instructions (or other memory for storing executable instructions) are stored that, when executed by one or more processors of a computer system (e.g., as a result of execution), cause the computer system to perform operations described herein.In at least one embodiment, a set of non-transitory computer-readable storage media comprises a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media of the plurality of non-transitory computer-readable storage media lack all code, while a plurality of non-transitory computer-readable storage media collectively store all code. In at least one embodiment, executable instructions are executed such that different instructions are executed by different processors—for example, a non-transitory computer-readable storage medium stores instructions, and a central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions.In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of instructions.

[0313] Accordingly, in at least one embodiment, computer systems are configured to implement one or more services that individually or collectively perform operations of the processes described herein, and such computer systems are configured with applicable hardware and / or software that enable the operations to be performed. Further, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, it is a distributed computer system that includes multiple devices that operate differently such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.

[0314] The use of examples or exemplary language (e.g., "such as") is intended only to better illustrate embodiments of the disclosure and does not limit the scope of the disclosure unless otherwise stated. No language in the specification should be construed to indicate any unclaimed element as essential to the practice of the disclosure.

[0315] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference to the same extent as if each reference were individually and expressly indicated to be incorporated by reference and set forth herein in its entirety.

[0316] The terms "coupled" and "connected," and their derivatives, may be used in the description and claims. It should be understood that these terms are not synonymous. Rather, in certain examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0317] Unless explicitly stated otherwise, terms such as "processing", "computation", "calculating", "determining" or the like throughout the specification can be assumed to refer to actions and / or processes of a computer or computer system or similar electronic computing device that manipulate data represented as physical, e.g. electronic, quantities in the registers and / or memories of the computer system and / or convert that data into other data similarly represented as physical quantities in the memories, registers or other such devices for storing, transferring or displaying information of the computer system.

[0318] Similarly, the term "processor" may refer to a device or a portion of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, the term "software" processes may include, for example, software and / or hardware units that perform work over time, such as tasks, threads, and intelligent agents. Each process may also refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently.The terms “system” and “method” are used interchangeably here, insofar as a system can comprise one or more methods and methods can be considered as a system.

[0319] In at least one embodiment, an arithmetic logic unit is a set of combinational logic circuits that process one or more inputs to produce a result. In at least one embodiment, an arithmetic logic unit is used by a processor to perform mathematical operations such as addition, subtraction, or multiplication. In at least one embodiment, an arithmetic logic unit is used to implement logical operations such as logical AND / OR or XOR. In at least one embodiment, an arithmetic logic unit is stateless and consists of physical switching components such as semiconductor transistors arranged to form logic gates. In at least one embodiment, an arithmetic logic unit may operate internally as a stateful logic circuit with an associated clock.In at least one embodiment, an arithmetic logic unit may be implemented as an asynchronous logic circuit with internal state that is not maintained in an associated register set. In at least one embodiment, an arithmetic logic unit is used by a processor to combine operands stored in one or more registers of the processor and generate an output that may be stored by the processor in another register or a memory location.

[0320] In at least one embodiment, as a result of processing an instruction fetched by the processor, the processor provides one or more inputs or operands to an arithmetic logic unit, causing the arithmetic logic unit to generate a result based at least in part on an instruction code provided to the inputs of the arithmetic logic unit. In at least one embodiment, the instruction codes provided by the processor to the ALU are based at least in part on the instruction executed by the processor. In at least one embodiment, combinational logic in the ALU processes the inputs and generates an output that is placed on a bus within the processor.In at least one embodiment, the processor selects a destination register, memory location, device, or output memory location on the output bus such that the clocking of the processor causes the results generated by the ALU to be sent to the desired location.

[0321] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. The process of obtaining, acquiring, receiving, or inputting analog and digital data may be performed in a variety of ways, such as by receiving data as a parameter to a function call or an application programming interface call. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be performed by transmitting data over a serial or parallel interface. In another implementation, the process of acquiring, receiving, or inputting analog or digital data may be performed by transmitting data over a computer network from the providing entity to the acquiring entity.It may also refer to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be performed by passing data as an input or output parameter of a function call, a parameter of an application programming interface, or an interprocess communication mechanism.

[0322] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of this disclosure. Furthermore, although specific distributions of responsibilities are defined above for discussion purposes, various functions and responsibilities may be distributed and allocated in different ways depending on the circumstances.

[0323] Although the subject matter has been described in language referring to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, certain features and acts are disclosed as exemplary forms of implementing the claims.

Claims

[1] Processor comprising: one or more circuits for implementing an API to allocate power to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. [2] The processor of claim 1, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power supplied to the first and second processors via the first and second sockets is independently allocable. [3] The processor of claim 1, wherein the one or more circuits are to cause power to be allocated to a first processor of the one or more processors based at least in part on a proportion of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. [4] The processor of claim 1, wherein the one or more circuits select one or more of the one or more threads to execute on a first processor of the one or more processors based at least in part on power to be allocated to the first processor. [5] The processor of claim 1, wherein a first processor of the one or more processors comprises a plurality of processor cores. [6] The processor of claim 1, wherein the one or more circuits are to allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface ("API"). [7] The processor of claim 1, wherein the power allocation between at least a first and a second processor of the one or more processors is based at least in part on maintaining a power budget. [8] System comprising: one or more circuits to execute an API to allocate power to one or more processors based, at least in part, on one or more indications of the priority of one or more threads to be executed by the one or more processors. [9] The system of claim 8, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power supplied to the first and second processors via the first and second sockets is independently allocable. [10] The system of claim 8, wherein the one or more circuits are to cause power to be allocated to a first processor of the one or more processors based at least in part on a proportion of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. [11] The system of claim 8, wherein the one or more circuits are to select one or more of the one or more threads to execute on a first processor of the one or more processors based at least in part on the power to be allocated to the first processor. [12] The system of claim 8, wherein a first processor of the one or more processors comprises a plurality of processor cores. [13] The system of claim 8, wherein the one or more circuits are to allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface ("API"). [14] The method of claim 8, wherein the power allocation between at least a first and a second processor of the one or more processors is based at least in part on maintaining a power budget. [15] Procedure comprising: Use of a processor including one or more circuits to execute an API to allocate performance to one or more processors based at least in part on one or more indications of the priority of one or more threads to be executed by the one or more processors. [16] The method of claim 15, wherein a first processor of the one or more processors is connected to a device via a first socket and a second processor of the one or more processors is connected to the device via a second socket, and wherein power supplied to the first and second processors via the first and second sockets is to be allocated independently. [17] The method of claim 15, wherein the one or more circuits are to cause power allocation to a first processor of the one or more processors based at least in part on a proportion of high priority threads to be executed by the first processor compared to low priority threads to be executed by the first processor. [18] The method of claim 15, wherein the one or more circuits are to select one or more of the one or more threads for execution on a first processor of the one or more processors based at least in part on the power to be allocated to the first processor. [19] The method of claim 15, wherein a first processor of the one or more processors comprises a plurality of processor cores. [20] The method of claim 15, wherein the one or more circuits allocate power to the one or more processors based at least in part on utilization information provided via an application programming interface (“API”).