Techniques for modifying an executable graph to implement a workload associated with a new task graph

By modifying the executable graph to meet the needs of the new task graph, the problems of low resource utilization and poor performance in the prior art are solved, and more efficient workload implementation is achieved.

CN112817738BActive Publication Date: 2025-06-27NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011280064.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-31
Filing Date
2020-11-16
Publication Date
2025-06-27
Estimated Expiration
2041-04-28

AI Technical Summary

Technical Problem

The prior art has difficulty effectively using executable graphs to implement workloads associated with new task graphs, resulting in underutilization of resources and poor performance.

Method used

By modifying the executable graph, it can dynamically adapt to the needs of the new task graph, enabling flexible reconfiguration and optimization of workloads.

Benefits of technology

Improve resource utilization and system performance, enabling more efficient implementation of workloads associated with the new task map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112817738B_ABST
    Figure CN112817738B_ABST
Patent Text Reader

Abstract

The present disclosure relates to techniques for modifying an executable graph to implement a workload associated with a new task graph. Techniques for modifying an executable graph to implement a different workload. In at least one embodiment, the executable version of a first task graph is modified by applying a non-executable version of a second task graph to the executable version of the first task graph such that the executable version of the first task graph can implement a second workload of the non-executable version of the second task graph.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of Indian Provisional Patent Application titled "TECHNIQUES FOR MODIFYING AN EXECUTABLE GRAPH TO PERFORM A WORKLOAD ASSOCIATED WITH A NEW TASK GRAPH" filed on November 15, 2019 and having Serial No. 201941046676. The subject matter of the related application is hereby incorporated herein by reference. Technical Field

[0003] Described are various different embodiments which generally relate to parallel computing and, more particularly, to techniques for modifying an executable graph to perform a workload associated with a new task graph. Background Art

[0004] A task graph is a useful tool for modeling a workload corresponding to parallel computing tasks. In a task graph, computing tasks are modeled as nodes and dependencies between computing tasks are modeled as directed edges. To be available for causing one or more computing resources to perform the workload, the task graph is converted into an executable graph. The workload is then performed by configuring each of one or more computing resources in accordance with the executable graph to perform tasks, transferring data to one or more computing resources if necessary, and retrieving results from one or more computing resources. An executable graph can be used multiple times to cause the same workload to be performed multiple times. However, a given executable graph is typically available for only a single workload of the task graph from which it was created.

[0005] As previously mentioned, what is needed in the art are techniques for more effectively using executable graphs to perform workloads associated with new task graphs. Brief Description of the Drawings

[0006] In order to understand in a more detailed manner the ways in which the above-described features of the various different embodiments can be, some of the embodiments are illustrated in the drawings. However, it should be noted that the drawings merely illustrate typical embodiments of the inventive concept and should not be considered to limit its scope, as the invention can admit other equally effective embodiments.

[0007] Figure 1 Illustrates an exemplary data center in accordance with at least one embodiment;

[0008] Figure 2 The figure shows a processing system in accordance with at least one embodiment;

[0009] Figure 3 The figure shows a computer system in accordance with at least one embodiment;

[0010] Figure 4 The figure shows a system in accordance with at least one embodiment;

[0011] Figure 5 The figure shows an exemplary integrated circuit in accordance with at least one embodiment;

[0012] Figure 6 The figure shows a computing system in accordance with at least one embodiment;

[0013] Figure 7 The figure shows an APU in accordance with at least one embodiment;

[0014] Figure 8 The figure shows a CPU in accordance with at least one embodiment;

[0015] Figure 9 The figure shows an exemplary accelerator integrated slice in accordance with at least one embodiment;

[0016] Figures 10A - 10B The figure shows an exemplary graphics processor in accordance with at least one embodiment;

[0017] Figure 11A The figure shows a graphics core in accordance with at least one embodiment;

[0018] Figure 11B The figure shows a GPGPU in accordance with at least one embodiment;

[0019] Figure 12A The figure shows a parallel processor in accordance with at least one embodiment;

[0020] Figure 12B The figure shows a processing cluster in accordance with at least one embodiment;

[0021] Figure 12C The figure shows a graphics multiprocessor in accordance with at least one embodiment;

[0022] Figure 13 The figure shows a graphics processor in accordance with at least one embodiment;

[0023] Figure 14 The figure shows a processor in accordance with at least one embodiment;

[0024] Figure 15 The figure shows a processor in accordance with at least one embodiment;

[0025] Figure 16The figure shows a graphics processing unit core according to at least one embodiment;

[0026] Figure 17 The figure shows a PPU according to at least one embodiment;

[0027] Figure 18 The figure shows a GPC according to at least one embodiment;

[0028] Figure 19 The figure shows a streaming multiprocessor according to at least one embodiment;

[0029] Figure 20 The figure shows the software stack of a programming platform according to at least one embodiment;

[0030] Figure 21 The figure shows the CUDA implementation of the software stack of graph STACK (stack) A according to at least one embodiment;

[0031] Figure 22 The figure shows the ROCm implementation of the software stack of graph STACK A according to at least one embodiment;

[0032] Figure 23 The figure shows the OpenCL implementation of the software stack of graph STACK A according to at least one embodiment;

[0033] Figure 24 The figure shows the software supported by a programming platform according to at least one embodiment;

[0034] Figure 25 The figure shows compiling according to at least one embodiment Figures 20 - 23 the code executed on a programming platform;

[0035] Figure 26 The figure shows in more detail compiling according to at least one embodiment Figures 20 - 23 the code executed on a programming platform;

[0036] Figure 27 The figure shows transforming source code before compiling according to at least one embodiment;

[0037] Figure 28A The figure shows a system configured to compile and execute CUDA source code using different types of processing units according to at least one embodiment;

[0038] Figure 28B The figure shows a system configured to compile and execute Figure 28A the CUDA source code using a CPU and a CUDA-enabled GPU according to at least one embodiment;

[0039] Figure 28CThe figure shows a system configured to compile and execute CUDA source code using a CPU and a non-CUDA-enabled GPU according to at least one embodiment; Figure 28A ;

[0040] Figure 29 The figure shows an exemplary kernel transformed by a CUDA-to-HIP conversion tool according to at least one embodiment; Figure 28C ;

[0041] Figure 30 The figure shows in more detail a non-CUDA-enabled GPU according to at least one embodiment; Figure 28C ;

[0042] Figure 31 The figure shows how threads of an exemplary CUDA grid are mapped to different compute units according to at least one embodiment; Figure 30 ;

[0043] Figure 32 The figure shows a flowchart according to at least one embodiment;

[0044] Figure 33A -C The figure shows an exemplary task graph according to at least one embodiment;

[0045] Figure 34 The figure shows a flowchart according to at least one embodiment; and

[0046] Figure 35 The figure shows a flowchart according to at least one embodiment. DETAILED DESCRIPTION

[0047] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one of ordinary skill in the art that these inventive concepts may be practiced without one or more of these specific details.

[0048] Data Center

[0049] Figure 1 The figure shows an exemplary data center 100 according to at least one embodiment. In at least one embodiment, data center 100 includes, but is not limited to, a data center infrastructure layer 110, a framework layer 120, a software layer 130, and an application layer 140.

[0050] In at least one embodiment, as Figure 1As shown, the data center infrastructure layer 110 may include a resource coordinator 112, grouped computing resources 114, and node computing resources ("node C.R.") 116(1)-116(N), where "N" represents any integer, positive integer. In at least one embodiment, the node C.R. 116(1)-116(N) may include, but is not limited to, any number of central processing units ("CPU") or other processors (including accelerators, field programmable gate arrays ("FPGA"), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VM"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node C.R. from among 116(1)-116(N) may be servers having one or more of the computing resources mentioned above.

[0051] In at least one embodiment, the grouped computing resources 114 may include many racks (not shown) housed in data centers at different geographical locations or separate groupings of node C.R. (also not shown) housed within one or more racks. Separate groupings of node C.R. within the grouped computing resources 114 may include grouped computing, network, memory, or storage resources, which may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. including CPUs or processors may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any combination of any number of power modules, cooling modules, and network switches.

[0052] In at least one embodiment, the resource coordinator 112 may configure or otherwise control one or more of the node C.R. 116(1)-116(N) and / or the grouped computing resources 114. In at least one embodiment, the resource coordinator 112 may include a software design infrastructure ("SCT") management entity for the data center 100. In at least one embodiment, the resource coordinator 112 may include hardware, software, or some combination thereof.

[0053] In at least one embodiment, as Figure 1As shown, the framework layer 120 includes, but is not limited to, a job scheduler 132, a configuration manager 134, a resource manager 136, and a distributed file system 138. In at least one embodiment, the framework layer 120 may include a framework that supports software 152 of the software layer 130 and / or one or more applications 142 of the application layer 140. In at least one embodiment, the software 152 or the application 142 may respectively include web-based service software or applications, such as software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 120 may be, but is not limited to, a type of free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark"), which may use the distributed file system 138 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 132 may include a Spark driver that facilitates the scheduling of workloads supported by the layers of the data center 100. In at least one embodiment, the configuration manager 134 may be able to configure different layers such as the software layer 130 and the framework layer 120, including Spark and the distributed file system 138 for supporting large-scale data processing. In at least one embodiment, the resource manager 136 may be able to manage cluster or grouped computing resources that are mapped to or allocated for supporting the distributed file system 138 and the job scheduler 132. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 114 at the data center infrastructure layer 110. In at least one embodiment, the resource manager 136 may cooperate with the resource coordinator 112 to manage these mapped or allocated computing resources.

[0054] In at least one embodiment, the software 152 included in the software layer 130 may include software used by at least part of nodes C.R. 116(1)-116(N), the grouped computing resources 114, and / or the distributed file system 138 of the framework layer 120. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0055] In at least one embodiment, the applications 142 included in the application layer 140 may include one or more types of applications used by at least part of nodes C.R. 116(1)-116(N), the grouped computing resources 114, and / or the distributed file system 138 of the framework layer 120. One or more types of applications may include, but are not limited to, CUDA applications.

[0056] In at least one embodiment, any one of the configuration manager 134, the resource manager 136, and the resource coordinator 112 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions can save the data center operator of the data center 100 from making potentially bad configuration decisions and potentially avoid underutilized and / or poorly performing portions of the data center.

[0057] Computer-based system

[0058] The following figures non-limitingly illustrate exemplary computer-based systems that can be used to implement at least one embodiment.

[0059] Figure 2 FIG. illustrates a processing system 200 in accordance with at least one embodiment. In at least one embodiment, the processing system 200 includes one or more processors 202 and one or more graphics processors 208, and can be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 202 or processor cores 207. In at least one embodiment, the processing system 200 is a processing platform incorporated into a system-on-chip (“SoC”) used in mobile, handheld, or embedded devices.

[0060] In at least one embodiment, the processing system 200 can be included in or incorporated into a server-based gaming platform, a game console, a media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, the processing system 200 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 200 can also include, be coupled to, or be integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 200 is a television or set-top box device having one or more processors 202 and a graphics interface generated by one or more graphics processors 208.

[0061] In at least one embodiment, each of one or more processors 202 includes one or more processor cores 207 that process instructions which, when executed, implement operations for system and user software. In at least one embodiment, each of one or more processor cores 207 is configured to process a particular instruction set 209. In at least one embodiment, the instruction set 209 can facilitate complex instruction set computing (“CISC”), reduced instruction set computing (“RISC”), or computing via very long instruction words (“VLIW”). In at least one embodiment, each of the processor cores 207 can process a different instruction set 209, which can include instructions that facilitate emulating other instruction sets. In at least one embodiment, the processor cores 207 can also include other processing devices, such as a digital signal processor (“DSP”).

[0062] In at least one embodiment, the processor 202 includes a cache memory (“cache”) 204. In at least one embodiment, the processor 202 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various different components of the processor 202. In at least one embodiment, the processor 202 also uses an external cache (e.g., a level 3 (“L3”) cache or a last level cache (“LLC”)) (not shown), which can be shared among the processor cores 207 using known cache coherence techniques. In at least one embodiment, a register file 206 is additionally included in the processor 202, which can include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 206 can include general purpose registers or other registers.

[0063] In at least one embodiment, one or more processors 202 are coupled to one or more interface buses 210 to transfer communication signals such as address, data, or control signals between the processors 202 and other components in the processing system 200. In at least one embodiment, the interface bus 210 can be a processor bus, such as a version of the Direct Media Interface (“DMI”) bus, in one embodiment. In at least one embodiment, the interface bus 210 is not limited to the DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., “PCI”, PCI Express (“PCIe”), memory bus, or other types of interface buses). In at least one embodiment, the processor 202 includes an integrated memory controller 216 and a Platform Controller Hub 230. In at least one embodiment, the memory controller 216 facilitates communication between the memory devices of the processing system 200 and other components, while the Platform Controller Hub (“PCH”) 230 provides connections to input / output (“I / O”) devices via a local I / O bus.

[0064] In at least one embodiment, the memory device 220 can be a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, a Phase Change Memory device, or some other memory device with suitable performance for use as a processor memory. In at least one embodiment, the memory device 220 can operate as the system memory for the processing system 200 to store data and instructions 221 used when one or more processors 202 execute an application or process. In at least one embodiment, the memory controller 216 is also coupled to an optional external graphics processor 212, which can communicate with one or more graphics processors 208 in the processor 202 to perform graphics and media operations. In at least one embodiment, a display device 211 can be connected to the processor 202. In at least one embodiment, the display device 211 can include one or more of an internal display device such as in a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 211 can include a Head-Mounted Display (“HMD”), such as a stereoscopic display device used in Virtual Reality (“VR”) applications or Augmented Reality (“AR”) applications.

[0065] In at least one embodiment, the platform controller hub 230 enables peripherals to be connected to the memory device 220 and the processor 202 via a high-speed I / O bus. In at least one embodiment, the I / O includes, but is not limited to, an audio controller 246, a network controller 234, a firmware interface 228, a wireless transceiver 226, a touch sensor 225, a data storage device 224 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 224 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as PCI or PCIe. In at least one embodiment, the touch sensor 225 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 226 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as 3G, 4G, or a Long Term Evolution (“LTE”) transceiver. In at least one embodiment, the firmware interface 228 allows for communication with the system firmware and can be, for example, the Unified Extensible Firmware Interface (“UEFI”). In at least one embodiment, the network controller 234 can allow for a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 210. In at least one embodiment, the audio controller 246 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 200 includes an optional legacy I / O controller 240 for coupling legacy (e.g., Personal System 2 (“PS / 2”)) devices to the processing system 200. In at least one embodiment, the platform controller hub 230 can also be connected to one or more Universal Serial Bus (“USB”) controllers 242 to connect input devices such as a keyboard and mouse 243 combination, a camera 244, or other USB input devices.

[0066] In at least one embodiment, instances of the memory controller 216 and the platform controller hub 230 can be integrated into a discrete external graphics processor such as the external graphics processor 212. In at least one embodiment, the platform controller hub 230 and / or the memory controller 216 can be external to one or more processors 202. For example, in at least one embodiment, the processing system 200 can include an external memory controller 216 and a platform controller hub 230, which can be a memory controller hub and a peripheral controller hub within a system chipset configured to communicate with the processor 202.

[0067] Figure 3The figure shows a computer system 300 in accordance with at least one embodiment. In at least one embodiment, the computer system 300 can be a system with interconnected devices and components, a system-on-a-chip (SOC), or some combination thereof. In at least one embodiment, the computer system 300 is formed by a processor 302 that can include execution units for executing instructions. In at least one embodiment, the computer system 300 can include, but is not limited to, components such as the processor 302 to implement algorithms for processing data using execution units that include logic. In at least one embodiment, the computer system 300 can include a processor, such as a processor family, XeonTM, XScaleTM, and / or StrongARMTM, CoreTM, NervanaTM microprocessor, obtainable from Intel Corporation in Santa Clara, California, but other systems can also be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, the computer system 300 can execute a certain version of the WINDOWS operating system obtainable from Microsoft Corporation in Redmond, Washington, but other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.

[0068] In at least one embodiment, the computer system 300 can be used in other devices and embedded applications such as handheld devices. Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors (DSPs), SOCs, network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system that can implement one or more instructions.

[0069] In at least one embodiment, the computer system 300 can include, but is not limited to, the processor 302, which can include, but is not limited to, being configured to execute Compute Unified Device Architecture (“CUDA”) ( (developed by NVIDIA Corporation of Santa Clara, California) one or more execution units 308 of the program. In at least one embodiment, the CUDA program is at least a part of a software application written in the CUDA programming language. In at least one embodiment, the computer system 300 is a single-processor desktop or server system. In at least one embodiment, the computer system 300 can be a multi-processor system. In at least one embodiment, the processor 302 can include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing an instruction set combination, or any other processor device such as, for example, a digital signal processor. In at least one embodiment, the processor 302 can be coupled to a processor bus 310, which can transmit data signals between the processor 302 and other components in the computer system 300.

[0070] In at least one embodiment, the processor 302 can include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 304. In at least one embodiment, the processor 302 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory can reside outside the processor 302. In at least one embodiment, the processor 302 can also include a combination of internal and external caches. In at least one embodiment, the register file 306 can store different types of data in various different registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0071] In at least one embodiment, the execution units 308, including but not limited to the logic implementing integer and floating-point operations, also reside in the processor 302. The processor 302 can also include a microcode (“ucode”) read-only memory (“ROM”) storing microcode for certain macro instructions. In at least one embodiment, the execution units 308 can include logic for processing a packed instruction set 309. In at least one embodiment, by including a packed instruction set 309 in the instruction set of the general-purpose processor 302, together with the associated circuitry for executing instructions, operations used by many multimedia applications can be implemented using packed data in the general-purpose processor 302. In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by performing operations on packed data using a full-width processor data bus, which can eliminate the need to transfer smaller data units across the processor data bus to perform one or more operations on data elements one at a time.

[0072] In at least one embodiment, execution unit 308 can also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 300 can include, but is not limited to, memory 320. In at least one embodiment, memory 320 can be implemented as a DRAM device, an SRAM device, a flash memory device, or other memory devices. Memory 320 can store instructions 319 executable by processor 302 and / or data 321 represented by data signals.

[0073] In at least one embodiment, the system logic chip can be coupled to processor bus 310 and memory 320. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub (“MCH”) 316, and processor 302 can communicate with MCH 316 via processor bus 310. In at least one embodiment, MCH 316 can provide a high-bandwidth memory path 318 to memory 320, which is used for instruction and data storage and for storing graphics commands, data, and textures. In at least one embodiment, MCH 316 can direct data signals among processor 302, memory 320, and other components in computer system 300, and bridge data signals among processor bus 310, memory 320, and system I / O 322. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 316 can be coupled to memory 320 through high-bandwidth memory path 318, and graphics / video card 312 can be coupled to MCH 316 through an Accelerated Graphics Port (“AGP”) interconnect 314.

[0074] In at least one embodiment, computer system 300 can use system I / O 322 to couple MCH 316 to an I / O controller hub (“ICH”) 330, and the system I / O is a proprietary hub interface bus. In at least one embodiment, ICH 330 can provide a direct connection to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus for connecting peripherals to memory 320, the chipset, and processor 302. Examples can include, but are not limited to, audio controller 329, firmware hub (“flash BIOS”) 328, wireless transceiver 326, data storage device 324, legacy I / O controller 323 including user input interface 325 and keyboard interface, serial expansion port 327 such as USB, and network controller 334. Data storage device 324 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage devices.

[0075] In at least one embodiment, Figure 3 FIG. illustrates a system that includes interconnected hardware devices or "chips". In at least one embodiment, Figure 3 An exemplary SoC may be illustrated. In at least one embodiment, Figure 3 The devices illustrated therein may be interconnected using proprietary interconnects, standardized interconnects (such as PCIe), or some combination thereof. In at least one embodiment, one or more components of system 300 are interconnected using Compute Express Link ("CXL") interconnects.

[0076] Figure 4 FIG. illustrates system 400 in accordance with at least one embodiment. In at least one embodiment, system 400 is an electronic device that utilizes processor 410. In at least one embodiment, system 400 may be, for example and without limitation, a notebook, tower server, rack server, blade server, laptop, desktop computer, tablet, mobile device, phone, embedded computer, or any other suitable electronic device.

[0077] In at least one embodiment, system 400 may include, but is not limited to, processor 410 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 410 is coupled using a bus or interface such as an I2C bus, System Management Bus ("SMBus"), Low Pin Count ("LPC") bus, Serial Peripheral Interface ("SPI"), High Definition Audio ("HDA") bus, Serial Advanced Technology Attachment ("SATA") bus, USB (versions 1, 2, 3), or Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 4 FIG. illustrates a system that includes interconnected hardware devices or "chips". In at least one embodiment, Figure 4 An exemplary SoC may be illustrated. In at least one embodiment, Figure 4 The devices illustrated therein may be interconnected using proprietary interconnects, standardized interconnects (such as PCIe), or some combination thereof. In at least one embodiment, Figure 4 one or more components of are interconnected using CXL interconnects.

[0078] In at least one embodiment, Figure 4It may include a display 424, a touch screen 425, a touchpad 430, a Near Field Communication unit (“NFC”) 445, a sensor hub 440, a thermal sensor 446, a Fast Chipset (“EC”) 435, a Trusted Platform Module (“TPM”) 438, a BIOS / Firmware / Flash (“BIOS, FW Flash”) 422, a DSP 460, a Solid State Drive (“S23”) or a Hard Disk Drive (“HDD”) 420, a Wireless Local Area Network unit (“WLAN”) 450, a Bluetooth unit 452, a Wireless Wide Area Network unit (“WWAN”) 456, a Global Positioning System (“GPS”) 455, a camera such as a USB 3.0 camera (“USB 3.0 camera”) 454, or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 415 implemented, for example, to the LPDDR3 standard. Each of these components may be implemented in any suitable manner.

[0079] In at least one embodiment, other components may be communicatively coupled to the processor 410 via the components discussed above. In at least one embodiment, an accelerometer 441, an Ambient Light Sensor (“ALS”) 442, a compass 443, and a gyroscope 444 may be communicatively coupled to the sensor hub 440. In at least one embodiment, a thermal sensor 439, a fan 437, a keyboard 446, and a touchpad 430 may be communicatively coupled to the EC 435. In at least one embodiment, a speaker 463, a headset 464, and a microphone (“mic”) 465 may be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 464, which in turn may be communicatively coupled to the DSP 460. In at least one embodiment, the audio unit 464 may include, for example and without limitation, an audio encoder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a Subscriber Identity Module (“SIM”) 457 may be communicatively coupled to the WWAN unit 456. In at least one embodiment, components such as the WLAN unit 450, the Bluetooth unit 452, and the WWAN unit 456 may be implemented in a next-generation form factor (“NGFF”).

[0080] Figure 5The figure illustrates an exemplary integrated circuit 500 in accordance with at least one embodiment. In at least one embodiment, the exemplary integrated circuit 500 is a SoC that can be fabricated using one or more IP cores. In at least one embodiment, the integrated circuit 500 includes one or more application processors 505 (e.g., CPUs), at least one graphics processor 510, and may additionally include an image processor 515 and / or a video processor 520, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 500 includes peripherals or bus logic, including a USB controller 525, a UART controller 530, an SPI / SDIO controller 535, and an I2S / I2C controller 540. In at least one embodiment, the integrated circuit 500 may include a display device 545 coupled to one or more of a high-definition multimedia interface (“HDMI”) controller 550 and a mobile industry processor interface (“MIPI”) display interface 555. In at least one embodiment, storage may be provided by a flash memory subsystem 560 that includes a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 565 to access a 23RAM or SRAM memory device. In at least one embodiment, some integrated circuits additionally include an embedded security engine 570.

[0081] Figure 6 The figure illustrates a computing system 600 in accordance with at least one embodiment; in at least one embodiment, the computing system 600 includes a processing subsystem 601 having one or more processors 602 and a system memory 604 that communicate via an interconnect path that may include a memory hub 605. In at least one embodiment, the memory hub 605 may be a separate component within a chipset component or may be integrated into one or more of the processors 602. In at least one embodiment, the memory hub 605 is coupled to an I / O subsystem 611 via a communication link 606. In at least one embodiment, the I / O subsystem 611 includes an I / O hub 607 that may enable the computing system 600 to receive input from one or more input devices 608. In at least one embodiment, the I / O hub 607 may enable a display controller that may be included in one or more of the processors 602 to provide output to one or more display devices 610A. In at least one embodiment, one or more of the display devices 610A coupled to the I / O hub 607 may include a local, internal, or embedded display device.

[0082] In at least one embodiment, the processing subsystem 601 includes one or more parallel processors 612 coupled to a memory hub 605 via a bus or other communication link 613. In at least one embodiment, the communication link 613 can be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 612 form a parallel or vector processing system focused on computing, which can include a large number of processing cores and / or processing clusters, such as many integrated core processors. In at least one embodiment, the one or more parallel processors 612 form a graphics processing subsystem that can output pixels to one of one or more display devices 610A coupled via an I / O hub 607. In at least one embodiment, the one or more parallel processors 612 can also include a display controller and display interface (not shown) to allow for a direct connection to one or more display devices 610B.

[0083] In at least one embodiment, the system storage unit 614 can be connected to the I / O hub 607 to provide a storage mechanism for the computing system 600. In at least one embodiment, an I / O switch 616 can be used to provide an interface mechanism that allows for connections between the I / O hub 607 and other components, such as a network adapter 618 and / or a wireless network adapter 619 that can be integrated into the platform, as well as various other devices that can be added via one or more external devices 620. In at least one embodiment, the network adapter 618 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 619 can include one or more of Wi-Fi, Bluetooth, NFC devices, or other network devices that include one or more wireless radios.

[0084] In at least one embodiment, the computing system 600 can include other components not explicitly shown that can also be connected to the I / O hub 607, including USB or other port connections, optical storage drives, video capture devices, and so on. In at least one embodiment, the communication paths interconnecting the Figure 6 various different components can be implemented using any suitable protocol, such as a PCI-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols, such as an NVLink high-speed interconnect or interconnect protocol.

[0085] In at least one embodiment, one or more parallel processors 612 incorporate circuitry optimized for graphics and video processing, including, for example, video output circuitry, and form a graphics processing unit (“GPU”). In at least one embodiment, one or more parallel processors 612 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, components of the computing system 600 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 612, memory hub 605, processor 602, and I / O hub 607 may be integrated into a System-on-Chip (SoC) integrated circuit. In at least one embodiment, components of the computing system 600 may be integrated into a single package to form a System-in-Package (“SIP”) configuration. In at least one embodiment, at least a portion of the components of the computing system 600 may be integrated into a Multi-Chip Module (“MCM”) that may be interconnected with other multi-chip modules to form a modular computing system. In at least one embodiment, the I / O subsystem 611 and display device 610B are omitted from the computing system 600.

[0086] Processing system

[0087] The following figures non-limitingly illustrate exemplary processing systems that may be used to implement at least one embodiment.

[0088] Figure 7 FIG. illustrates an Accelerated Processing Unit (“APU”) 700 in accordance with at least one embodiment. In at least one embodiment, the APU 700 is developed by Advanced Micro Devices, Inc., Santa Clara, Calif. In at least one embodiment, the APU 700 may be configured to execute applications such as CUDA programs. In at least one embodiment, the APU 700 includes, but is not limited to, a Core Complex 710, a Graphics Complex 740, a Fabric 760, an I / O Interface 770, a Memory Controller 780, a Display Controller 792, and a Multimedia Engine 794. In at least one embodiment, the APU 700 may include, but is not limited to, any combination of any number of Core Complexes 710, any number of Graphics Complexes 750, any number of Display Controllers 792, and any number of Multimedia Engines 794. For purposes of explanation, where needed, multiple instances of like objects are labeled herein with the reference number identifying the object and a bracketed number identifying the instance.

[0089] In at least one embodiment, the core complex 710 is a CPU, the graphics complex 740 is a GPU, and the APU 700 is a processing unit that integrates 710 and 740 onto a single chip, without limitation. In at least one embodiment, some tasks can be assigned to the core complex 710, and other tasks can be assigned to the graphics complex 740. In at least one embodiment, the core complex 710 is configured to execute the master software associated with the APU 700, such as an operating system. In at least one embodiment, the core complex 710 is the main processor of the APU 700, controlling and coordinating the operations of other processors. In at least one embodiment, the core complex 710 issues commands that control the operations of the graphics complex 740. In at least one embodiment, the core complex 710 can be configured to execute host-executable code exported from CUDA source code, and the graphics complex 740 can be configured to execute device-executable code exported from CUDA source code.

[0090] In at least one embodiment, the core complex 710 includes, but is not limited to, cores 720(1)-720(4) and the L3 cache 730. In at least one embodiment, the core complex 710 can include any combination of any number of cores 720, as well as any number and type of caches, without limitation. In at least one embodiment, the cores 720 are configured to execute instructions of a specific instruction set architecture (“I20”). In at least one embodiment, each core 720 is a CPU core.

[0091] In at least one embodiment, each core 720 includes, but is not limited to, an instruction fetch / decoder unit 722, an integer execution engine 724, a floating-point execution engine 726, and an L2 cache 728. In at least one embodiment, the instruction fetch / decoder unit 722 fetches instructions, decodes such instructions, generates micro-operations, and dispatches separate micro-instructions to the integer execution engine 724 and the floating-point execution engine 726. In at least one embodiment, the instruction fetch / decoder unit 722 can concurrently dispatch one micro-instruction to the integer execution engine 724 and another micro-instruction to the floating-point execution engine 726. In at least one embodiment, the integer execution engine 724 performs integer and memory operations, without limitation. In at least one embodiment, the floating-point execution engine 726 performs floating-point and vector operations, without limitation. In at least one embodiment, the instruction fetch / decoder unit 722 dispatches micro-instructions to a single execution engine that replaces both the integer execution engine 724 and the floating-point execution engine 726.

[0092] In at least one embodiment, each core 720(i), where i is an integer representing a particular instance of core 720, may access the L2 cache 728(i) included in core 720(i). In at least one embodiment, each core 720 included in a core complex 710(j), where j is an integer representing a particular instance of core complex 710, is connected to other cores 720 included in core complex 710(j) via the L3 cache 730(j) included in core complex 710(j). In at least one embodiment, the cores 720 included in a core complex 710(j), where j is an integer representing a particular instance of core complex 710, may access all of the L3 caches 730(j) included in core complex 710(j). In at least one embodiment, the L3 cache 730 may include, but is not limited to, any number of slices.

[0093] In at least one embodiment, the graphics complex 740 may be configured to perform computational operations in a highly parallel manner. In at least one embodiment, the graphics complex 740 is configured to perform graphics pipeline operations such as draw commands, pixel operations, geometric calculations, and other operations associated with rendering an image to a display. In at least one embodiment, the graphics complex 740 is configured to perform operations that are not related to graphics. In at least one embodiment, the graphics complex 740 is configured to perform both operations related to graphics and operations not related to graphics.

[0094] In at least one embodiment, the graphics complex 740 includes, but is not limited to, any number of compute units 750 and an L2 cache 742. In at least one embodiment, the compute units 750 share the L2 cache 742. In at least one embodiment, the L2 cache 742 is partitioned. In at least one embodiment, the graphics complex 740 includes, but is not limited to, any number of compute units 750 and any number (including zero) and type of caches. In at least one embodiment, the graphics complex 740 includes, but is not limited to, any number of dedicated graphics hardware.

[0095] In at least one embodiment, each computing unit 750 includes, but is not limited to, any number of SIMD units 752 and shared memory 754. In at least one embodiment, each SIMD unit 752 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, each computing unit 750 can execute any number of thread blocks, but each thread block is executed on a single computing unit 750. In at least one embodiment, a thread block includes, but is not limited to, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 752 executes different warps. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predication can be used to disable one or more threads in a warp. In at least one embodiment, a lane is a type of thread. In at least one embodiment, a work item is a type of thread. In at least one embodiment, a wavefront is a type of warp. In at least one embodiment, different wavefronts in a thread block can synchronize together and communicate via the shared memory 754.

[0096] In at least one embodiment, fabric 760 is a system interconnect that facilitates data and control transfer across the core complex 710, graphics complex 740, I / O interface 770, memory controller 780, display controller 792, and multimedia engine 794. In at least one embodiment, in addition to or instead of fabric 760, APU 700 can include, but is not limited to, any number and type of system interconnects that facilitate data and control transfer across any number and type of directly or indirectly linked components that can be internal or external to APU 700. In at least one embodiment, I / O interface 770 represents any number and type of I / O interfaces (e.g., PCI, PCI Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various different types of peripheral devices are coupled to I / O interface 770. In at least one embodiment, peripheral devices coupled to I / O interface 770 can include, but are not limited to, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0097] In at least one embodiment, the display controller AMD92 displays images on one or more display devices such as a liquid crystal display (“LCD”) device. In at least one embodiment, the multimedia engine 240 includes, but is not limited to, any number and type of circuitry related to multimedia, such as a video decoder, a video encoder, an image signal processor, and the like. In at least one embodiment, the memory controller 780 facilitates data transfer between the APU 700 and the unified system memory 790. In at least one embodiment, the core complex 710 and the graphics complex 740 share the unified system memory 790.

[0098] In at least one embodiment, the APU 700 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 780 and memory devices (such as shared memory 754) that can be dedicated to one component or shared among multiple components. In at least one embodiment, the APU 700 implements a cache subsystem that includes, but is not limited to, one or more cache memories (such as L2 cache 828, L3 cache 730, and L2 cache 742), each of which can be dedicated to any number of components (such as cores 720, core complex 710, SIMD units 752, compute units 750, and graphics complex 740) or shared among these components.

[0099] Figure 8 The figure shows a CPU 800 according to at least one embodiment. In at least one embodiment, the CPU 800 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, the CPU 800 can be configured to execute application programs. In at least one embodiment, the CPU 800 is configured to execute master software such as an operating system. In at least one embodiment, the CPU 800 issues commands to control the operation of an external GPU (not shown). In at least one embodiment, the CPU 800 can be configured to execute host-executable code derived from CUDA source code, and the external GPU can be configured to execute device-executable code derived from such CUDA source code. In at least one embodiment, the CPU 800 includes, but is not limited to, any number of core complexes 810, fabric 860, I / O interfaces 870, and memory controllers 880.

[0100] In at least one embodiment, the core complex 810 includes, but is not limited to, cores 820(1)-820(4) and an L3 cache 830. In at least one embodiment, the core complex 810 may include any number of cores 820 and any number and type of caches in any combination. In at least one embodiment, the cores 820 are configured to execute instructions of a particular I20. In at least one embodiment, each core 820 is a CPU core.

[0101] In at least one embodiment, each core 820 includes, but is not limited to, a fetch / decode unit 822, an integer execution engine 824, a floating-point execution engine 826, and an L2 cache 828. In at least one embodiment, the fetch / decode unit 822 fetches instructions, decodes such instructions, generates micro-operations, and dispatches individual micro-instructions to the integer execution engine 824 and the floating-point execution engine 826. In at least one embodiment, the fetch / decode unit 822 may concurrently dispatch one micro-instruction to the integer execution engine 824 and another micro-instruction to the floating-point execution engine 826. In at least one embodiment, the integer execution engine 824 performs integer and memory operations without limitation. In at least one embodiment, the floating-point engine 826 performs floating-point and vector operations without limitation. In at least one embodiment, the fetch-decode unit 822 dispatches micro-instructions to a single execution engine in place of both the integer execution engine 824 and the floating-point execution engine 826.

[0102] In at least one embodiment, each core 820(i), where i is an integer representing a particular instance of the core 820, may access the L2 cache 828(i) included in the core 820(i). In at least one embodiment, each core 820 included in the core complex 810(j), where j is an integer representing a particular instance of the core complex 810, is connected to other cores 820 in the core complex 810(j) via the L3 cache 830(j) included in the core complex 810(j). In at least one embodiment, the cores 820 included in the core complex 810(j), where j is an integer representing a particular instance of the core complex 810, may access all of the L3 caches 830(j) included in the core complex 810(j). In at least one embodiment, the L3 cache 830 may include, but is not limited to, any number of slices.

[0103] In at least one embodiment, fabric 860 is a system interconnect that facilitates data and control transfers across core complexes 810(1)-810(N) (where N is an integer greater than 0), I / O interface 870, and memory controller 880. In at least one embodiment, in addition to or instead of fabric 860, CPU 800 may include, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components, which may be internal or external to CPU 800. In at least one embodiment, I / O interface 870 represents any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various different types of peripheral devices are coupled to I / O interface 870. In at least one embodiment, peripheral devices coupled to I / O interface 870 may include, but are not limited to, a display, keyboard, mouse, printer, scanner, joystick or other type of game controller, media recording device, external storage device, network interface card, etc.

[0104] In at least one embodiment, memory controller 880 facilitates data transfer between CPU 800 and system memory 890. In at least one embodiment, core complex 810 and graphics complex 840 share system memory 890. In at least one embodiment, CPU 800 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 880 and memory devices that may be dedicated to one component or shared among multiple components. In at least one embodiment, CPU 800 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 828 and L3 cache 830), each of which may be dedicated to any number of components (e.g., cores 820 and core complex 810) or shared among these components.

[0105] Figure 9The figure shows an exemplary accelerator integrated slice 990 in accordance with at least one embodiment. As used herein, a "slice" includes a designated portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, context management, and interrupt management services on behalf of a plurality of graphics processing engines included in a graphics acceleration module. Each of these graphics processing engines may include a separate GPU. Alternatively, these graphics processing engines may include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module may be a GPU having a plurality of graphics processing engines. In at least one embodiment, the graphics processing engines may be separate GPUs integrated into a common package, line card, or chip.

[0106] The application valid address space 982 within the system memory 914 stores process elements 983. In one embodiment, the process elements 983 are stored in response to a GPU call 981 from an application 980 executing on the processor 907. The process elements 983 contain the process state of the corresponding application 980. The work descriptor ("WD") 984 contained in the process element 983 may be a single job requested by the application, or may contain a pointer to a job queue. In at least one embodiment, the WD 984 is a pointer to a job request queue in the application valid address space 982.

[0107] The graphics acceleration module 946 and / or the separate graphics processing engines may be shared by all processes or a subset of processes in the system. In at least one embodiment, an infrastructure may be included for establishing process state in a virtualized environment and sending the WD 984 to the graphics acceleration module 946 to initiate a job.

[0108] In at least one embodiment, the dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 946 or a separate graphics processing engine. Since the graphics acceleration module 946 is owned by a single process, when the graphics acceleration module 946 is allocated, the hypervisor initializes the accelerator integrated circuit for the owned partition, and the operating system initializes the accelerator integrated circuit for the owned process.

[0109] In operation, the WD fetch unit 991 in the accelerator integration slice 990 fetches the next WD 984, which includes an indication of work to be completed by one or more graphics processing engines of the graphics acceleration module 946. As shown, data from the WD 984 can be stored in the register 945 and used by the memory management unit (“MMU”) 939, the interrupt management circuit 947, and / or the context management circuit 948. For example, one embodiment of the MMU 939 includes a segmentation / page walk circuit system for accessing the segmentation / page table 986 within the OS virtual address space 985. The interrupt management circuit 947 can process interrupt events (“INT”) 992 received from the graphics acceleration module 946. When performing a graphics operation, the effective address 993 generated by the graphics processing engine is converted to a real address by the MMU 939.

[0110] In one embodiment, the same set of registers 945 is replicated for each graphics processing engine and / or graphics acceleration module 946 and can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integration slice 990. Exemplary registers that can be initialized by the hypervisor are shown in Table 1.

[0111] Table 1 Hypervisor-Initialized Registers

[0112] 1 Fragment Control Register 2 Real Address (RA) Scheduler Process Area Pointer 3 Privilege Mask Override Register 4 Interrupt Vector Table Entry Offset 5 Interrupt Vector Table Entry Limit 6 Status Register 7 Logical Partition ID 8 Real Address (RA) Supervisor Accelerator Utilization Record Pointer 9 Memory Descriptor Register

[0113] Exemplary registers that can be initialized by the operating system are shown in Table 2.

[0114] Table 2 Operating-System-Initialized Registers

[0115] 1 Process and Thread Identification 2 Effective Address (EA) Context Save / Restore Pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) Memory Segment Table Pointer 5 Privilege Mask 6 Work Descriptor

[0116] In one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or a particular graphics processing engine. It contains all the information required for the graphics processing engine to complete the work, or it can be a pointer to a memory location where the application has established a command queue of work to be done.

[0117] Figures 10A - 10B The figure shows an exemplary graphics processor in accordance with at least one embodiment. In at least one embodiment, any of the exemplary graphics processors can be fabricated using one or more IP cores. Other logic and circuitry can also be included in addition to what is shown. In at least one embodiment, additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores are included. In at least one embodiment, the exemplary graphics processor is used in a SoC.

[0118] Figure 10AThe figure shows an exemplary graphics processor 1010 of a SoC integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. Figure 10B The figure shows an additional exemplary graphics processor 1040 of a SoC integrated circuit that can be fabricated using one or more IP cores. In at least one embodiment, Figure 10A the graphics processor 1010 is a low-power graphics processor core. In at least one embodiment, Figure 10B the graphics processor 1040 is a higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 1010, 1040 can be Figure 5 a variant of the graphics processor 510.

[0119] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A - 1015N (e.g., 1015A, 1015B, 1015C, 1015D, up to 1015N - 1 and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic such that the vertex processor 1005 is optimized to perform operations for a vertex shader program, while one or more fragment processors 1015A - 1015N perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. In at least one embodiment, the vertex processor 1005 implements the vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1015A - 1015N use the primitives and vertex data generated by the vertex processor 1005 to produce a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1015A - 1015N are optimized to execute a fragment shader program as provided in the OpenGL API, which can be used to implement operations similar to a pixel shader program as provided in the Direct 3D API.

[0120] In at least one embodiment, the graphics processor 1010 additionally includes one or more MMUs 1020A - 1020B, caches 1025A - 1025B, and circuit interconnects 1030A - 1030B. In at least one embodiment, one or more MMUs 1020A - 1020B provide virtual - physical address mapping for the graphics processor 1010, including for the vertex processor 1005 and / or the fragment processors 1015A - 1015N, which may reference vertex or image / texture data stored in memory in addition to referencing vertex or image / texture data stored in one or more of the caches 1025A - 1025B. In at least one embodiment, one or more MMUs 1020A - 1020B may be synchronized with other MMUs within the system, the other MMUs including one or more MMUs associated with Figure 5 one or more application processors 505, image processors 515, and / or video processors 520 such that each of the processors 505 - 520 may participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1030A - 1030B enable the graphics processor 1010 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0121] In at least one embodiment, the graphics processor 1040 includes Figure 10A one or more MMUs 1020A - 1020B, caches 1025A - 1025B, and circuit interconnects 1030A - 1030B of the graphics processor 1010. In at least one embodiment, the graphics processor 1040 includes one or more shader cores 1055A - 1055N (e.g., 1055A, 1055B, 1055C, 1055D, 1055E, 1055F, up to 1055N - 1 and 1055N), which provide a unified shader core architecture in which a single core or type of core can execute all types of programmable shader code, including shader program code implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores may vary. In at least one embodiment, the graphics processor 1040 includes an inter - core task manager 1045, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1055A - 1055N and a tiling unit 1058 that accelerates tiling operations for tile - based rendering, in which the rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence in the scene or to optimize the use of internal caches.

[0122] Figure 11AThe figure shows a graphics core 1100 according to at least one embodiment. In at least one embodiment, the graphics core 1100 may be included in the Figure 5 graphics processor 510. In at least one embodiment, the graphics core 1100 may be a unified shader core 1055A - 1055N as in Figure 10B . In at least one embodiment, the graphics core 1100 includes a shared instruction cache 1102, texture units 1118, and cache / shared memory 1120 for shared execution resources within the graphics core 1100. In at least one embodiment, the graphics core 1100 may include multiple shards 1101A - 1101N or partitions for each core, and the graphics processor may include multiple instances of the graphics core 1100. The shards 1101A - 1101N may include support logic including local instruction caches 1104A - 1104N, thread schedulers 1106A - 1106N, thread dispatchers 1108A - 1108N, and register sets 1110A - 1110N. In at least one embodiment, the shards 1101A - 1101N may include a set of additional functional units (“AFU”) 1112A - 1112N, floating - point units (“FPU”) 1114A - 1114N, integer arithmetic logic units (“ALU”) 1116 - 1116N, address calculation units (“ACU”) 1113A - 1113N, double - precision floating - point units (“DPFPU”) 1115A - 1115N, and matrix processing units (“MPU”) 1117A - 1117N.

[0123] In at least one embodiment, the FPUs 1114A - 1114N may perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, while the DPFPU 1115A - 1115N performs double - precision (64 - bit) floating - point operations. In at least one embodiment, the ALUs 1116A - 1116N may perform variable - precision integer operations at 8 - bit, 16 - bit, and 32 - bit precisions and may be configured for mixed - precision operations. In at least one embodiment, the MPUs 1117A - 1117N may also be configured for mixed - precision matrix operations including half - precision floating - point and 8 - bit integer operations. In at least one embodiment, the MPUs 1117A - 1117N may perform a variety of matrix operations to accelerate CUDA programs, including enabling support for accelerated general matrix - matrix multiplication (“GEMM”). In at least one embodiment, the AFUs 1112A - 1112N may perform additional logical operations not supported by the floating - point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0124] Figure 11BThe figure shows a general purpose graphics processing unit (“GPGPU”) 1130 in accordance with at least one embodiment. In at least one embodiment, the GPGPU 1130 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, the GPGPU 1130 can be configured such that highly parallel computing operations can be implemented by a GPU array. In at least one embodiment, the GPGPU 1130 can be directly linked to other instances of the GPGPU 1130 to create a multi-GPU cluster that improves the execution time of CUDA programs. In at least one embodiment, the GPGPU 1130 includes a host interface 1132 that allows for connection to a host processor. In at least one embodiment, the host interface 1132 is a PCIe interface. In at least one embodiment, the host interface 1132 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the GPGPU 1130 receives commands from the host processor and distributes execution threads associated with those commands to a set of compute clusters 1136A - 1136H using a global scheduler 1134. In at least one embodiment, the compute clusters 1136A - 1136H share a cache memory 1138. In at least one embodiment, the cache memory 1138 can act as a higher-level cache for the cache memories within the compute clusters 1136A - 1136H.

[0125] In at least one embodiment, the GPGPU 1130 includes memories 1144A - 1144B coupled to the compute clusters 1136A - 1136H via a set of memory controllers 1142A - 1142B. In at least one embodiment, the memories 1144A - 1144B can include various different types of memory devices, including DRAM or graphics random access memory, such as synchronous graphics random access memory (“SGRAM”), including graphics double data rate (“GDDR”) memory.

[0126] In at least one embodiment, each of the compute clusters 1136A - 1136H includes a set of graphics cores, such as Figure 11A graphics core 1100, which can include multiple types of integer and floating point logic units that can implement computing operations at a range of precisions, including those suitable for computations associated with CUDA programs. For example, in at least one embodiment, at least a subset of the floating point units in each of the compute clusters 1136A - 1136H can be configured to implement 16-bit or 32-bit floating point operations, while different subsets of floating point units can be configured to implement 64-bit floating point operations.

[0127] In at least one embodiment, multiple instances of GPGPU 1130 may be configured to operate as a compute cluster. Compute clusters 1136A - 1136H may implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of GPGPU 1130 communicate via host interface 1132. In at least one embodiment, GPGPU 1130 includes an I / O hub 1139 that couples GPGPU 1130 to a GPU link 1140 that allows for a direct connection to other instances of GPGPU 1130. In at least one embodiment, GPU link 1140 is coupled to a dedicated GPU - GPU bridge that allows for communication and synchronization between multiple instances of GPGPU 1130. In at least one embodiment, GPU link 1140 is coupled to a high - speed interconnect to transfer data to and receive data from other GPGPU 1130s or parallel processors. In at least one embodiment, multiple instances of GPGPU 1130 are located in separate data processing systems and communicate via a network device that is accessible via host interface 1132. In at least one embodiment, in addition to or in place of host interface 1132, GPU link 1140 may be configured to allow for a connection to a host processor. In at least one embodiment, GPGPU 1130 may be configured to execute CUDA programs.

[0128] Figure 12A The figure illustrates a parallel processor 1200 in accordance with at least one embodiment. In at least one embodiment, the various different components of parallel processor 1200 may be implemented using one or more integrated circuit devices such as programmable processors, application - specific integrated circuits (“ASICs”), or FPGAs.

[0129] In at least one embodiment, parallel processor 1200 includes a parallel processing unit 1202. In at least one embodiment, parallel processing unit 1202 includes an I / O unit 1204 that allows for communication with other devices including other instances of parallel processing unit 1202. In at least one embodiment, I / O unit 1204 may be directly connected to other devices. In at least one embodiment, I / O unit 1204 is connected to other devices such as memory hub 605 via the use of a hub or switch interface. In at least one embodiment, the connection between memory hub 605 and I / O unit 1204 forms a communication link. In at least one embodiment, I / O unit 1204 is connected to host interface 1206 and memory crossbar 1216, where host interface 1206 receives commands for performing processing operations and memory crossbar 1216 receives commands for performing memory operations.

[0130] In at least one embodiment, when host interface 1206 receives a command buffer via I / O unit 1204, host interface 1206 may direct the work operations to implement those commands to front end 1208. In at least one embodiment, front end 1208 is coupled to scheduler 1210, which is configured to distribute commands or other work items to processing array 1212. In at least one embodiment, scheduler 1210 ensures that processing array 1212 is properly configured and in an active state before tasks are distributed to processing array 1212. In at least one embodiment, scheduler 1210 is implemented via firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1210 can be configured to implement complex scheduling and work distribution operations at both coarse-grained and fine-grained levels, allowing for fast preemption and context switching of threads executing on processing array 1212. In at least one embodiment, host software can attest to the workload scheduled on processing array 1212 via one of a plurality of graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across processing array 1212 by scheduler 1210 logic within the microcontroller including scheduler 1210.

[0131] In at least one embodiment, processing array 1212 may include up to "N" clusters (e.g., cluster 1214A, cluster 1214B, up to cluster 1214N). In at least one embodiment, each cluster 1214A - 1214N of processing array 1212 may execute a large number of concurrent threads. In at least one embodiment, scheduler 1210 may use a variety of different scheduling and / or work distribution algorithms to allocate work to clusters 1214A - 1214N of processing array 1212, and the algorithms may vary according to the workload that occurs for each type of program or computation. In at least one embodiment, scheduling may be dynamically handled by scheduler 1210, or may be partially assisted by compiler logic during the compilation of program logic configured to be executed by processing array 1212. In at least one embodiment, different clusters 1214A - 1214N of processing array 1212 may be assigned to process different types of programs or to implement different types of computations.

[0132] In at least one embodiment, processing array 1212 may be configured to implement a variety of different types of parallel processing operations. In at least one embodiment, processing array 1212 is configured to implement general-purpose parallel computing operations. For example, in at least one embodiment, processing array 1212 may include logic for performing processing tasks including filtering video and / or audio data, implementing modeling operations including physical operations, and implementing data transformations.

[0133] In at least one embodiment, the processing array 1212 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing array 1212 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing array 1212 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1202 may transfer data from the system memory via the I / O unit 1204 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (such as the parallel processor memory 1222) during processing and then written back to the system memory.

[0134] In at least one embodiment, when the parallel processing unit 1202 is used to perform graphics processing, the scheduler 1210 may be configured to divide the processing workload into tasks of approximately equal size to better allow the distribution of graphics processing operations to the multiple clusters 1214A - 1214N of the processing array 1212. In at least one embodiment, portions of the processing array 1212 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations to generate a reproduced image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 1214A - 1214N may be stored in a buffer to allow the intermediate data to be transferred between the clusters 1214A - 1214N for further processing.

[0135] In at least one embodiment, the processing array 1212 may receive processing tasks to be executed via the scheduler 1210, which receives commands defining the processing tasks from the front end 1208. In at least one embodiment, the processing tasks may include indices of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as status parameters and commands defining how the data will be processed (such as what program to execute). In at least one embodiment, the scheduler 1210 may be configured to obtain the index corresponding to the task or may receive the index from the front end 1208. In at least one embodiment, the front end 1208 may be configured to ensure that the processing array 1212 is configured in a valid state before the workload specified in the incoming command buffer (such as a batch buffer, push buffer, etc.) is launched.

[0136] In at least one embodiment, each of one or more instances of parallel processing unit 1202 may be coupled to parallel processor memory 1222. In at least one embodiment, parallel processor memory 1222 may be accessed via memory crossbar 1216, which may receive memory requests from processing array 1212 as well as I / O unit 1204. In at least one embodiment, memory crossbar 1216 may access parallel processor memory 1222 via memory interface 1218. In at least one embodiment, memory interface 1218 may include a plurality of partitioning units (e.g., partitioning unit 1220A, partitioning unit 1220B, up to partitioning unit 1220N), each of which may be coupled to a portion (e.g., memory unit) of parallel processor memory 1222. In at least one embodiment, the number of partitioning units 1220A - 1220N is configured to be equal to the number of memory units such that first partitioning unit 1220A has a corresponding first memory unit 1224A, second partitioning unit 1220B has a corresponding memory unit 1224B, and Nth partitioning unit 1220N has a corresponding Nth memory unit 1224N. In at least one embodiment, the number of partitioning units 1220A - 1220N may not be equal to the number of memory devices.

[0137] In at least one embodiment, memory units 1224A - 1224N may include various different types of memory devices, including DRAM or graphics random access memory such as SGRAM, including GDDR memory. In at least one embodiment, memory units 1224A - 1224N may also include 3D stacked memory, including but not limited to high bandwidth memory (“HBM”). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory units 1224A - 1224N, allowing partitioning units 1220A - 1220N to write portions of each rendering target in parallel to efficiently utilize the available bandwidth of parallel processor memory 1222. In at least one embodiment, local instances of parallel processor memory 1222 may be excluded in favor of a unified memory design that utilizes system memory in conjunction with local cache memory.

[0138] In at least one embodiment, any one of clusters 1214A - 1214N of processing array 1212 can process data to be written to any memory cell 1224A - 1224N within parallel processor memory 1222. In at least one embodiment, memory crossbar 1216 can be configured to transfer the output of each cluster 1214A - 1214N to any partition unit 1220A - 1220N or to another cluster 1214A - 1214N, which can perform additional processing operations on the output. In at least one embodiment, each cluster 1214A - 1214N can communicate with memory interface 1218 through memory crossbar 1216 to read from or write to various different external memory devices. In at least one embodiment, memory crossbar 1216 has a connection to memory interface 1218 to communicate with I / O unit 1204 and a connection to a local instance of parallel processor memory 1222, such that processing units within different clusters 1214A - 1214N can communicate with system memory or other memory not local to parallel processing unit 1202. In at least one embodiment, memory crossbar 1216 can use virtual channels to separate traffic flows between clusters 1214A - 1214N and partition units 1220A - 1220N.

[0139] In at least one embodiment, multiple instances of parallel processing unit 1202 can be provided on a single external card, or multiple external cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1202 can be configured to interoperate even if the different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 1202 can include floating - point units with higher precision relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 1202 or parallel processor 1200 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0140] Figure 12BThe figure shows a processing cluster 1294 in accordance with at least one embodiment. In at least one embodiment, the processing cluster 1294 is included within a parallel processing unit. In at least one embodiment, the processing cluster 1294 is one of the processing clusters 1214A - 1214N of FIG. 12. In at least one embodiment, the processing cluster 1294 can be configured to execute a number of threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction multiple data ("SIMD") instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread ("SIMT") techniques are used to support the parallel execution of a large number of generally synchronized threads using a common instruction unit, which is configured to issue instructions to a set of processing engines within each processing cluster 1294.

[0141] In at least one embodiment, the operation of the processing cluster 1294 can be controlled via a pipeline manager 1232 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1232 receives instructions from the scheduler 1210 of FIG. 12 and manages the execution of those instructions via the graphics multiprocessor 1234 and / or the texture unit 1236. In at least one embodiment, the graphics multiprocessor 1234 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various different types of SIMT parallel processors of different architectures can be included within the processing cluster 1294. In at least one embodiment, one or more instances of the graphics multiprocessor 1234 can be included within the processing cluster 1294. In at least one embodiment, the graphics multiprocessor 1234 can process data, and a data crossbar 1240 can be used to distribute the processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, the pipeline manager 1232 can facilitate the distribution of the processed data by specifying the destination of the processed data to be distributed via the data crossbar 1240.

[0142] In at least one embodiment, each graphics multiprocessor 1234 within the processing cluster 1294 can include the same set of functional execution logic (e.g., arithmetic logic units, load / store units ("LSU"), etc.). In at least one embodiment, the functional execution logic can be configured in a pipeline manner where new instructions can be issued before previous instructions are completed. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, shifts, and the calculation of various different algebraic functions. In at least one embodiment, different operations can be implemented using the same functional unit hardware, and any combination of functional units can exist.

[0143] In at least one embodiment, the instructions transmitted to processing cluster 1294 constitute threads. In at least one embodiment, a set of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 1234. In at least one embodiment, the thread group can include fewer threads than the number of processing engines within graphics multiprocessor 1234. In at least one embodiment, when the thread group includes fewer threads than the number of processing engines, one or more of these processing engines may be idle during the cycles in which the thread group is being processed. In at least one embodiment, the thread group can also include more threads than the number of processing engines within graphics multiprocessor 1234. In at least one embodiment, when the thread group includes more threads than the number of processing engines within graphics multiprocessor 1234, the processing can be implemented in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed concurrently on graphics multiprocessor 1234.

[0144] In at least one embodiment, graphics multiprocessor 1234 includes an internal cache memory that implements load and store operations. In at least one embodiment, graphics multiprocessor 1234 can discard the internal cache and use the cache memory (e.g., L1 cache 1248) within processing cluster 1294. In at least one embodiment, each graphics multiprocessor 1234 also has access to a partition unit (e.g., Figure 12A partition units 1220A - 1220N) within a level 2 (“L2”) cache that is shared among all processing clusters 1294 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1234 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 1202 can be used as global memory. In at least one embodiment, processing cluster 1294 includes multiple instances of graphics multiprocessor 1234 that can share common instructions and data that may be stored in L1 cache 1248.

[0145] In at least one embodiment, each processing cluster 1294 may include a memory management unit (MMU) 1245 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1245 may reside within the memory interface 1218 of FIG. 12. In at least one embodiment, the MMU 1245 includes a set of page table entries ("PTEs") that map virtual addresses to the physical addresses of tiles and optionally cache line indices. In at least one embodiment, the MMU 1245 may include a translation lookaside buffer ("TLB") or cache that may reside within the graphics multiprocessor 1234, or the L1 cache 1248, or the processing cluster 1294. In at least one embodiment, the physical addresses are processed to distribute surface data access locations to allow for efficient request interleaving among partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0146] In at least one embodiment, the processing cluster 1294 may be configured such that each graphics multiprocessor 1234 is coupled to a texture unit 1236 to perform texture mapping operations, e.g., determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multiprocessor 1234, and, if needed, from the L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 1234 outputs the processed task to the data crossbar 1240 to provide the processed task to another processing cluster 1294 for further processing or to store the processed task in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 1216. In at least one embodiment, the pre-raster operation unit ("preROP") 1242 is configured to receive data from the graphics multiprocessor 1234 and direct the data to the ROP units, which may be located within partition units as described herein (e.g., the partition units 1220A - 1220N of FIG. 12). In at least one embodiment, the preROP 1242 may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0147] Figure 12C FIG. illustrates a graphics multiprocessor 1296 in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1296 is Figure 12BThe graphics multiprocessor 1234. In at least one embodiment, the graphics multiprocessor 1296 is coupled to the pipeline manager 1232 of the processing cluster 1294. In at least one embodiment, the graphics multiprocessor 1296 has an execution pipeline, including but not limited to an instruction cache 1252, an instruction unit 1254, an address mapping unit 1256, a register file 1258, one or more GPGPU cores 1262, and one or more LSUs 1266. The GPGPU cores 1262 and the LSUs 1266 are coupled to the cache memory 1272 and the shared memory 1270 via a memory and cache interconnect 1268.

[0148] In at least one embodiment, the instruction cache 1252 receives a stream of instructions to be executed from the pipeline manager 1232. In at least one embodiment, the instructions are cached in the instruction cache 1252 and are dispatched by the instruction unit 1254 for execution. In at least one embodiment, the instruction unit 1254 can dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within the GPGPU core 1262. In at least one embodiment, instructions can access any one of the local, shared, or global address spaces by specifying an address within a unified address space. In at least one embodiment, the address mapping unit 1256 can be used to translate an address within the unified address space into a different memory address that can be accessed by the LSU 1266.

[0149] In at least one embodiment, the register file 1258 provides a register set for the functional units of the graphics multiprocessor 1296. In at least one embodiment, the register file 1258 provides temporary storage for the operands of the data paths of the functional units (e.g., GPGPU cores 1262, LSUs 1266) connected to the graphics multiprocessor 1296. In at least one embodiment, the register file 1258 is partitioned among each functional unit such that each functional unit is assigned a dedicated portion of the register file 1258. In at least one embodiment, the register file 1258 is partitioned among different thread groups being executed by the graphics multiprocessor 1296.

[0150] In at least one embodiment, each of the GPGPU cores 1262 may include an FPU and / or an integer ALU for executing instructions of the graphics multiprocessor 1296. The GPGPU cores 1262 may be architecturally similar or may be architecturally different. In at least one embodiment, a first portion of the GPGPU core 1262 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core 1262 includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or allow for implementation of variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1296 may additionally include one or more fixed-function or special-function units to implement specific functions such as copy rectangle or pixel blend operations. In at least one embodiment, one or more of the GPGPU cores 1262 may also include fixed or special-function logic.

[0151] In at least one embodiment, the GPGPU core 1262 includes SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 1262 may physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core 1262 may be generated by a shader compiler at compile time or may be automatically generated when executing a program written and compiled for a single-program multiple-data (“SPMD”) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model may be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations may be executed in parallel via a single SIMD8 logic unit.

[0152] In at least one embodiment, the memory and cache interconnect 1268 is an interconnect network that connects each functional unit of the graphics multiprocessor 1296 to the register file 1258 and the shared memory 1270. In at least one embodiment, the memory and cache interconnect 1268 is a crossbar interconnect that allows the LSU 1266 to implement load and store operations between the shared memory 1270 and the register file 1258. In at least one embodiment, the register file 1258 can operate at the same frequency as the GPGPU core 1262, so the data transfer latency between the GPGPU core 1262 and the register file 1258 is very low. In at least one embodiment, the shared memory 1270 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1296. In at least one embodiment, the cache memory 1272 can be used as, for example, a data cache to cache texture data transferred between the functional units and the texture unit 1236. In at least one embodiment, the shared memory 1270 can also be used as a program-managed cache. In at least one embodiment, in addition to the automatically cached data stored in the cache memory 1272, threads executing on the GPGPU core 1262 can also programmatically store data into the shared memory.

[0153] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various different general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated into the same package or chip as the core and communicatively coupled to the core via a processor bus / interconnect within the package or chip. In at least one embodiment, regardless of how the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions included in the WD. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0154] Figure 13The figure shows a graphics processor 1300 in accordance with at least one embodiment. In at least one embodiment, the graphics processor 1300 includes a ring interconnect 1302, a pipeline front end 1304, a media engine 1337, and graphics cores 1380A - 1380N. In at least one embodiment, the ring interconnect 1302 couples the graphics processor 1300 to other processing units, including other graphics processors or one or more general - purpose processor cores. In at least one embodiment, the graphics processor 1300 is one of many processors integrated into a multi - core processing system.

[0155] In at least one embodiment, the graphics processor 1300 receives batch commands via the ring interconnect 1302. In at least one embodiment, incoming commands are interpreted by a command streamer 1303 in the pipeline front end 1304. In at least one embodiment, the graphics processor 1300 includes scalable execution logic for 3D geometry processing and media processing implemented via the graphics cores 1380A - 1380N. In at least one embodiment, for 3D geometry processing commands, the command streamer 1303 provides the commands to a geometry pipeline 1336. In at least one embodiment, for at least some media processing commands, the command streamer 1303 provides the commands to a video front end 1334 coupled to the media engine 1337. In at least one embodiment, the media engine 1337 includes a video quality engine (“VQE”) 1330 for video and image post - processing and a multi - format encoding / decoding (“MFX”) engine 1333 for hardware - accelerated media data encoding and decoding. In at least one embodiment, each of the geometry pipeline 1336 and the media engine 1337 generates execution threads for the thread execution resources provided to at least one graphics core 1380A.

[0156] In at least one embodiment, the graphics processor 1300 includes scalable thread execution resources characterized by modular graphics cores 1380A - 1380N (sometimes referred to as core slices), each modular graphics core having a plurality of sub - cores 1350A - 550N, 1360A - 1360N (sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 1300 can have any number of graphics cores 1380A through 1380N. In at least one embodiment, the graphics processor 1300 includes a graphics core 1380A having at least a first sub - core 1350A and a second sub - core 1360A. In at least one embodiment, the graphics processor 1300 is a low - power processor having a single sub - core (e.g., sub - core 1350A). In at least one embodiment, the graphics processor 1300 includes a plurality of graphics cores 1380A - 1380N, each graphics core including a set of first sub - cores 1350A - 1350N and a set of second sub - cores 1360A - 1360N. In at least one embodiment, each of the first sub - cores 1350A - 1350N includes at least a first set of execution units (“EUs”) 1352A - 1352N and media / texture samplers 1354A - 1354N. In at least one embodiment, each of the second sub - cores 1360A - 1360N includes at least a second set of execution units 1362A - 1362N and samplers 1364A - 1364N. In at least one embodiment, each of the sub - cores 1350A - 1350N, 1360A - 1360N shares a set of shared resources 1370A - 1370N. In at least one embodiment, the shared resources 1370 include shared cache memory and pixel operation logic.

[0157] Figure 14 FIG. illustrates a processor 1400 in accordance with at least one embodiment. In at least one embodiment, the processor 1400 can include, but is not limited to, logic circuitry for implementing instructions. In at least one embodiment, the processor 1400 can implement instructions, including x86 instructions, ARM instructions, special instructions for ASICs, and the like. In at least one embodiment, the processor 1400 can include registers for storing packed data, such as the 64 - bit wide MMXTM registers in a microprocessor enabled with MMX technology from Intel Corporation, Santa Clara, California. In at least one embodiment, MMX registers in both integer and floating - point forms can operate on packed data elements accompanied by SIMD and Streaming SIMD Extensions (“SSE”) instructions. In at least one embodiment, 128 - bit wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or later versions (collectively referred to as “SSEx”) technology can hold such packed data operands. In at least one embodiment, the processor 1410 can implement instructions for accelerating CUDA programs.

[0158] In at least one embodiment, the processor 1400 includes an in-order front end (“front end”) 1401 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 1401 may include a number of units. In at least one embodiment, the instruction prefetcher 1426 fetches instructions from memory and feeds the instructions to an instruction decoder 1428, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1428 decodes the received instructions into one or more operations for execution, referred to as “microinstructions” or “micro-operations” (also referred to as “micro op” or “uop”). In at least one embodiment, the instruction decoder 1428 parses the instructions into opcodes and corresponding data and control fields that can be used by the microarchitecture to perform the operations. In at least one embodiment, the trace cache 1430 may assemble the decoded uops into a program-ordered sequence or trace in a uop queue 1434 for execution. In at least one embodiment, when the trace cache 1430 encounters a complex instruction, the microcode ROM 1432 provides the uops required to complete the operation.

[0159] In at least one embodiment, some instructions may be converted into a single micro-op, while other instructions require several micro-ops to complete the entire operation. In at least one embodiment, if more than four micro-ops are required to complete a certain instruction, then the instruction decoder 1428 may access the microcode ROM 1432 to implement the instruction. In at least one embodiment, instructions may be decoded into a small number of micro-ops for processing at the instruction decoder 1428. In at least one embodiment, if a certain number of micro-ops are required to complete an operation, the instruction may be stored within the microcode ROM 1432. In at least one embodiment, the trace cache 1430 references an entry point programmable logic array (“PLA”) to determine the correct microinstruction pointer for reading the microcode sequence for completing one or more instructions from the microcode ROM 1432. In at least one embodiment, after the microcode ROM 1432 completes the micro-op sequencing of the instruction, the front end 1401 of the machine may continue to fetch micro-ops from the trace cache 1430.

[0160] In at least one embodiment, an out-of-order execution engine (“out-of-order engine”) 1403 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has a number of buffers to organize and reorder the instruction stream as it proceeds down the pipeline and is scheduled for execution to optimize performance. The out-of-order execution engine 1403 includes, but is not limited to, an allocator / register renamer 1440, a memory uop queue 1442, an integer / float uop queue 1444, a memory scheduler 1446, a fast scheduler 1402, a slow / general floating-point scheduler (“slow / general FP scheduler”) 1404, and a simple floating-point scheduler (“simple FP scheduler”) 1406. In at least one embodiment, the fast scheduler 1402, the slow / general floating-point scheduler 1404, and the simple floating-point scheduler 1406 are also collectively referred to herein as “uop schedulers 1402, 1404, 1406”. The allocator / register renamer 1440 allocates the machine buffers and resources required for each uop to execute. In at least one embodiment, the allocator / register renamer 1440 renames logical registers to entries in the register file. In at least one embodiment, the allocator / register renamer 1440 also allocates entries for each uop in one of two uop queues, the uop queues being a memory uop queue 1442 for memory operations and an integer / float uop queue 1444 for non-memory operations, prior to the memory scheduler 1446 and the uop schedulers 1402, 1404, 1406. In at least one embodiment, the uop schedulers 1402, 1404, 1406 determine when a uop is ready to execute based on the availability of its dependent input register operands sources being ready and the execution resources required for the uop to complete its operation. In at least one embodiment, the fast scheduler 1402 of at least one embodiment may schedule every half master clock cycle, while the slow / general floating-point scheduler 1404 and the simple floating-point scheduler 1406 may schedule once per master processor clock cycle. In at least one embodiment, the uop schedulers 1402, 1404, 1406 arbitrate dispatch ports to schedule uops for execution.

[0161] In at least one embodiment, execution block b11 includes, but is not limited to, integer register file / bypass network 1408, floating-point register file / bypass network (“FP register file / bypass network”) 1410, address generation units (“AGU”) 1412 and 1414, fast ALUs 1416 and 1418, slow ALU 1420, floating-point ALU (“FP”) 1422, and floating-point move unit (“FP move”) 1424. In at least one embodiment, integer register file / bypass network 1408 and floating-point register file / bypass network 1410 are also referred to herein as “register files 1408, 1410”. In at least one embodiment, AGUs 1412 and 1414, fast ALUs 1416 and 1418, slow ALU 1420, floating-point ALU 1422, and floating-point move unit 1424 are also referred to herein as “execution units 1412, 1414, 1416, 1418, 1420, 1422, and 1424”. In at least one embodiment, an execution block may include, but is not limited to, any combination of any number (including zero) and type of register files, bypass networks, address generation units, and execution units.

[0162] In at least one embodiment, register files 1408, 1410 may be arranged between uop schedulers 1402, 1404, 1406 and execution units 1412, 1414, 1416, 1418, 1420, 1422, and 1424. In at least one embodiment, integer register file / bypass network 1408 performs integer operations. In at least one embodiment, floating-point register file / bypass network 1410 performs floating-point operations. In at least one embodiment, each of register files 1408, 1410 may include, but is not limited to, a bypass network that may bypass a just-completed result that has not yet been written to the register file or forward it to a new dependent uop. In at least one embodiment, register files 1408, 1410 may transfer data to each other. In at least one embodiment, integer register file / bypass network 1408 may include, but is not limited to, two separate register files, one for low-order 32-bit data and another for high-order 32-bit data. In at least one embodiment, floating-point register file / bypass network 1410 may include, but is not limited to, 128-bit wide entries, since floating-point instructions typically have operands with widths ranging from 64 bits to 128 bits.

[0163] In at least one embodiment, execution units 1412, 1414, 1416, 1418, 1420, 1422, and 1424 may execute instructions. In at least one embodiment, register files 1408, 1410 store integer and floating-point data operands required for microinstructions to be executed. In at least one embodiment, processor 1400 may include, but is not limited to, any number and combination of execution units 1412, 1414, 1416, 1418, 1420, 1422, 1424. In at least one embodiment, floating-point ALU 1422 and floating-point move unit 1424 may perform floating-point, MMX, SIMD, AVX, and SSE or other operations. In at least one embodiment, floating-point ALU 1422 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro ops. In at least one embodiment, instructions involving floating-point values may be processed using floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 1416, 1418. In at least one embodiment, fast ALUs 1416, 1418 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, the most complex integer operations are transferred to slow ALU 1420 because slow ALU 1420 may include, but is not limited to, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by AGUs 1412, 1414. In at least one embodiment, fast ALU 1416, fast ALU 1418, and slow ALU 1420 may perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 1416, fast ALU 1418, and slow ALU 1420 may be implemented to support a variety of data bit sizes, including 16, 32, 128, 256, and so on. In at least one embodiment, floating-point ALU 1422 and floating-point move unit 1424 may be implemented to have a series of operands of various widths of bits. In at least one embodiment, floating-point ALU 1422 and floating-point move unit 1424 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0164] In at least one embodiment, the uop schedulers 1402, 1404, 1406 dispatch dependent operations before the completion of the execution of the parent load. In at least one embodiment, since uops can be speculatively scheduled and executed in the processor 1400, the processor 1400 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be in - flight dependent operations in the pipeline that leave temporarily incorrect data for the scheduler. In at least one embodiment, a replay mechanism tracks and re - executes instructions that use incorrect data. In at least one embodiment, dependent operations may need to be replayed, and independent operations may be allowed to complete. In at least one embodiment, the scheduler and the replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0165] In at least one embodiment, the term "register" may refer to an on - board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from the programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuitry. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented through circuitry within the processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, combinations of dedicated and dynamically allocated physical registers, and so on. In at least one embodiment, integer registers store 32 - bit integer data. The register file of at least one embodiment also contains eight multimedia SIMD registers for packed data.

[0166] Figure 15 FIG. illustrates a processor 1500 in accordance with at least one embodiment. In at least one embodiment, the processor 1500 includes, but is not limited to, one or more processor cores ("cores") 1502A - 1502N, an integrated memory controller 1514, and an integrated graphics processor 1508. In at least one embodiment, the processor 1500 may include additional cores up to and including the additional processor core 1502N represented by the dashed box. In at least one embodiment, each of the processor cores 1502A - 1502N includes one or more internal cache units 1504A - 1504N. In at least one embodiment, each processor core also has access to one or more shared cache units 1506.

[0167] In at least one embodiment, the internal cache units 1504A - 1504N and the shared cache unit 1506 represent a cache memory hierarchy within the processor 1500. In at least one embodiment, the cache memory units 1504A - 1504N can include at least one level of instruction and data cache within each processor core and one or more levels of shared mid - level cache, such as L2, L3, level 4 (“L4”), or other levels of cache, where the highest level of cache before the external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence between the various different cache units 1506 and 1504A - 1504N.

[0168] In at least one embodiment, the processor 1500 may also include a set of one or more bus controller units 1516 and a system agent core 1510. In at least one embodiment, the one or more bus controller units 1516 manage a collection of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 1510 provides management functions for the various different processor components. In at least one embodiment, the system agent core 1510 includes one or more integrated memory controllers 1514 that manage access to various different external memory devices (not shown).

[0169] In at least one embodiment, one or more of the processor cores 1502A - 1502N include support for simultaneous multithreading. In at least one embodiment, the system agent core 1510 includes components for coordinating and operating the processor cores 1502A - 1502N during multithreaded processing. In at least one embodiment, the system agent core 1510 may additionally include a power control unit (“PCU”) that includes logic and components for regulating one or more power states of the processor cores 1502A - 1502N and the graphics processor 1508.

[0170] In at least one embodiment, the processor 1500 additionally includes a graphics processor 1508 that performs graphics processing operations. In at least one embodiment, the graphics processor 1508 is coupled to the shared cache unit 1506 and the system agent core 1510 that includes one or more integrated memory controllers 1514. In at least one embodiment, the system agent core 1510 also includes a display controller 1511 that drives the output of the graphics processor to one or more coupled displays. In at least one embodiment, the display controller 1511 can also be a separate module coupled to the graphics processor 1508 via at least one interconnect, or can be integrated within the graphics processor 1508.

[0171] In at least one embodiment, a ring-based interconnect unit 1512 is used to couple the internal components of the processor 1500. In at least one embodiment, a replaceable interconnect unit, such as a point-to-point interconnect line, a switched interconnect line, or other technologies, may be used. In at least one embodiment, the graphics processor 1508 is coupled to the ring interconnect 1512 via an I / O link 1513.

[0172] In at least one embodiment, the I / O link 1513 represents at least one of a variety of I / O interconnect lines, including a package I / O interconnect line that facilitates communication between various different processor components and a high-performance embedded memory module 1518 such as an eDRAM module. In at least one embodiment, each of the processor cores 1502A - 1502N and the graphics processor 1508 uses the embedded memory module 1518 as a shared LLC.

[0173] In at least one embodiment, the processor cores 1502A - 1502N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1502A - 1502N are heterogeneous in terms of I20, where one or more of the processor cores 1502A - 1502N execute a common instruction set, while one or more of the other cores among the processor cores 1502A - 1502N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 1502A - 1502N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled to one or more cores with lower power consumption. In at least one embodiment, the processor 1500 may be implemented on one or more chips or implemented as a SoC integrated circuit.

[0174] Figure 16 The figure shows a graphics processor core 1600 according to at least one of the described embodiments. In at least one embodiment, the graphics processor core 1600 is included in a graphics core array. In at least one embodiment, the graphics processor core 1600, sometimes referred to as a core slice, may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 1600 is an example of a graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 1600 may include a fixed function block 1630 coupled to a plurality of sub-cores 1601A - 1601F, also referred to as sub-slices, which include modular blocks of general and fixed function logic.

[0175] In at least one embodiment, the fixed function block 1630 includes a geometry / fixed function pipeline 1636, which may be shared by all sub-cores in the graphics processor 1600, for example, in a lower performance and / or lower power graphics processor implementation. In at least one embodiment, the geometry / fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front-end unit, a thread spawner and a thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0176] In at least one embodiment, the fixed function block 1630 also includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. The graphics SoC interface 1637 provides an interface between the graphics core 1600 and other processor cores within the SoC integrated circuit. In at least one embodiment, the graphics microcontroller 1638 is a programmable sub-processor that can be configured to manage various different functions of the graphics processor 1600, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 1639 includes logic that facilitates decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, the media pipeline 1639 implements media operations via requests to the computing or sampling logic within the sub-cores 1601-1601F.

[0177] In at least one embodiment, the SoC interface 1637 enables the graphics core 1600 to communicate with a general-purpose application processor core (such as a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared LLC memory, system RAM, and / or embedded on-chip or package DRAM. In at least one embodiment, the SoC interface 1637 may also allow for communication with fixed function devices within the SoC, such as a camera imaging pipeline, and allow for the use and / or implementation of global memory atoms that can be shared between the graphics core 1600 and the CPU within the SoC. In at least one embodiment, the SoC interface 1637 may also implement power management control for the graphics core 1600 and allow for the implementation of an interface between the clock domain of the graphics core 1600 and other clock domains within the SoC. In at least one embodiment, the SoC interface 1637 allows for the receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, the commands and instructions may be dispatched to the media pipeline 1639 when media operations are to be implemented, or to the geometry and fixed function pipelines (such as the geometry and fixed function pipeline 1636, the geometry and fixed function pipeline 1614) when graphics processing operations are to be implemented.

[0178] In at least one embodiment, the graphics microcontroller 1638 may be configured to implement various different scheduling and management tasks for the graphics core 1600. In at least one embodiment, the graphics microcontroller 1638 may perform graphics and / or compute workload scheduling on the execution unit (EU) arrays 1602A-1602F within the sub-cores 1601A-1601F and various different graphics parallel engines within 1604A-1604F. In at least one embodiment, host software executing on a CPU core of an SoC including the graphics core 1600 may submit a workload to one of a plurality of graphics processor doorbells, which invokes a scheduling operation on an appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload to run next, submitting the workload to a command streamer, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 1638 may also facilitate low-power or idle states of the graphics core 1600, provide independence from the operating system and / or graphics driver software on the system to the graphics core 1600, and the ability to save and restore registers within the graphics core 1600 across low-power state transitions.

[0179] In at least one embodiment, the graphics core 1600 may have more or fewer sub-cores 1601A-1601F than illustrated, up to N modular sub-cores. In at least one embodiment, for each group of N sub-cores, the graphics core 1600 may also include shared functional logic 1610, shared and / or cache memory 1612, a geometry / fixed-function pipeline 1614, and additional fixed-function logic 1616 to accelerate various different graphics and compute processing operations. In at least one embodiment, the shared functional logic 1610 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) that may be shared by each N sub-cores within the graphics core 1600. The shared and / or cache memory 1612 may be an LLC for the N sub-cores 1601A-1601F within the graphics core 1600 and may also be used as a shared memory accessible by multiple sub-cores. In at least one embodiment, the geometry / fixed-function pipeline 1614 may be included instead of the geometry / fixed-function pipeline 1636 within the fixed-function block 1630 and may include the same or similar logic units.

[0180] In at least one embodiment, the graphics core 1600 includes additional fixed function logic 1616, which may include various different fixed function acceleration logics for use by the graphics core 1600. In at least one embodiment, the additional fixed function logic 1616 includes an additional geometry pipeline for position-only shading. In position-only shading, there are at least two geometry pipelines, and in the full geometry pipeline within the geometry / fixed function pipelines 1616, 1636, the cull pipeline is an additional geometry pipeline that can be included within the additional fixed function logic 1616. In at least one embodiment, the cull pipeline is a trimmed-down version of the full geometry pipeline. In at least one embodiment, the full pipeline and the cull pipeline can execute different instances of a certain application, each instance having a separate context. In at least one embodiment, position-only shading can hide the long culling runs of discarded triangles, enabling shading to complete earlier in some instances. For example, in at least one embodiment, the cull pipeline logic within the additional fixed function logic 1616 can execute the position shader in parallel with the main application and generally generate key results faster than the full pipeline because the cull pipeline fetches and shades the position attributes of vertices without performing rasterization and reproducing pixels to the frame buffer. In at least one embodiment, the cull pipeline can use the generated key results to calculate the visibility information for all triangles regardless of whether those triangles are culled. In at least one embodiment, the full pipeline (which may be referred to as the replay pipeline in this instance) can consume the visibility information to skip the culled triangles in order to shade only the visible triangles that are ultimately passed to the rasterization stage.

[0181] In at least one embodiment, the additional fixed function logic 1616 may also include general processing acceleration logic such as fixed function matrix multiplication logic to accelerate CUDA programs.

[0182] In at least one embodiment, each graphics sub-core 1601A - 1601F includes a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-cores 1601A - 1601F include multiple EU arrays 1602A - 1602F, 1604A - 1604F, thread dispatch and inter-thread communication ("TD / IC") logic 1603A - 1603F, 3D (e.g., texture) samplers 1605A - 1605F, media samplers 1606A - 1606F, shader processors 1607A - 1607F, and shared local memory ("SLM") 1608A - 1608F. Each of the EU arrays 1602A - 1602F, 1604A - 1604F includes multiple execution units that are GPGPUs capable of performing floating-point and integer / fixed-point logical operations for servicing graphics, media, or compute operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 1603A - 1603F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 1605A - 1605F can read texture or other 3D graphics-related data into memory. In at least one embodiment, the 3D sampler can read texture data differently based on the configured sample state and the texture format associated with a given texture. In at least one embodiment, the media samplers 1606A - 1606F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 1601A - 1601F can alternatively include a unified 3D and media sampler. In at least one embodiment, the threads executing on the execution units within each of the sub-cores 1601A - 1601F can utilize the shared local memory 1608A - 1608F within each sub-core such that the threads executing within a thread group can execute using a common pool of on-chip memory.

[0183] Figure 17The figure shows a parallel processing unit (“PPU”) 1700 in accordance with at least one embodiment. In at least one embodiment, the PPU 1700 is configured with machine-readable code that, when executed by the PPU 1700, causes the PPU 1700 to perform some or all of the processes and techniques described herein. In at least one embodiment, the PPU 1700 is a multi-threaded processor that is implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique that is designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) in parallel across multiple threads. In at least one embodiment, a thread refers to a thread of execution and is an instantiation of a set of instructions configured to be executed by the PPU 1700. In at least one embodiment, the PPU 1700 is a GPU that is configured to implement a graphics rendering pipeline for processing three-dimensional (“3D”) graphics data to generate two-dimensional (“2D”) image data for display on a display device such as an LCD device. In at least one embodiment, the PPU 1700 is utilized to perform computations such as linear algebra operations and machine learning operations. Figure 17 An exemplary parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture that may be implemented in at least one embodiment.

[0184] In at least one embodiment, one or more PPU 1700s are configured to accelerate high-performance computing (“HPC”), data center, and machine learning applications. In at least one embodiment, one or more PPU 1700s are configured to accelerate CUDA programs. In at least one embodiment, PPU 1700 includes, but is not limited to, I / O unit 1706, front-end unit 1710, scheduler unit 1712, work distribution unit 1714, hub 1716, crossbar (“Xbar”) 1720, one or more general processing clusters (“GPC”) 1718, and one or more partition units (“memory partition units”) 1722. In at least one embodiment, PPU 1700 is connected to a host processor or other PPU 1700 via one or more high-speed GPU interconnects (“GPU interconnects”) 1708. In at least one embodiment, PPU 1700 is connected to a host processor or other peripheral devices via interconnect 1702. In at least one embodiment, PPU 1700 is connected to local memory including one or more memory devices (“memory”) 1704. In at least one embodiment, memory device 1704 includes, but is not limited to, one or more dynamic random access memory (DRAM) devices. In at least one embodiment, one or more DRAM devices are configured as and / or configurable as a high-bandwidth memory (“HBM”) subsystem having multiple DRAM dies stacked within each device.

[0185] In at least one embodiment, high-speed GPU interconnect 1708 may refer to a wired multi-lane communication link used by the system to expand and include one or more PPU 1700s in combination with one or more CPUs, support cache coherence between the PPU 1700 and the CPU, and CPU master control. In at least one embodiment, data and / or commands are transmitted to / from other units of PPU 1700, such as one or more copy engines, video encoders, video decoders, power management units, and Figure 17 other components that may not be explicitly illustrated in

[0186] In at least one embodiment, I / O unit 1706 is configured to receive from a host processor (not shown in Figure 17The figure shows the transmission and reception of communications (such as commands, data). In at least one embodiment, the I / O unit 1706 communicates with the host processor directly via the system bus 1702 or through one or more intermediate devices such as a memory bridge. In at least one embodiment, the I / O unit 1706 may communicate with one or more other processors such as one or more of the PPU 1700 via the system bus 1702. In at least one embodiment, the I / O unit 1706 implements a PCIe interface for communication via the PCIe bus. In at least one embodiment, the I / O unit 1706 implements an interface for communicating with external devices.

[0187] In at least one embodiment, the I / O unit 1706 decodes data packets received via the system bus 1702. In at least one embodiment, at least some of the data packets represent commands configured to cause the PPU 1700 to perform various different operations. In at least one embodiment, the I / O unit 1706 transmits the decoded commands to various different other units of the PPU 1700 as specified by the commands. In at least one embodiment, the commands are transmitted to the front-end unit 1710 and / or to the hub 1716 or other units of the PPU 1700, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown in Figure 17 the figure). In at least one embodiment, the I / O unit 1706 is configured to route communications between and among various different logical units of the PPU 1700.

[0188] In at least one embodiment, a command stream of a workload is provided to the PPU 1700 for processing in a program encoding buffer executed by the host processor. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in a memory that can be accessed (e.g., read / written) by both the host processor and the PPU 1700 - the host interface unit may be configured to access the buffer in the system memory connected to the system bus 1702 via a memory request transmitted through the I / O unit 1706 via the system bus 1702. In at least one embodiment, the host processor writes the command stream to the buffer and then transmits a pointer pointing to the start of the command stream to the PPU 1700, such that the front-end unit 1710 receives pointers to one or more command streams and manages the one or more command streams, reads commands from the command streams, and forwards the commands to various different units of the PPU 1700.

[0189] In at least one embodiment, the front-end unit 1710 is coupled to a scheduler unit 1712 that configures the various different GPCs 1718 to handle tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 1712 is configured to track status information related to the various different tasks managed by the scheduler unit 1712, where the status information may indicate which GPC 1718 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and so on. In at least one embodiment, the scheduler unit 1712 manages the execution of multiple tasks on one or more of the GPCs 1718.

[0190] In at least one embodiment, the scheduler unit 1712 is coupled to a work distribution unit 1714 that is configured to dispatch tasks for execution on the GPCs 1718. In at least one embodiment, the work distribution unit 1714 tracks a certain number of scheduled tasks received from the scheduler unit 1714, and the work distribution unit 1714 manages a pending task pool and an active task pool for each GPC 1718. In at least one embodiment, the pending task pool includes a certain number of slots (e.g., 32 slots) containing tasks assigned to be processed by a particular GPC 1718; the active task pool may include a certain number of slots (e.g., 4 slots) for tasks being actively processed by the GPC 1718, such that when one of the GPCs 1718 finishes executing a task, that task is evicted from the active task pool for the GPC 1718, and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 1718. In at least one embodiment, if an active task is idle on the GPC 1718, e.g., when waiting to resolve data dependencies, then that active task is evicted from the GPC 1718 and returned to the pending task pool, while another task from the pending task pool is selected and scheduled for execution on the GPC 1718.

[0191] In at least one embodiment, the work distribution unit 1714 communicates with one or more GPCs 1718 via an XBar 1720. In at least one embodiment, the XBar 1720 is an interconnection network that couples many of the units of the PPU 1700 to other units of the PPU 1700 and can be configured to couple the work distribution unit 1714 to a particular GPC 1718. In at least one embodiment, one or more other units of the PPU 1700 can also be connected to the XBar 1720 via a hub 1716.

[0192] In at least one embodiment, tasks are managed by a scheduler unit 1712 and dispatched by a work distribution unit 1714 to one of the GPCs 1718. The GPC 1718 is configured to process the tasks and generate results. In at least one embodiment, the results can be consumed by other tasks within the GPC 1718, routed to a different GPC 1718 via the XBar 1720, or stored in the memory 1704. In at least one embodiment, the results can be written to the memory 1704 via a partitioning unit 1722 that implements a memory interface for reading and writing data to / from the memory 1704. In at least one embodiment, the results can be transmitted via a high-speed GPU interconnect 1708 to another PPU 1704 or CPU. In at least one embodiment, the PPU 1700 includes, but is not limited to, a number of partitioning units 1722 equal to the number of separate and distinct memory devices 1704 coupled to the PPU 1700.

[0193] In at least one embodiment, a host processor executes a driver kernel that implements an application programming interface (“API”) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 1700. In at least one embodiment, multiple computing applications are executed simultaneously by the PPU 1700, and the PPU 1700 provides isolation, quality of service (“QoS”), and an independent address space for the multiple computing applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 1700, and the driver kernel outputs the tasks to one or more streams being processed by the PPU 1700. In at least one embodiment, each task includes one or more groups of related threads that can be referred to as a warp. In at least one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, cooperating threads can refer to multiple threads that include instructions for implementing a task and exchange data via shared memory.

[0194] Figure 18 The figure illustrates a GPC 1800 in accordance with at least one embodiment. In at least one embodiment, the GPC 1800 is Figure 17 the GPC 1718. In at least one embodiment, each GPC 1800 includes, but is not limited to, a number of hardware units for processing tasks, and each GPC 1800 includes, but is not limited to, a pipeline manager 1802, a pre-raster operation unit (“PROP”) 1804, a raster engine 1808, a work distribution crossbar (“WDX”) 1816, an MMU 1818, one or more data processing clusters (“DPC”) 1806, and any suitable combination of components.

[0195] In at least one embodiment, the operation of the GPC 1800 is controlled by the pipeline manager 1802. In at least one embodiment, the pipeline manager 1802 manages the configuration of one or more DPCs 1806 for processing tasks assigned to the GPC 1800. In at least one embodiment, the pipeline manager 1802 configures at least one of the one or more DPCs 1806 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, the DPC 1806 is configured to execute a vertex shader program on a programmable streaming multiprocessor (“SM”) 1814. In at least one embodiment, the pipeline manager 1802 is configured to route data packets received from a work distribution unit to appropriate logic units within the GPC 1800, and in at least one embodiment, some data packets may be routed to fixed function hardware units in the PROP 1804 and / or the raster engine 1808, while other data packets may be routed to the DPC 1806 for processing by the primitive engine 1812 or the SM 1814. In at least one embodiment, the pipeline manager 1802 configures at least one of the DPCs 1806 to implement a compute pipeline. In at least one embodiment, the pipeline manager 1802 configures at least one of the DPCs 1806 to execute at least a portion of a CUDA program.

[0196] In at least one embodiment, the PROP unit 1804 is configured to route data generated by the raster engine 1808 and the DPC 1806 to, such as in connection with the above Figure 17A raster operation ("ROP") unit in a partitioning unit such as the memory partitioning unit 1722 described in more detail. In at least one embodiment, the PROP unit 1804 is configured to be optimized for color mixing, organize pixel data, perform address translation, and so on. In at least one embodiment, the raster engine 1808 includes, but is not limited to, a number of fixed function hardware units configured to perform various different raster operations, and in at least one embodiment, the raster engine 1808 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile merge engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are transmitted to the coarse raster engine to generate coverage information for the primitives (e.g., an x,y coverage mask for a tile); the output of the coarse raster engine is transmitted to the culling engine where fragments associated with the primitives that do not pass the z-test are culled, and is transmitted to the clipping engine where fragments located outside the viewing frustum are clipped. In at least one embodiment, the fragments that survive culling and clipping are passed to the fine raster engine to generate attributes for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of the raster engine 1808 includes fragments to be processed by any suitable entity, such as a fragment shader implemented within the DPC 1806.

[0197] In at least one embodiment, each DPC 1806 included in the GPC 1800 includes, but is not limited to, an M-type pipeline controller ("MPC") 1810, a primitive engine 1812, one or more SMs 1814, and any suitable combination thereof. In at least one embodiment, the MPC 1810 controls the operation of the DPC 1806 and routes data packets received from the pipeline manager 1802 to the appropriate units within the DPC 1806. In at least one embodiment, data packets associated with vertices are routed to the primitive engine 1812, which is configured to fetch vertex attributes associated with the vertices from memory; in contrast, data packets associated with shader programs can be transmitted to the SM 1814.

[0198] In at least one embodiment, SM 1814 includes, but is not limited to, programmable streaming processors configured to process tasks represented by a number of threads. In at least one embodiment, SM 1814 is multi-threaded and configured to concurrently execute multiple threads (e.g., 32 threads) from a particular thread group, and implements a SIMD architecture where each thread in a thread group (e.g., a warp) is configured to process different data sets based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instruction. In at least one embodiment, SM 1814 implements a SIMT architecture where each thread in a thread group is configured to process different data sets based on the same instruction set, but where individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, allowing for concurrency between warps and serial execution within a warp when threads within the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, allowing for the same concurrency between all threads within and between the number of threads. In at least one embodiment, an execution state is maintained for each individual thread, and threads executing the same instruction can converge and execute in parallel for increased efficiency. At least one embodiment of SM 1814 is described in conjunction with Figure 19 in more detail.

[0199] In at least one embodiment, MMU 1818 provides an interface between GPC 1800 and a memory partition unit (e.g., Figure 17 partition unit 1722), and MMU 1818 provides virtual address-physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, MMU 1818 provides one or more translation lookaside buffers (TLBs) for implementing the translation of virtual addresses to physical addresses in memory.

[0200] Figure 19 FIG. illustrates a streaming multi-processor (“SM”) 1900 in accordance with at least one embodiment. In at least one embodiment, SM 1900 is Figure 18SM 1814. In at least one embodiment, SM 1900 includes, but is not limited to, an instruction cache 1902, one or more scheduler units 1904, a register file 1908, one or more processing cores (“cores”) 1910, one or more special function units (“SFUs”) 1912, one or more LSUs 1914, an interconnect network 1916, a shared memory / L1 cache 1918, and any suitable combination thereof. In at least one embodiment, a work distribution unit dispatches tasks for execution on a GPC of a parallel processing unit (PPU), each task being assigned to a specific data processing cluster (DPC) within the GPC, and if the task is associated with a shader program, then the task is assigned to one of the SM 1900s. In at least one embodiment, the scheduler unit 1904 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to the SM 1900. In at least one embodiment, the scheduler unit 1904 schedules thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 1904 manages multiple different thread blocks, assigns warps to different thread blocks, and then during each clock cycle dispatches instructions from multiple different cooperating groups to various different functional units (e.g., processing cores 1910, SFUs 1912, and LSUs 1914).

[0201] In at least one embodiment, a "cooperative group" may refer to a programming model for organizing groups of communication threads, which allows developers to express the granularity of thread communication and enables a richer and more efficient expression of parallel decomposition. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks for the execution of parallel algorithms. In at least one embodiment, the API of a conventional programming model provides a single simple construct for synchronizing cooperative threads: a barrier for all threads across a thread block (e.g., the syncthreads() function). However, in at least one embodiment, a programmer can define groups of threads with a finer granularity than a thread block and synchronize within the defined groups to allow for greater performance, design flexibility, and software reuse in the form of collective group-wide functional interfaces. In at least one embodiment, software groups enable a programmer to explicitly define groups of threads at sub-block and multi-block granularities and perform collective operations such as thread synchronization within the cooperative group. In at least one embodiment, the sub-block granularity is as small as a single thread. In at least one embodiment, the programming model supports clean composition across software boundaries such that libraries and utility functions can synchronize safely within their local context without making assumptions about convergence. In at least one embodiment, cooperative group primitives allow for the implementation of new modes of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire grid of thread blocks.

[0202] In at least one embodiment, the dispatch unit 1906 is configured to transmit instructions to one or more of the functional units, and the scheduler unit 1904 includes two dispatch units 1906 that allow, among other things, two different instructions from the same warp to be dispatched during each clock cycle. In at least one embodiment, each scheduler unit 1904 includes a single dispatch unit 1906 or additional dispatch units 1906.

[0203] In at least one embodiment, each SM 1900 includes, but is not limited to, a register file 1908 that provides a register set for the functional units of the SM 1900. In at least one embodiment, the register file 1908 is partitioned among each of the functional units such that each functional unit is assigned a dedicated portion of the register file 1908. In at least one embodiment, the register file 1908 is partitioned among different warps being executed by the SM 1900, and the register file 1908 provides temporary storage for the operands of the data paths connected to the functional units. In at least one embodiment, each SM 1900 includes, but is not limited to, a plurality of L processing cores 1910. In at least one embodiment, the SM 1900 includes, but is not limited to, a large number (e.g., 128 or more) of different processing cores 1910. In at least one embodiment, each processing core 1910 includes, but is not limited to, fully pipelined, single-precision, double-precision, and / or mixed-precision processing units, which include, but are not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In at least one embodiment, the processing core 1910 includes, but is not limited to, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0204] In at least one embodiment, the tensor cores are configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in the processing core 1910. In at least one embodiment, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolutional operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4x4 matrix and performs matrix multiplication and accumulation operations D = A X B + C, where A, B, C, and D are 4x4 matrices.

[0205] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core operates on 16-bit floating-point input data using 32-bit floating-point accumulation. In at least one embodiment, for a 4x4x4 matrix multiplication, the 16-bit floating-point multiplication uses 64 operations and results in a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point operations. In at least one embodiment, the tensor core is used to perform much larger two-dimensional or higher-dimensional matrix operations composed of these smaller elements. In at least one embodiment, APIs such as the CUDA-C++ API expose dedicated matrix load, matrix multiplication and accumulation, and matrix store operations for efficient use of the tensor core from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes that a matrix of size 16x16 spans all 32 threads of a warp.

[0206] In at least one embodiment, each SM 1900 includes, but is not limited to, M SFUs 1912 that perform special functions (such as attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU 1912 includes, but is not limited to, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU 1912 includes, but is not limited to, a texture unit configured to perform texture mapping filtering operations. In at least one embodiment, the texture unit is configured to load texture maps (such as 2D texel arrays) and sample texture maps from memory to generate sampled texture values used in shader programs executed by the SM 1900. In at least one embodiment, the texture maps are stored in the shared memory / L1 cache 1918. In at least one embodiment, the texture unit uses mip maps (such as texture maps with varying levels of detail) to implement texture operations such as filtering operations. In at least one embodiment, each SM 1900 includes, but is not limited to, two texture units.

[0207] In at least one embodiment, each SM 1900 includes, but is not limited to, N LSUs 1914 that implement load and store operations between a shared memory / L1 cache 1918 and a register file 1908. In at least one embodiment, each SM 1900 includes, but is not limited to, an interconnect network 1916 that connects each of the functional units to the register file 1908 and connects the LSUs 1914 to the register file 1908 and the shared memory / L1 cache 1918. In at least one embodiment, the interconnect network 1916 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 1908 and connect the LSUs 1914 to the register file 1908 and memory locations in the shared memory / L1 cache 1918.

[0208] In at least one embodiment, the shared memory / L1 cache 1918 is an on-chip memory array that allows data storage and communication between the SM 1900 and the primitive engine and between threads in the SM 1900. In at least one embodiment, the shared memory / L1 cache 1918 includes, but is not limited to, a storage capacity of 128 KB and is on the path from the SM 1900 to the partitioning unit. In at least one embodiment, the shared memory / L1 cache 1918 is used to cache reads and writes. In at least one embodiment, one or more of the shared memory / L1 cache 1918, the L2 cache, and the memory are backing stores.

[0209] In at least one embodiment, combining data caching and shared memory functionality into a single memory block provides improved performance for both types of memory access. In at least one embodiment, the capacity is used as or can be used by programs that do not use the shared memory as a cache. For example, if the shared memory is configured to use half of its capacity, then texture and load / store operations can use the remaining capacity. In at least one embodiment, integration into the shared memory / L1 cache 1918 enables the shared memory / L1 cache 1918 to be used as a high-throughput conduit for streaming data, while providing high-bandwidth and low-latency access to data that is frequently reused. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function GPU is bypassed, creating a much simpler programming model. In at least one embodiment and in a general-purpose parallel computing configuration, the work distribution unit directly allocates and distributes thread blocks to the DPC. In at least one embodiment, the threads in a block execute the same program, using the unique thread ID in the computation to ensure that each thread generates a unique result, using the SM 1900 to execute the program and perform the computation, using the shared memory / L1 cache 1918 to communicate between threads, and using the LSU 1914 to read and write global memory through the shared memory / L1 cache 1918 and the memory partition unit. In at least one embodiment, when configured for general-purpose parallel computing, the SM 1900 write scheduler unit 1904 can be used to issue commands to start new work on the DPC.

[0210] In at least one embodiment, the PPU is included in or coupled to a desktop computer, laptop computer, tablet computer, server, supercomputer, smart phone (e.g., a wireless handheld device), PDA, digital camera, vehicle, head-mounted display, handheld electronic device, and the like. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a SoC together with one or more other devices, such as additional PPUs, memories, RISC CPUs, MMUs, digital-to-analog converters ("DACs"), and the like.

[0211] In at least one embodiment, the PPU can be included on a graphics card that includes one or more memory devices. In at least one embodiment, the graphics card can be configured to interface with a PCIe interface on the motherboard of a desktop computer. In at least one embodiment, the PPU can be an integrated GPU ("iGPU") included in a motherboard chipset.

[0212] Software constructs for general computing

[0213] The following figures illustrate, without limitation, an exemplary software architecture for implementing at least one embodiment.

[0214] Figure 20 The figure shows a software stack of a programming platform according to at least one embodiment. In at least one embodiment, the programming platform is a platform for leveraging hardware acceleration on a computing system for computing tasks. In at least one embodiment, the programming platform may be accessible to software developers through libraries, compiler directives, and / or extensions to programming languages. In at least one embodiment, the programming platform may be, but is not limited to, CUDA, Radeon Open Compute Platform (“ROCm”), OpenCL (OpenCLTM is developed by the Khronos Group), SYCL, or Intel One API.

[0215] In at least one embodiment, the software stack 2000 of the programming platform provides an execution environment for an application 2001. In at least one embodiment, the application 2001 may include any computer software capable of being launched on the software stack 2000. In at least one embodiment, the application 2001 may include, but is not limited to, artificial intelligence (“AI”) / machine learning (“ML”) applications, high-performance computing (“HPC”) applications, virtual desktop infrastructure (“VDI”), or data center workloads.

[0216] In at least one embodiment, the application 2001 and the software stack 2000 run on hardware 2007. In at least one embodiment, the hardware 2007 may include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support the programming platform. In at least one embodiment, for example, for CUDA, the software stack 2000 may be vendor-specific and only compatible with devices from a specific vendor. In at least one embodiment, for example, for OpenCL, the software stack 2000 may be used with devices from different vendors. In at least one embodiment, the hardware 2007 includes a host connected to one or more devices that can be accessed via application programming interface (“API”) calls to perform computing tasks. In at least one embodiment, the devices within the hardware 2007 may include, but are not limited to, GPUs, FPGAs, AI engines, or other computing devices (but may also include CPUs) and their memories, as opposed to the host within the hardware 2007 that may include, but is not limited to, CPUs (but may also include computing devices) and their memories.

[0217] In at least one embodiment, the software stack 2000 of the programming platform includes, but is not limited to, a number of libraries 2003, a runtime 2005, and a device kernel driver 2006. In at least one embodiment, each of the libraries 2003 may include data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, the libraries 2003 may include, but are not limited to, pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. In at least one embodiment, the libraries 2003 include functions optimized for execution on one or more types of devices. In at least one embodiment, the libraries 2003 may include, but are not limited to, functions for performing mathematics, deep learning, and / or other types of operations on the device. In at least one embodiment, the libraries 2103 are associated with corresponding APIs 2102, which may include one or more APIs that expose the functions implemented in the libraries 2103.

[0218] In at least one embodiment, as discussed in more detail below in conjunction with Figures 25 - 27 The application 2001 is written as source code, which is compiled into executable code. In at least one embodiment, the executable code of the application 2001 can run at least in part on the execution environment provided by the software stack 2000. In at least one embodiment, during the execution of the application 2001, as opposed to the host, code that needs to run on the device can be accessed. In at least one embodiment, in such a case, the runtime 2005 can be called to load and start the necessary code on the device. In at least one embodiment, the runtime 2005 can include any technically feasible runtime system capable of supporting the execution of the application S01.

[0219] In at least one embodiment, the runtime 2005 is implemented as one or more runtime libraries associated with a corresponding API shown as API 2004. In at least one embodiment, one or more of such runtime libraries may include, but are not limited to, functions for memory management, execution control, device management, error handling, and / or synchronization among others. In at least one embodiment, the memory management functions may include, but are not limited to, functions for allocating, deallocating, and copying device memory and transferring data between host memory and device memory. In at least one embodiment, the execution control functions may include, but are not limited to, starting a function on the device (sometimes referred to as a "kernel" when the function is a global function callable from the host) and setting property values in buffers maintained by the runtime library for a given function to be executed on the device.

[0220] In at least one embodiment, the runtime library and corresponding API 2004 can be implemented in any technically feasible manner. In at least one embodiment, one (or any number of) APIs can expose a set of low-level functions for fine-grained control of the device, while another (or any number of) APIs can expose a higher-level set of such functions. In at least one embodiment, the high-level runtime API can be built on top of the low-level API. In at least one embodiment, one or more of the runtime APIs can be language-specific APIs that are layered on top of language-independent runtime APIs.

[0221] In at least one embodiment, the device kernel driver 2006 is configured to facilitate communication with the underlying device. In at least one embodiment, the device kernel driver 2006 can provide low-level functions that APIs such as API 2004 and / or other software rely on. In at least one embodiment, the device kernel driver 2006 can be configured to compile intermediate representation (“IR”) code into binary code at runtime. In at least one embodiment, for CUDA, the device kernel driver 2006 can compile parallel thread execution (“PTX”) IR code that is not specific to the hardware into binary code for a specific target device (cached compiled binary code) at runtime, which is sometimes also referred to as “finalized” code. In at least one embodiment, doing so can allow the finalized code to run on the target device, which may not have existed when the source code was initially compiled into PTX code. Alternatively, in at least one embodiment, the device source code can be compiled into binary code offline without the device kernel driver 2006 compiling the IR code at runtime.

[0222] Figure 21 illustrates a CUDA implementation of the Figure 20 software stack 2000 according to at least one embodiment. In at least one embodiment, the CUDA software stack 2100 on which an application 2001 can be launched includes a CUDA library 2103, a CUDA runtime 2105, a CUDA driver 2107, and a device kernel driver 2108. In at least one embodiment, the CUDA software stack 2100 executes on hardware 2109, which can include a GPU that supports CUDA and is developed by NVIDIA Corporation of Santa Clara, California.

[0223] In at least one embodiment, the application 2101, the CUDA runtime 2105, and the device kernel driver 2108 can be implemented respectively in combination with the above Figure 20The described application 2001, runtime 2005, and device kernel driver 2006 have similar functions. In at least one embodiment, the CUDA driver 2107 includes a library (libcuda.so) that implements the CUDA driver API 2106. In at least one embodiment, similar to the CUDA runtime API 2104 implemented by the CUDA runtime library (cudart), the CUDA driver API 2106 can non - restrictively expose functions for, among other things, memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability. In at least one embodiment, the CUDA driver API 2106 is different from the CUDA runtime API 2104 because the CUDA runtime API 2104 simplifies device code management by providing implicit initialization, context (similar to a process) management, and module (similar to a dynamically loaded library) management. In at least one embodiment, in contrast to the high - level CUDA runtime API 2104, the CUDA driver API 2106 is a low - level API that provides more fine - grained control of the device, particularly in terms of context and module loading. In at least one embodiment, the CUDA driver API 2106 can expose functions for context management that the CUDA runtime API 2104 does not expose. In at least one embodiment, the CUDA driver API 2106 is also language - independent and supports, for example, OpenCL in addition to the CUDA runtime API 2104. Further, in at least one embodiment, the development libraries, including the CUDA runtime 2105, can be considered separate from the driver components, including the user - mode CUDA driver 2107 and the kernel - mode device driver 2108 (sometimes also referred to as the "display" driver).

[0224] In at least one embodiment, the CUDA library 2103 can include, but is not limited to, math libraries, deep - learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries that parallel computing applications such as the application 2101 may utilize. In at least one embodiment, the CUDA library 2103 can include math libraries, such as the cuBLAS library, which is, among other things, an implementation of the basic linear algebra subprograms ("BLAS") for performing linear algebra operations, the cuFFT library for computing the fast Fourier transform ("FFT"), and the cuRAND library for generating random numbers. In at least one embodiment, the CUDA library 2103 can include deep - learning libraries, such as the cuDNN library for primitives of deep neural networks and the TensorRT platform for high - performance deep - learning inference, among other things.

[0225] Figure 22 illustrates in accordance with at least one embodiment ofFigure 20 The ROCm implementation of the software stack 2000. In at least one embodiment, the ROCm software stack 2200 on which an application 2201 can be launched includes a language runtime 2203, a system runtime 2205, a thunk 2207, a ROCm kernel driver 2208, and a device kernel driver 2209. In at least one embodiment, the ROCm software stack 2200 executes on hardware 2210, which may include a GPU that supports ROCm and is developed by AMD in Santa Clara, California.

[0226] In at least one embodiment, the application 2201 may implement functions similar to those of the application 2001 discussed above in connection with Figure 20 In addition, in at least one embodiment, the language runtime 2203 and the system runtime 2205 may implement functions similar to those of the runtime 2005 discussed above in connection with Figure 20 In at least one embodiment, the difference between the language runtime 2203 and the system runtime 2205 is that the system runtime 2205 is a language-independent runtime that implements the ROCr system runtime API 2204 and utilizes the Heterogeneous System Architecture (“HAS”) runtime API. In at least one embodiment, the HAS runtime API is a thin user-mode API that exposes interfaces for accessing and interacting with an AMD GPU, including functions for, among other things, memory management, execution control via architectural dispatching of kernels, error handling, system and agent information, and runtime initialization and shutdown. In at least one embodiment, in contrast to the system runtime 2205, the language runtime 2203 is an implementation of a language-specific runtime API 2202 that is layered on top of the ROCr system runtime API 2204. In at least one embodiment, the language runtime API may include, but is not limited to, among other things, the Heterogeneous compute Interface for Portability (“HIP”) language runtime API, the Heterogeneous Compute Compiler (“HCC”) language runtime API, or the OpenCL API. In particular, the HIP language is an extension of the C++ programming language and has a functionally similar version of the CUDA mechanism. And in at least one embodiment, the HIP language runtime API includes functions similar to those of the CUDA runtime API 2104 discussed above in connection with Figure 21 such as functions for, among other things, memory management, execution control, device management, error handling, and synchronization.

[0227] In at least one embodiment, the Real-to-OCaml Compiler (ROCt) 2207 is an interface that can be used to interact with the underlying ROCm driver 2208. In at least one embodiment, the ROCm driver 2208 is a ROCk driver, which is a combination of the AMDGPU driver and the HAS kernel driver (amdkfd). In at least one embodiment, the AMDGPU driver is a device kernel driver for GPUs developed by AMD, which implements functions similar to those of the device kernel driver 2006 discussed above in conjunction with Figure 20 the device kernel driver 2006. In at least one embodiment, the HAS kernel driver is a driver that allows different types of processors to more efficiently share system resources via hardware features.

[0228] In at least one embodiment, various different libraries (not shown) can be included in the ROCm software stack 2200 above the language runtime 2203 and provide functions similar to those of the CUDA libraries 2103 discussed above in conjunction with Figure 21 the CUDA libraries 2103. In at least one embodiment, the various different libraries can include, but are not limited to, math, deep learning, and / or other libraries, such as, among others, the hipBLAS library that implements functions similar to those of CUDA cuBLAS, and the rocFFT library for computing FFTs similar to CUDA cuFFT.

[0229] Figure 23 FIG. illustrates an OpenCL implementation of the software stack 2000 of STACK A according to at least one embodiment. In at least one embodiment, the OpenCL software stack 2300 on which an application 2301 can be launched includes an OpenCL framework 2305, an OpenCL runtime 2306, and a driver 2307. In at least one embodiment, the OpenCL software stack 2300 executes on hardware 2109 that is not specific to a vendor. In at least one embodiment, since OpenCL is supported by devices developed by different vendors, a specific OpenCL driver may be required to interoperate with hardware from such vendors.

[0230] In at least one embodiment, the application 2301, the OpenCL runtime 2306, the device kernel driver 2307, and the hardware 2308 can respectively implement functions similar to those of the application 2001, the runtime 2005, the device kernel driver 2006, and the hardware 2007 discussed above in conjunction with Figure 20 the application 2001, the runtime 2005, the device kernel driver 2006, and the hardware 2007. In at least one embodiment, the application 2301 further includes an OpenCL kernel 2302 having code to be executed on the device.

[0231] In at least one embodiment, OpenCL defines a "platform" that allows a host to control devices connected to the host. In at least one embodiment, the OpenCL framework provides platform layer APIs and runtime APIs shown as platform API 2303 and runtime API 2305. In at least one embodiment, the runtime API 2305 uses contexts to manage the execution of kernels on devices. In at least one embodiment, each identified device can be associated with a respective context, and the runtime API 2305 can use that context to manage, among other things, command queues, program objects, and kernel objects, and shared memory objects for that device. In at least one embodiment, the platform API 2303 exposes functions that allow, among other things, selection and initialization of devices using device contexts, submission of work to devices via command queues, and enabling implementation of data transfers to / from devices. Additionally, in at least one embodiment, the OpenCL framework provides a variety of built-in functions (not shown), including, among other things, mathematical functions, relational functions, and image processing functions.

[0232] In at least one embodiment, a compiler 2304 is also included in the OpenCL framework 2305. In at least one embodiment, source code can be compiled offline before execution of an application or online during execution of the application. In at least one embodiment, in contrast to CUDA and ROCm, OpenCL applications can be compiled online by the compiler 2304, which is included to represent any number of compilers that can be used to compile source code and / or IR code, such as standard portable intermediate representation ("SPIR-V") code, into binary code. Alternatively, in at least one embodiment, OpenCL applications can be compiled offline before execution of such applications.

[0233] Figure 24 The figure shows software supported by a programming platform according to at least one embodiment. In at least one embodiment, the programming platform 2404 is configured to support a variety of programming models 2403, middleware, and / or libraries 2402, and frameworks 2401 that applications 2400 may depend on. In at least one embodiment, the application 2400 can be an AI / ML application implemented using a deep learning framework such as, for example, MXNet, PyTorch, or TensorFlow, which may depend on libraries such as cuDNN, NVIDIA Collective Communications Library ("NCCL"), and / or NVIDIA Developer Data Loading Library ("DALI") CUDA libraries to provide accelerated computing on the underlying hardware.

[0234] In at least one embodiment, the programming platform 2404 can be the one described above in connection with Figure 21 , Figure 22 andFigure 23 One of the described CUDA, ROCm, or OpenCL platforms. In at least one embodiment, the programming platform 2404 supports multiple programming modules 2403, which are abstractions of the underlying computing system that allow the expression of algorithms and data structures. In at least one embodiment, the programming module 2403 may expose the characteristics of the underlying hardware to improve performance. In at least one embodiment, the programming module 2403 may include, but is not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism (“C++AMP”), Open Multi-Processing (“OpenMP”), Open Accelerator (“OpenACC”), and / or Vulkan Compute.

[0235] In at least one embodiment, the library and / or middleware 2402 provides an implementation of the abstraction of the programming module 2404. In at least one embodiment, such libraries include data and programming code that can be used by computer programs and exploited during software development. In at least one embodiment, such middleware includes software that provides services other than the services available from the programming platform 2404 to the application. In at least one embodiment, the library and / or middleware 2402 may include, but is not limited to, cuBLAS, cuFFT, cuRAND, and other CUDA libraries, or rocBLAS, rocFFT, rocRAND, and other ROCm libraries. Additionally, in at least one embodiment, the library and / or middleware 2402 may include the NCCL and the ROCm Communication Collective Library (“RCCL”) libraries that provide communication routines for the GPU, the MIOpen library for deep learning acceleration, and / or feature libraries for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.

[0236] In at least one embodiment, the application framework 2401 depends on the library and / or middleware 2402. In at least one embodiment, each of the application frameworks 2401 is a software framework used to implement the standard structure of application software. Returning to the AI / ML example discussed above, in at least one embodiment, the AI / ML application may be implemented using frameworks such as the Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks.

[0237] Figure 25 Illustrated in accordance with at least one embodiment compiled in Figures 20 - 23Code executed on one of the programming platforms. In at least one embodiment, the compiler 2501 receives source code 2500 that includes both host code and device code. In at least one embodiment, the compiler 2501 is configured to convert the source code 2500 into host-executable code 2502 for execution on the host and device-executable code 2503 for execution on the device. In at least one embodiment, the source code 2500 can be compiled either offline before executing the application or online during the execution of the application.

[0238] In at least one embodiment, the source code 2500 can include code in any programming language supported by the compiler 2501, such as C++, C, Fortran, and the like. In at least one embodiment, the source code 2500 can be included in a single source file that mixes host code and device code, and the location of the device code is indicated in the single source file. In at least one embodiment, the single source file can be a.cu file that includes CUDA code or a.hip.cpp file that includes HIP code. Alternatively, in at least one embodiment, the source code 2500 can include multiple source code files instead of a single source file, with the host code and device code separated into them.

[0239] In at least one embodiment, the compiler 2501 is configured to compile the source code 2500 into host-executable code 2502 for execution on the host and device-executable code 2503 for execution on the device. In at least one embodiment, the compiler 2501 performs operations including parsing the source code 2500 into an abstract syntax tree (AST), performing optimizations, and generating executable code. In at least one embodiment where the source code 2500 includes a single source file, as discussed in more detail below with respect to Figure 26 the compiler 2501 can separate the device code and host code in such a single source file, compile the device code and host code into device-executable code 2503 and host-executable code 2502 respectively, and link the device-executable code 2503 and host-executable code 2502 together in a single file.

[0240] In at least one embodiment, the host-executable code 2502 and the device-executable code 2503 can be in any suitable format, such as binary code and / or IR code. In at least one embodiment, in the case of CUDA, the host-executable code 2502 can include native object code, and the device-executable code 2503 can include code in PTX intermediate representation. In at least one embodiment, in the case of ROCm, both the host-executable code 2502 and the device-executable code 2503 can include target binary code.

[0241] Figure 26 is a more detailed illustration of code compiled according to at least one embodiment and executed on one of the programming platforms of Figures 20 - 23 In at least one embodiment, compiler 2601 is configured to receive source code 2600, compile source code 2600, and output an executable file 2608. In at least one embodiment, source code 2600 is a single source file such as a.cu file,.hip.cpp file, or another format file that includes both host and device code. In at least one embodiment, compiler 2601 may be, but is not limited to, the NVIDIA CUDA compiler ("NVCC") for compiling CUDA code in a.cu file, or the HCC compiler for compiling HIP code in a.hip.cpp file.

[0242] In at least one embodiment, compiler 2601 includes a compiler front end 2602, a host compiler 2605, a device compiler 2606, and a linker 2609. In at least one embodiment, compiler front end 2602 is configured to separate device code 2604 from host code 2603 in source code 2600. In at least one embodiment, device code 2604 is compiled by device compiler 2606 into device executable code 2608, which as described may include binary code or IR code. Additionally, in at least one embodiment, host code 2603 is compiled by host compiler 2605 into host executable code 2607. In at least one embodiment, for NVCC, host compiler 2605 may be, but is not limited to, a general C / C++ compiler that outputs native object code, and device compiler 2606 may be, but is not limited to, a compiler based on the low-level virtual machine ("LLVM") that derives from the LLVM compiler infrastructure and outputs PTX code or binary code. In at least one embodiment, for HCC, both host compiler 2605 and device compiler 2606 may be, but is not limited to, LLVM-based compilers that output target binary code.

[0243] In at least one embodiment, after source code 2600 is compiled into host executable code 2607 and device executable code 2608, linker 2609 links the host and device executable codes 2607 and 2608 together in executable file 2610. In at least one embodiment, the native object code for the host and the PTX or binary code for the device can be linked together in an executable and linkable format ("ELF") file, which is a container format for storing object code.

[0244] Figure 27The figure shows the transformation of source code before compiling the source code according to at least one embodiment. In at least one embodiment, the source code 2700 is passed through a transformation tool 2701 that transforms the source code 2700 into transformed source code 2702. In at least one embodiment, a compiler 2703 is used to compile the transformed source code 2702 into host-executable code 2704 and device-executable code 2705 in a process similar to that discussed above in connection with Figure 25 the compiler 2501 compiling the source code 2500 into host-executable code 2502 and device-executable code 2503.

[0245] In at least one embodiment, the transformation performed by the transformation tool 2701 is used to port the source 2700 for execution in an environment different from the environment in which it was originally intended to run. In at least one embodiment, the transformation tool 2701 may include, but is not limited to, a HIP converter for "hipifying" CUDA code intended for the CUDA platform into HIP code that can be compiled and executed on the ROCm platform. In at least one embodiment, as discussed in more detail below in connection with Figures 28A - 29 the transformation of the source code 2700 may include parsing the source code 2700 and converting calls to APIs provided by one programming model (e.g., CUDA) into calls to APIs provided by another programming model (e.g., HIP). Returning to the example of hipifying CUDA, in at least one embodiment, calls to the CUDA runtime API, the CUDA driver API, and / or the CUDA library may be converted into corresponding HIP API calls. In at least one embodiment, the automated transformation performed by the transformation tool 2701 may sometimes be incomplete and require additional manual effort to fully port the source code 2700.

[0246] Configuring the GPU for General-Purpose Computing

[0247] The following figures non-limitingly illustrate an exemplary architecture for compiling and executing computational source code according to at least one embodiment.

[0248] Figure 28AThe figure shows a system 28A00 configured to compile and execute CUDA source code 2810 using different types of processing units, in accordance with at least one embodiment. In at least one embodiment, the system 28A00 includes, but is not limited to, CUDA source code 2810, a CUDA compiler 2850, host executable code 2870(1), host executable code 2870(2), CUDA device executable code 2884, a CPU 2890, a CUDA-enabled GPU 2894, a GPU 2892, a CUDA-HIP conversion tool 2820, HIP source code 2830, a HIP compiler driver 2840, an HCC 2860, and HCC device executable code 2882.

[0249] In at least one embodiment, the CUDA source code 2810 is a collection of human-readable code in the CUDA programming language. In at least one embodiment, the CUDA code is human-readable code in the CUDA programming language. In at least one embodiment, the CUDA programming language is an extension of the C++ programming language that includes, but is not limited to, mechanisms for defining device code and differentiating between device code and host code. In at least one embodiment, device code is a type of source code that can be executed in parallel on a device after compilation. In at least one embodiment, the device can be a processor optimized for parallel instruction processing, such as a CUDA-enabled GPU 2890, a GPU 28192, or another GPGPU, etc. In at least one embodiment, host code is source code that can be executed on a host after compilation. In at least one embodiment, the host is a processor optimized for sequential instruction processing, such as a CPU 2890.

[0250] In at least one embodiment, the CUDA source code 2810 includes, but is not limited to, any number (including zero) of global functions 2812, any number (including zero) of device functions 2814, any number (including zero) of host functions 2816, and any number (including zero) of host / device functions 2818. In at least one embodiment, the global functions 2812, device functions 2814, host functions 2816, and host / device functions 2818 may be mixed in the CUDA source code 2810. In at least one embodiment, each of the global functions 2812 may be executed on the device and called from the host. In at least one embodiment, one or more of the global functions 2812 may thus serve as an entry point to the device. In at least one embodiment, each of the global functions 2812 is a kernel. In at least one embodiment and in a technique called dynamic parallelism, one or more of the global functions 2812 define kernels that may be executed on the device and called from such a device. In at least one embodiment, the kernel is executed N times in parallel by N different threads on the device during execution (where N is any positive integer).

[0251] In at least one embodiment, each of the device functions 2814 is executed on the device and callable only from such a device. In at least one embodiment, each of the host functions 2816 is executed on the host and callable only from such a host. In at least one embodiment, each of the host / device functions 2816 defines both a host version of the function that may be executed on the host and callable only from such a host and a device version of the function that may be executed on the device and callable only from such a device.

[0252] In at least one embodiment, the CUDA source code 2810 may also include, but is not limited to, any number of calls to any number of functions defined via the CUDA runtime API 2802. In at least one embodiment, the CUDA runtime API 2802 may include, but is not limited to, any number of functions for performing operations on the host to allocate and deallocate device memory, transfer data between host memory and device memory, manage a system with multiple devices, and so on. In at least one embodiment, the CUDA source code 2810 may also include any number of calls to any number of functions specified in any number of other CUDA APIs. In at least one embodiment, a CUDA API may be any API designed for use with CUDA code. In at least one embodiment, CUDA APIs include, but are not limited to, the CUDA runtime API 2802, the CUDA driver API, APIs for any number of CUDA libraries, and so on. In at least one embodiment and with respect to the CUDA runtime API 2802, the CUDA driver API is a lower-level API but provides finer-grained control of the device. In at least one embodiment, examples of CUDA libraries include, but are not limited to, cuBLAS, cuFFT, cuRAND, cuDNN, and so on.

[0253] In at least one embodiment, the CUDA compiler 2850 compiles the input CUDA code (e.g., the CUDA source code 2810) to generate host-executable code 2870(1) and CUDA device-executable code 2884. In at least one embodiment, the CUDA compiler 2850 is NVCC. In at least one embodiment, the host-executable code 2870(1) is a compiled version of the host code included in the input source code that is executable on the CPU 2890. In at least one embodiment, the CPU 2890 may be any processor optimized for sequential instruction processing.

[0254] In at least one embodiment, the CUDA device executable code 2884 is a compiled version of the device code included in the input source code that is executable on a CUDA-enabled GPU 2894. In at least one embodiment, the CUDA device executable code 2884 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device executable code 2884 includes, but is not limited to, IR code such as PTX code, which is further compiled at runtime by the device driver into binary code for a specific target device (e.g., the CUDA-enabled GPU 2894). In at least one embodiment, the CUDA-enabled GPU 2894 can be any processor optimized for parallel instruction processing and supporting CUDA. In at least one embodiment, the CUDA-enabled GPU 2894 is developed by NVIDIA Corporation of Santa Clara, California.

[0255] In at least one embodiment, the CUDA-HIP conversion tool 2820 is configured to convert the CUDA source code 2810 into a functionally similar HIP source code 2830. In at least one embodiment, the HIP source code 2830 is a collection of human-readable code in the HIP programming language. In at least one embodiment, HIP code is human-readable code in the HIP programming language. In at least one embodiment, the HIP programming language is an extension of the C++ programming language that includes, but is not limited to, functionally similar versions of the CUDA mechanisms that define device code and distinguish device code from host code. In at least one embodiment, the HIP programming language may include a functional subset of the CUDA programming language. In at least one embodiment, for example, the HIP programming language includes, but is not limited to, a mechanism for defining a global function 2812, but such a HIP programming language may lack support for dynamic parallelism and thus the global function 2812 defined in the HIP code may only be callable from the host.

[0256] In at least one embodiment, the HIP source code 2830 includes, but is not limited to, any number (including zero) of global functions 2812, any number (including zero) of device functions 2814, any number (including zero) of host functions 2816, and any number (including zero) of host / device functions 2818. In at least one embodiment, the HIP source code 2830 may also include any number of calls to any number of functions specified in the HIP runtime API 2832. In at least one embodiment, the HIP runtime API 2832 includes, but is not limited to, a functionally similar version of a subset of the functions included in the CUDA runtime API 2802. In at least one embodiment, the HIP source code 2830 may also include any number of calls to any number of functions specified in any number of other HIP APIs. In at least one embodiment, a HIP API may be any API designed for use with HIP code and / or ROCm. In at least one embodiment, HIP APIs include, but are not limited to, the HIP runtime API 2832, the HIP driver API, the APIs for any number of HIP libraries, the APIs for any number of ROCm libraries, and so on.

[0257] In at least one embodiment, the CUDA-HIP conversion tool 2820 converts each kernel call in the CUDA code from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the CUDA code to any number of other functionally similar HIP calls. In at least one embodiment, a CUDA call is a call to a function specified in the CUDA API, and a HIP call is a call to a function specified in the HIP API. In at least one embodiment, the CUDA-HIP conversion tool 2820 converts any number of calls to functions specified in the CUDA runtime API 2802 to any number of calls to functions specified in the HIP runtime API 2832.

[0258] In at least one embodiment, the CUDA-HIP conversion tool 2820 is a tool called hipify-perl that performs a text-based conversion process. In at least one embodiment, the CUDA-HIP conversion tool 2820 is a tool called hipify-clang that performs a more complex and more robust conversion process relative to hipify-perl, which involves using clang (a compiler front-end) to parse the CUDA code and then converting the resulting symbols. In at least one embodiment, properly converting CUDA code to HIP code may require modifications (such as manual editing) in addition to those implemented by the CUDA-HIP conversion tool 2820.

[0259] In at least one embodiment, the HIP compiler driver 2840 is a front end that determines the target device 2846 and then configures a compiler compatible with the target device 2846 to compile the HIP source code 2830. In at least one embodiment, the target device 2846 is a processor optimized for parallel instruction processing. In at least one embodiment, the HIP compiler driver 2840 can determine the target device 2846 in any technically feasible manner.

[0260] In at least one embodiment, if the target device 2846 is CUDA compatible (e.g., a CUDA-enabled GPU 2894), then the HIP compiler driver 2840 generates HIP / NVCC compilation commands 2842. In at least one embodiment and as described in more detail in conjunction with Figure 28B the HIP / NVCC compilation commands 2842 configure the CUDA compiler 2850 to compile the HIP source code 2830 using, without limitation, HIP-CUDA translation headers and the CUDA runtime library. In at least one embodiment and in response to the HIP / NVCC compilation commands 2842, the CUDA compiler 2850 generates host-executable code 2870(1) and CUDA device-executable code 2884.

[0261] In at least one embodiment, if the target device 2846 is not CUDA compatible, then the HIP compiler driver 2840 generates HIP / HCC compilation commands 2844. In at least one embodiment and as described in more detail in conjunction with Figure 28C the HIP / HCC compilation commands 2844 configure the HCC 2860 to compile the HIP source code 2830 using, without limitation, HCC headers and the HIP / HCC runtime library. In at least one embodiment and in response to the HIP / HCC compilation commands 2844, the HCC 2860 generates host-executable code 2870(2) and HCC device-executable code 2882. In at least one embodiment, the HCC device-executable code 2882 is a compiled version of the device code included in the HIP source code 2830 that is executable on the GPU 2892. In at least one embodiment, the GPU 2892 can be any processor optimized for parallel instruction processing, not CUDA compatible, and HCC compatible. In at least one embodiment, the GPU 2892 is developed by AMD of Santa Clara, California. In at least one embodiment, the GPU 2892 is a non-CUDA-enabled GPU 2892.

[0262] For purposes of explanation only, Figure 28AThree different streams are depicted that can be implemented in at least one embodiment to compile CUDA source code 2810 for execution on a CPU 2890 and different devices. In at least one embodiment, a direct CUDA stream compiles the CUDA source code 2810 without converting the CUDA source code 2810 into HIP source code 2830 for execution on a CPU 2890 and a CUDA-enabled GPU 2894. In at least one embodiment, an indirect CUDA stream converts the CUDA source code 2810 into HIP source code 2830 and then compiles the HIP source code 2830 for execution on a CPU 2890 and a CUDA-enabled GPU 2894. In at least one embodiment, a CUDA / HCC stream converts the CUDA source code 2810 into HIP source code 2830 and then compiles the HIP source code 2830 for execution on a CPU 2890 and a GPU 2892.

[0263] The direct CUDA stream, which can be implemented in at least one embodiment, is depicted via a dashed line and a series of bubbles labeled A1 - A3. In at least one embodiment and as depicted using the bubble labeled A1, a CUDA compiler 2850 receives the CUDA source code 2810 and a CUDA compilation command 2848 that configures the CUDA compiler 2850 to compile the CUDA source code 2810. In at least one embodiment, the CUDA source code 2810 used in the direct CUDA stream is written in the CUDA programming language based on a programming language different from C++ (such as C, Fortran, Python, Java, etc.). In at least one embodiment and in response to the CUDA compilation command 2848, the CUDA compiler 2850 generates host-executable code 2870(1) and CUDA device-executable code 2884 (depicted using the bubble labeled A2). In at least one embodiment and as depicted using the bubble labeled A3, the host-executable code 2870(1) and the CUDA device-executable code 2884 can be executed on a CPU 2890 and a CUDA-enabled GPU 2894, respectively. In at least one embodiment, the CUDA device-executable code 2884 can include, but is not limited to, binary code. In at least one embodiment, the CUDA device-executable code 2884 includes, but is not limited to, PTX code and is further compiled into binary code for a specific target device at runtime.

[0264] An indirect CUDA stream, which can be implemented in at least one embodiment, is depicted via dotted lines and a series of bubbles labeled B1 - B6. In at least one embodiment and as depicted by the bubble labeled B1, the CUDA-HIP conversion tool 2820 receives the CUDA source code 2810. In at least one embodiment and as depicted by the bubble labeled B2, the CUDA-HIP conversion tool 2820 converts the CUDA source code 2810 into HIP source code 2830. In at least one embodiment and as depicted by the bubble labeled B3, the HIP compiler driver 2840 receives the HIP source code 2830 and determines that the target device 2846 is CUDA-enabled.

[0265] In at least one embodiment and as depicted by the bubble labeled B4, the HIP compiler driver 2840 generates HIP / NVCC compilation commands 2842 and transmits both the HIP / NVCC compilation commands 2842 and the HIP source code 2830 to the CUDA compiler 2850. In at least one embodiment and as described in more detail Figure 28B below, the HIP / NVCC compilation commands 2842 configure the CUDA compiler 2850 to compile the HIP source code 2830 using, without limitation, HIP-CUDA conversion headers and the CUDA runtime library. In at least one embodiment and in response to the HIP / NVCC compilation commands 2842, the CUDA compiler 2850 generates host-executable code 2870(1) and CUDA device-executable code 2884 (depicted by the bubble labeled B5). In at least one embodiment and as depicted by the bubble labeled B6, the host-executable code 2870(1) and the CUDA device-executable code 2884 can be executed on the CPU 2890 and the CUDA-enabled GPU 2894, respectively. In at least one embodiment, the CUDA device-executable code 2884 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device-executable code 2884 includes, but is not limited to, PTX code and is further compiled into binary code for a specific target device at runtime.

[0266] The CUDA / HCC flow that can be implemented in at least one embodiment is depicted by solid lines and a series of bubbles labeled C1 - C6. In at least one embodiment and as depicted by the bubble labeled C1, the CUDA - HIP conversion tool 2820 receives the CUDA source code 2810. In at least one embodiment and as depicted by the bubble labeled C2, the CUDA - HIP conversion tool 2820 converts the CUDA source code 2810 into HIP source code 2830. In at least one embodiment and as depicted by the bubble labeled C3, the HIP compiler driver 2840 receives the HIP source code 2830 and determines that the target device 2846 is not CUDA - enabled.

[0267] In at least one embodiment, the HIP compiler driver 2840 generates a HIP / HCC compilation command 2844 and transmits both the HIP / HCC compilation command 2844 and the HIP source code 2830 to HCC 2860 (depicted by the bubble labeled C4). In at least one embodiment and as described in more detail in conjunction with Figure 28C the HIP / HCC compilation command 2844 configures HCC 2860 to compile the HIP source code 2830 using, without limitation, the HCC headers and the HIP / HCC runtime libraries. In at least one embodiment and in response to the HIP / HCC compilation command 2844, HCC 2860 generates host - executable code 2870(2) and HCC device - executable code 2882 (depicted by the bubble labeled C5). In at least one embodiment and as depicted by the bubble labeled C6, the host - executable code 2870(2) and the HCC device - executable code 2882 can be executed on the CPU 2890 and the GPU 2892, respectively.

[0268] In at least one embodiment, after converting CUDA source code 2810 to HIP source code 2830, HIP compiler driver 2840 can then be used to generate executable code for CUDA-enabled GPU 2894 or GPU 2892 without re-executing CUDA-HIP conversion tool 2820. In at least one embodiment, CUDA-HIP conversion tool 2820 converts CUDA source code 2810 to HIP source code 2830, which is then stored in memory. In at least one embodiment, HIP compiler driver 2840 then configures HCC 2860 to generate host executable code 2870(2) and HCC device executable code 2882 based on HIP source code 2830. In at least one embodiment, HIP compiler driver 2840 then configures CUDA compiler 2850 to generate host executable code 2870(1) and CUDA device executable code 2884 based on the stored HIP source code 2830.

[0269] Figure 28B FIG. illustrates a system 2804 configured to compile and execute Figure 28A CUDA source code 2810 using CPU 2890 and CUDA-enabled GPU 2894 in accordance with at least one embodiment. In at least one embodiment, system 2804 includes, but is not limited to, CUDA source code 2810, CUDA-HIP conversion tool 2820, HIP source code 2830, HIP compiler driver 2840, CUDA compiler 2850, host executable code 2870(1), CUDA device executable code 2884, CPU 2890, and CUDA-enabled GPU 2894.

[0270] In at least one embodiment and as previously described herein in connection with Figure 28A CUDA source code 2810 includes, but is not limited to, any number (including zero) of global functions 2812, any number (including zero) of device functions 2814, any number (including zero) of host functions 2816, and any number (including zero) of host / device functions 2818. In at least one embodiment, CUDA source code 2810 also includes, but is not limited to, any number of calls to any number of functions specified in any number of CUDA APIs.

[0271] In at least one embodiment, the CUDA-HIP conversion tool 2820 converts CUDA source code 2810 into HIP source code 2830. In at least one embodiment, the CUDA-HIP conversion tool 2820 converts each kernel call in the CUDA source code 2810 from CUDA syntax to HIP syntax, and converts any number of other CUDA calls in the CUDA source code 2810 into any number of other functionally similar HIP calls.

[0272] In at least one embodiment, the HIP compiler driver 2840 determines that the target device 2846 is CUDA-enabled and generates a HIP / NVCC compilation command 2842. In at least one embodiment, the HIP compiler driver 2840 then configures the CUDA compiler 2850 to compile the HIP source code 2830 via the HIP / NVCC compilation command 2842. In at least one embodiment, as part of configuring the CUDA compiler 2850, the HIP compiler driver 2840 provides access to the HIP-CUDA conversion header 2852. In at least one embodiment, the HIP-CUDA conversion header 2852 converts any number of mechanisms (e.g., functions) specified in any number of HIP APIs into any number of mechanisms specified in any number of CUDA APIs. In at least one embodiment, the CUDA compiler 2850 uses the HIP-CUDA conversion header 2852 in conjunction with the CUDA runtime library 2854 corresponding to the CUDA runtime API 2802 to generate host-executable code 2870(1) and CUDA device-executable code 2884. In at least one embodiment, the host-executable code 2870(1) and the CUDA device-executable code 2884 can then be executed on the CPU 2890 and the CUDA-enabled GPU 2894, respectively. In at least one embodiment, the CUDA device-executable code 2884 includes, but is not limited to, binary code. In at least one embodiment, the CUDA device-executable code 2884 includes, but is not limited to, PTX code, and is further compiled into binary code for a specific target device at runtime.

[0273] Figure 28C The figure shows configured to compile and execute using the CPU 2890 and a non-CUDA-enabled GPU 2892 in accordance with at least one embodiment Figure 28AThe CUDA source code 2810 of the system 2806. In at least one embodiment, the system 2806 includes, but is not limited to, the CUDA source code 2810, the CUDA-HIP conversion tool 2820, the HIP source code 2830, the HIP compiler driver 2840, HCC 2860, the host executable code 2870(2), the HCC device executable code 2882, the CPU 2890, and the GPU 2892.

[0274] In at least one embodiment and as previously described herein in connection with Figure 28A the CUDA source code 2810 includes, but is not limited to, any number (including zero) of global functions 2812, any number (including zero) of device functions 2814, any number (including zero) of host functions 2816, and any number (including zero) of host / device functions 2818. In at least one embodiment, the CUDA source code 2810 also includes, but is not limited to, any number of calls to any number of functions specified in any number of CUDA APIs.

[0275] In at least one embodiment, the CUDA-HIP conversion tool 2820 converts the CUDA source code 2810 into the HIP source code 2830. In at least one embodiment, the CUDA-HIP conversion tool 2820 converts each kernel call in the CUDA source code 2810 from CUDA syntax to HIP syntax and converts any number of other CUDA calls in the source code 2810 into any number of other functionally similar HIP calls.

[0276] In at least one embodiment, the HIP compiler driver 2840 then determines that the target device 2846 is not CUDA-enabled and generates the HIP / HCC compilation command 2844. In at least one embodiment, the HIP compiler driver 2840 then configures HCC 2860 to execute the HIP / HCC compilation command 2844 to compile the HIP source code 2830. In at least one embodiment, the HIP / HCC compilation command 2844 configures HCC 2860 to generate the host executable code 2870(2) and the HCC device executable code 2882 using, without limitation, the HIP / HCC runtime library 2858 and the HCC headers 2856. In at least one embodiment, the HIP / HCC runtime library 2858 corresponds to the HIP runtime API 2832. In at least one embodiment, the HCC headers 2856 include, but are not limited to, any number and type of interoperability mechanisms for HIP and HCC. In at least one embodiment, the host executable code 2870(2) and the HCC device executable code 2882 can be executed on the CPU 2890 and the GPU 2892, respectively.

[0277] Figure 29 The figure shows an exemplary kernel transformed by the CUDA-HIP transformation tool 2820 according to at least one embodiment. In at least one embodiment, the CUDA source code 2810 divides the overall problem that a given kernel is designed to solve into relatively coarse sub-problems that can be solved independently using thread blocks. In at least one embodiment, each thread block includes, but is not limited to, any number of threads. In at least one embodiment, each sub-problem is divided into relatively fine-grained fragments that can be solved cooperatively in parallel by the threads within a thread block. In at least one embodiment, the threads within a thread block can cooperate by sharing data via shared memory and by synchronizing their execution to coordinate memory access. Figure 28C In at least one embodiment, the CUDA source code 2810 organizes the thread blocks associated with a given kernel into a one-dimensional, two-dimensional, or three-dimensional grid of thread blocks. In at least one embodiment, each thread block includes, but is not limited to, any number of threads, and the grid includes, but is not limited to, any number of thread blocks.

[0278] In at least one embodiment, a kernel is a function in device code that is qualified using the "_global_" declaration specifier. In at least one embodiment, the dimensions of the grid of the kernel to be executed for a given kernel call and associated stream are specified using CUDA kernel launch syntax 2910. In at least one embodiment, the CUDA kernel launch syntax 2910 is specified as "KernelName<<<GridSize,BlockSize,SharedMemorySize,Stream>>>(KernelArguments);". In at least one embodiment, the execution configuration syntax is the "<<<...>>>" construct inserted between the kernel name ("KernelName") and the list of kernel arguments in parentheses ("KernelArguments"). In at least one embodiment, the CUDA kernel launch syntax 2910 includes, but is not limited to, CUDA launch function syntax rather than execution configuration syntax.

[0279]

[0280] ​In at least one embodiment, "GridSize" is of type dim3 and specifies the dimensions and size of the grid. In at least one embodiment, the dim3 type is a CUDA-defined structure that includes, but is not limited to, unsigned integers x, y, and z. In at least one embodiment, if z is not specified, then z defaults to 1. In at least one embodiment, if y is not specified, then y defaults to 1. In at least one embodiment, the number of thread blocks in the grid is equal to the product of GridSize.x, GridSize.y, and GridSize.z. In at least one embodiment, "BlockSize" is of type dim3 and specifies the dimensions and size of each thread block. In at least one embodiment, the number of threads per thread block is equal to the product of BlockSize.x, BlockSize.y, and BlockSize.z. In at least one embodiment, each thread executing the kernel is given a unique thread ID that is accessible within the kernel via a built-in variable (e.g., "threadIdx").

[0281] In at least one embodiment and with respect to CUDA kernel launch syntax 2910, "SharedMemorySize" is an optional parameter that specifies the number of bytes of shared memory dynamically allocated per thread block in addition to statically allocated memory for a given kernel call. In at least one embodiment and with respect to CUDA kernel launch syntax 2910, SharedMemorySize defaults to zero. In at least one embodiment and with respect to CUDA kernel launch syntax 2910, "Stream" is an optional parameter that specifies the associated stream and defaults to zero to specify the default stream. In at least one embodiment, a stream is an ordered sequence of commands (possibly issued by different host threads). In at least one embodiment, different streams can execute commands out of order or concurrently with respect to each other.

[0282] In at least one embodiment, the CUDA source code 2810 includes, but is not limited to, a kernel definition for an exemplary kernel “MatAdd” and a main function. In at least one embodiment, the main function is host code that executes on the host and includes, but is not limited to, a kernel call that causes the kernel MatAdd to execute on the device. In at least one embodiment and as shown, the kernel MatAdd adds two matrices A and B of size NxN, where N is a positive integer, and stores the result in matrix C. In at least one embodiment, the main function defines the threadsPerBlock variable as 16x16 and the numBlocks variable as N / 16 x N / 16. In at least one embodiment, the main function then specifies the kernel call “MatAdd<<<numBlocks,threadsPerBlock>>>(A,B,C);”. In at least one embodiment and according to the CUDA kernel launch syntax 2910, the kernel MatAdd executes using a thread block grid with dimensions N / 16 x N / 16, where each thread block has dimensions 16x16. In at least one embodiment, each thread block includes 256 threads, the grid is created with enough blocks so that each matrix element has one thread, and each thread in such a grid executes the kernel MatAdd to perform one pairwise addition.

[0283] In at least one embodiment, when converting the CUDA source code 2810 to HIP source code 2830, the CUDA-HIP conversion tool 2820 converts each kernel call in the CUDA source code 2810 from the CUDA kernel launch syntax 2910 to the HIP kernel launch syntax 2920, and converts any number of other CUDA calls in the source code 2810 to any number of other functionally similar HIP calls. In at least one embodiment, the HIP kernel launch syntax 2920 is specified as “hipLaunchKernelGGL(KernelName,GridSize,BlockSize,SharedMemorySize,Stream,KernelArguments);”. In at least one embodiment, each of KernelName, GridSize, BlockSize, ShareMemorySize, Stream, and KernelArguments in the HIP kernel launch syntax 2920 has the same meaning as in the CUDA kernel launch syntax 2910 (as previously described herein). In at least one embodiment, the parameters SharedMemorySize and Stream are required in the HIP kernel launch syntax 2920 and are optional in the CUDA kernel launch syntax 2910.

[0284] In at least one embodiment, in addition to the kernel call that causes the kernel MatAdd to execute on the device, Figure 29 the portion of the HIP source code 2830 depicted in Figure 29 is the same as the portion of the CUDA source code 2810 depicted in

[0285] Figure 30 More specifically illustrated is a non-CUDA enabled GPU 2892 in accordance with at least one embodiment. Figure 28C In at least one embodiment, the GPU 2892 is developed by Santa Clara AMD Corporation. In at least one embodiment, the GPU 2892 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, the GPU 2892 is configured to perform graphics pipeline operations such as draw commands, pixel operations, geometric calculations, and other operations associated with reproducing an image to a display. In at least one embodiment, the GPU 2892 is configured to perform operations that are not related to graphics. In at least one embodiment, the GPU 2892 is configured to perform both operations related to graphics and operations not related to graphics. In at least one embodiment, the GPU 2892 can be configured to execute the device code included in the HIP source code 2830.

[0286] In at least one embodiment, GPU 2892 includes, but is not limited to, any number of programmable processing units 3020, a command processor 3010, an L2 cache 3022, a memory controller 3070, DMA engines 3080(1), a system memory controller 3082, DMA engines 3080(2), and a GPU controller 3084. In at least one embodiment, each programmable processing unit 3020 includes, but is not limited to, a workload manager 3030 and any number of compute units 3040. In at least one embodiment, the command processor 3010 reads commands from one or more command queues (not shown) and distributes the commands to the workload manager 3030. In at least one embodiment, for each programmable processing unit 3020, the associated workload manager 3030 distributes work to the compute units 3040 included in the programmable processing unit 3020. In at least one embodiment, each compute unit 3040 can execute any number of thread blocks, but each thread block is executed on a single compute unit 3040. In at least one embodiment, a workgroup is a type of thread block.

[0287] In at least one embodiment, each compute unit 3040 includes, but is not limited to, any number of SIMD units 3050 and a shared memory 3060. In at least one embodiment, each SIMD unit 3050 implements a SIMD architecture and is configured to perform operations in parallel. In at least one embodiment, each SIMD unit 3050 includes, but is not limited to, a vector ALU 3052 and a vector register file 3054. In at least one embodiment, each SIMD unit 3050 executes different warps. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in the warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predicates can be used to disable one or more threads in a warp. In at least one embodiment, a lane is a type of thread. In at least one embodiment, a work item is a type of thread. In at least one embodiment, a wavefront is a type of warp. In at least one embodiment, different wavefronts in a thread block can synchronize together and communicate via the shared memory 3060.

[0288] In at least one embodiment, the programmable processing unit 3020 is referred to as a "shader engine". In at least one embodiment, each programmable processing unit 3020 includes, but is not limited to, any number of dedicated graphics hardware in addition to the compute units 3040. In at least one embodiment, each programmable processing unit 3020 includes, but is not limited to, any number (including zero) of geometry processors, any number (including zero) of rasterizers, any number (including zero) of rendering backends, a workload manager 3030, and any number of compute units 3040.

[0289] In at least one embodiment, the compute units 3040 share the L2 cache 3022. In at least one embodiment, the L2 cache 3022 is partitioned. In at least one embodiment, the GPU memory 3090 is accessible by all compute units 3040 in the GPU 2892. In at least one embodiment, the memory controller 3070 and the system memory controller 3082 facilitate data transfer between the GPU 2892 and the host, and the DMA engine 3080(1) allows for asynchronous memory transfer between the GPU 2892 and such a host. In at least one embodiment, the memory controller 3070 and the GPU controller 3084 facilitate data transfer between the GPU 2892 and other GPU 2892s, and the DMA engine 3080(2) allows for asynchronous memory transfer between the GPU 2892 and other GPU 2892s.

[0290] In at least one embodiment, the GPU 2892 includes, but is not limited to, any number and type of system interconnects that facilitate data and control transfer across any number and type of directly or indirectly linked components, which may be internal or external to the GPU 2892. In at least one embodiment, the GPU 2892 includes, but is not limited to, any number and type of I / O interfaces (e.g., PCIe) coupled to any number and type of peripheral devices. In at least one embodiment, the GPU 2892 may include, but is not limited to, any number (including zero) of display engines and any number (including zero) of multimedia engines. In at least one embodiment, the GPU 2892 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers (e.g., the memory controller 3070 and the system memory controller 3082) and memory devices (e.g., the shared memory 3060) that may be dedicated to one component or shared among multiple components. In at least one embodiment, the GPU 2892 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., the L2 cache 3022), each of which may be dedicated to any number of components (e.g., the SIMD units 3050, the compute units 3040, and the programmable processing unit 3020) or shared among these components.

[0291] Figure 31 The figure shows how threads of an exemplary CUDA grid 3120 are mapped in accordance with at least one embodiment as Figure 30different computing units 3040. In at least one embodiment and for illustrative purposes only, the grid 3120 has a GridSize of BX x BY x 1 and a BlockSize of TX x TY x 1. In at least one embodiment, the grid 3120 thus includes, but is not limited to, (BX*BY) thread blocks 3130, and each thread block 3130 includes, but is not limited to, (TX*TY) threads 3140. The threads 3140 are depicted as curvilinear arrows in Figure 31 as shown.

[0292] In at least one embodiment, the grid 3120 is mapped to a programmable processing unit 3020(1) that includes, but is not limited to, computing units 3040(1)-3040(C). In at least one embodiment and as shown, (BJ*BY) thread blocks 3130 are mapped to computing unit 3040(1), and the remaining thread blocks 3130 are mapped to computing unit 3040(2). In at least one embodiment, each thread block 3130 may include, but is not limited to, any number of warps, and each warp is mapped to Figure 30 a different SIMD unit 3050.

[0293] In at least one embodiment, the warps in a given thread block 3130 can synchronize together and communicate through a shared memory 3060 included in the associated computing unit 3040. For example and in at least one embodiment, the warps in thread block 3130(BJ,1) can synchronize together and communicate through shared memory 3060(1). For example and in at least one embodiment, the warps in thread block 3130(BJ+1,1) can synchronize together and communicate through shared memory 3060(2).

[0294] Task Graph and Executable Graph

[0295] In at least one embodiment, the use of dedicated computing resources and / or an array of parallel computing resources provides an opportunity to execute complex workloads more efficiently and quickly. In at least one embodiment, the execution of many image processing, machine learning, and related workloads can generally be accelerated by using one or more dedicated computing resources operating in parallel and / or two or more computing resources. In at least one embodiment, a host processor, such as a host central processing unit (“CPU”), is responsible for preparing one or more tasks and distributing them to one or more other computing resources, and then retrieving the results of one or more tasks from one or more other computing resources. In at least one embodiment, as long as one or more other computing resources can complete one or more tasks faster than the host CPU could directly implement one or more tasks, then using one or more other computing resources to implement one or more tasks is generally considered advantageous.

[0296] In at least one embodiment, offloading one or more tasks to one or more other computing resources can be overhead-intensive. In at least one embodiment, among other possibilities, the overhead can include selecting one or more other computing resources, configuring and / or transmitting executable instructions to one or more other computing resources, transmitting data to one or more other computing resources, and retrieving results from one or more other computing resources. In at least one embodiment, if the overhead is not carefully managed, the cost of the overhead can significantly reduce the efficiency of using one or more other computing resources to implement computational tasks.

[0297] In at least one embodiment, a method for reducing overhead is to consider multiple tasks as an integrated whole, rather than considering each of the multiple tasks individually. In at least one embodiment, there are dependencies between the various different tasks in a workload, where some tasks generate data required by other tasks, some tasks need to be completed before other tasks can start, some tasks need to share exclusive use resources with other tasks, and / or some tasks need to execute in synchronization with other tasks. In at least one embodiment, some computing resources are more suitable for certain tasks than others. In at least one embodiment, techniques that consider the dependencies between tasks and / or consider which combinations of computing resources are most computationally efficient generally can reduce the overhead and computational time associated with executing multiple tasks across multiple computing resources, as compared to managing each of the multiple tasks in isolation.

[0298] In at least one embodiment, multiple tasks can be organized into an integrated whole by using a task graph. In at least one embodiment, in a task graph, a workload including multiple tasks is organized as a directed graph, where each node corresponds to a task to be implemented, and each directed edge between two nodes corresponds to a data dependency, an execution dependency, or some other dependency between the two nodes. In at least one embodiment, a dependency can indicate when the task of one node must be completed before the task of another node can start. In at least one embodiment, a dependency can indicate when one node must wait for data from another node before the node can start and / or continue its task. In at least one embodiment, once the task graph is ready, it is converted into an executable version of the task graph ("executable graph"). In at least one embodiment, instantiation can be used to generate the executable graph from the task graph. In at least one embodiment, since converting the task graph into an executable graph provides access to the entire task graph, various different optimizations can be implemented, which can reduce the overall execution time of the workload.

[0299] In at least one embodiment, the executable graph can be used multiple times so that the same workload can be implemented by computing resources without having to regenerate it from the task graph. One drawback of using the executable graph to implement a workload in at least one embodiment is that the executable graph is limited to the same workload of the task graph from which it is generated.

[0300] Modify the executable graph to implement the workload of a new task graph

[0301] Figure 32 The figure shows a flowchart according to at least one embodiment. Although the method steps 3210 - 3260 of method 3200 are described in connection with Figures 1 - 31 the system and techniques of, it should be understood that any system configured to implement steps 3210 - 3260 is consistent with at least one embodiment. In at least one embodiment, one or more of the steps 3210 - 3260 of method 3200 can be at least partially implemented as executable code stored on a non - transitory tangible machine - readable medium, which, when run by one or more processors such as Figures 1 - 31 any processor or computing device, can cause the one or more processors to implement one or more of the steps 3210 - 3260.

[0302] In at least one embodiment, method 3200 begins at step 3210, where a task graph for a workload is constructed. In at least one embodiment, the task graph can be explicitly constructed by making one or more application programming interface (API) or library function calls. The API or library function calls are used to create each node of the task graph and associate the nodes with computational tasks. In at least one embodiment, the computational tasks can correspond to kernel tasks to be implemented by computing resources, host tasks corresponding to callback functions implemented on a host CPU used to launch an executable graph, memory setup tasks, memory copy tasks, and so on. In at least one embodiment, the same and / or different API or library function calls can be used to add directed edges to the task graph to specify data, execution, and / or other dependencies between two tasks. In at least one embodiment, a directed edge can be indirectly added to the task graph based on a data dependency where data generated by one task is source data for another task. In at least one embodiment, alternatively, the task graph can be indirectly constructed by recording a task flow. In at least one embodiment, the task flow is recorded by creating a task flow for recording the task graph and / or using one or more API or library function calls to convert an existing task flow into a task flow for recording the task graph, making one or more additional API or library function calls that specify tasks to be added to the task flow, and then using one or more third API or library function calls to complete the conversion of the task flow to the task graph. In at least one embodiment, the task graph is a non-executable task graph.

[0303] In at least one embodiment, when each task is added as a node to the task graph, an order of creation is associated with each node. In at least one embodiment, the order of creation associated with each node corresponds to an ordinal number that indicates whether each node is the first, second, third, etc. node to be created and added to the task graph. In at least one embodiment, the order of creation of each node can correspond to a count value of a counter that is initialized (e.g., to 0 or 1) when the task graph is initially created and incremented whenever a new node is added to the task graph.

[0304] Figure 33A -Figure C illustrates an exemplary task graph according to at least one embodiment. In at least one embodiment, task graphs 3310, 3320, and / or 3330 can be consistent with CUDA for specifying one or more CUDA tasks. Figure One In at least one embodiment, task graphs 3310, 3320, and / or 3330 can be consistent with HIP for specifying one or more HIP tasks. Figure OneRefer to. As shown in FIG. 33, the task graph 3310 is a directed graph including eight nodes / tasks shown as having labels A - G and a terminating circle. The task graph 3300 starts with task A. The directed edges between task A and tasks B and F respectively, shown as arrows, correspond to the dependencies between task A and tasks B and F. More specifically, the directed edge between task A and task B indicates that task A must be completed before task B can start, and the directed edge between task A and task F indicates that task A must be completed before task F can start. Since there is no directed edge between tasks B and F, the execution of tasks B and F is independent of each other, which means that tasks B and F can be executed in any order relative to each other, including: task B is executed before task F, task F is executed before task B, tasks B and F are at least partially concurrent, tasks B and F are executed in an interleaved manner relative to each other, and / or any combination thereof. Additional task dependencies in the task graph 3300 include: task B must be completed before task C can start, task B must be completed before task D can start, both tasks C and D must be completed before task E can start, and task F must be completed before task G can start. In addition, the task graph 3300 indicates that tasks E and G must be completed before reaching the task end to complete the workload of the task graph 3300. In at least one embodiment, control is returned to the host CPU by the task end. In at least one embodiment, when reaching the task end, another task (not shown) can be started.

[0305] In at least one embodiment and not shown in Figure 33A each of tasks A - G can correspond to a sub - graph including two or more tasks / nodes. In at least one embodiment, the sub - graph can correspond to one or more nodes and / or dependencies created by API and / or library function calls implementing a multi - task workload. In at least one embodiment, Figure 33B FIG. shows a sub - graph 3320 corresponding to node D. The sub - graph 3320 includes four tasks labeled D1 - D4. The sub - graph 3320 further shows that task D1 must be completed before tasks D2 and D3 can start, and both tasks D2 and D3 must be completed before task D4 can start. In addition, since there is no dependency between tasks D2 and D3, tasks D2 and D3 can be executed in any order relative to each other. The sub - graph 3320 further indicates that task D4 must be completed before task D is completed.

[0306] Back reference Figure 32, in at least one embodiment and at step 3220, an executable graph is generated from a task graph. In at least one embodiment, step 3220 is implemented via instantiation. In at least one embodiment, step 3220 is implemented via compilation. In at least one embodiment, the generation of the executable graph includes assigning each task / node of the task graph to a corresponding computing resource and / or a subset of corresponding computing resources, such as Figures 1 - 31 any computing resource, graphics processor, fragment processor, core, execution unit, cluster, multiprocessor, computing unit, and / or streaming multiprocessor, etc. In at least one embodiment, the corresponding computing resource can be selected based on the task to be implemented, such as whether the task is a kernel task, host task, memory setup task, memory copy task, etc. In at least one embodiment, the corresponding computing resource can additionally and / or alternatively be selected based on the computation to be performed by the task, the number of resources such as threads, warps, etc. required by the task, and / or other criteria. In at least one embodiment, the corresponding computing resource can additionally and / or alternatively be selected based on the dependencies between the tasks in the task graph.

[0307] In at least one embodiment, creation order information, such as an ordinal number, associated with each task / node in the task graph is associated with the corresponding task / node in the executable graph.

[0308] In at least one embodiment, when generating an executable graph from a task graph, one or more optimizations may be implemented. In at least one embodiment, the optimizations may include rearranging tasks to exploit parallelism between tasks. In at least one embodiment, rearranging tasks may include one or more of the following: determining how much computing resource to use, transferring tasks between computing resources, and so on. In at least one embodiment, the optimizations may include redistributing tasks to computing resources, for example by assigning a task having data dependencies for another task to the same computing resource as the other task, so as to reduce the overhead caused by data dependencies. In at least one embodiment, redistributing tasks may reduce the execution time because the result of another task is available for a task without having to transfer the result to a different computing resource for use by the task. In at least one embodiment, redistributing tasks may reduce the execution time because the completion of another task can be detected more quickly, allowing a task to start more quickly after the completion of another task. In at least one embodiment, tasks may be assigned to computing resources to reduce the computing cost or the latency of transferring the result of another task to the computing resource assigned to the previous task. In at least one embodiment, the optimizations may include rearranging tasks to better balance and / or optimize the capabilities of computing resources, such as the number of threads, warps, compute thread arrays, cores, streaming multiprocessors, etc. available to each computing resource. In at least one embodiment, the task graph may be converted to an executable graph by making one or more API or library function calls.

[0309] In at least one embodiment, at step 3230, the executable graph is launched one or more times so that the workload is implemented. In at least one embodiment, once the executable graph is ready, the executable graph may be launched to cause one or more computing resources to implement the workload of the task graph. In at least one embodiment, the executable graph may be launched using one or more API or library function calls. In at least one embodiment, for each task / node in the executable graph, launching the executable graph may include one or more of the following: sending the executable code corresponding to each task to the associated computing resource, sending a configuration to each associated computing resource based on one or more attributes of each task, transferring (if necessary) the data for each task to each associated computing resource, and / or when each associated computing resource completes each task, transferring the result of each task back to the host CPU or one or more computing resources that need the result of each task as data for the next corresponding task to be implemented. In at least one embodiment, when all tasks in the executable graph are executed by the associated computing resources, the workload of the task graph has been implemented. In at least one embodiment, the executable graph may be launched as many times as implementing the same workload.

[0310] In at least one embodiment, at step 3240, another task graph is constructed for another workload. In at least one embodiment, step 3240 is substantially similar to step 3210, except for constructing another task graph. In at least one embodiment, the workload of the other task graph may be different from the workload of the task graph constructed during process 3210. In at least one embodiment, the other task graph is a non-executable task graph.

[0311] In at least one embodiment, at step 3250, the other task graph is applied to the executable graph. In at least one embodiment, the goal of step 3250 is to modify the executable graph such that the executable graph can be used by one or more computing resources to implement the workload of the other task graph instead of implementing the workload of the task graph constructed during step 3210. In at least one embodiment, when step 3250 successfully applies the other task graph to the executable graph, the executable graph can be modified without incurring computational costs and latencies such as converting the other task graph into a new executable graph by using step 3220. In at least one embodiment, one or more API and / or library function calls can be used to apply the other task graph to the executable graph.

[0312] In at least one embodiment, step 3250 may be helpful when creating any task graph from a task flow, and / or the code used to create any task graph includes one or more API and / or library function calls such that the topology of the task graph and / or the parameters of the tasks in the task graph are generally not known to the developer who writes the code to construct the task graph. In at least one embodiment, the developer may not know if any API and / or library function calls create new task nodes and / or create new dependencies between task nodes. In at least one embodiment, the developer may have a simpler understanding of the task graph, such as shown in task graph 3310, even though one or more nodes of the task graph may have embedded subgraphs included therein, such as shown in task graph 3330, where node D has been replaced by embedded subgraph 3320. In at least one embodiment, this may be further exacerbated when one or more API and / or library function calls also make one or more additional API or library function calls.

[0313] In at least one embodiment, the developer may not be able to know if two task graphs have the same topology and / or if the parameters of the tasks in the two task graphs allow one task graph to be applied to the executable graph generated from the other task graph. In at least one embodiment, step 3460 solves these problems by using method 3400 as shown in Figure 34 as shown, Figure 34The figure shows a flowchart according to at least one embodiment. In at least one embodiment and within the context of step 3460, the new task graph corresponds to the task graph constructed during step 3440, and the executable graph corresponds to the executable graph generated during step 3420. Although the method steps 3410 - 3460 of method 3400 are described in connection with Figures 1 - 31 the systems and techniques of, it should be understood that any system configured to implement steps 3410 - 3460 is consistent with at least one embodiment. In at least one embodiment, one or more of the steps 3410 - 3460 of method 3400 may be at least partially implemented as executable code stored on a non - transitory tangible machine - readable medium, which when run by one or more processors such as Figures 1 - 31 any processor or computing device of, may cause the one or more processors to implement one or more of steps 3410 - 3460.

[0314] In at least one embodiment, method 3400 begins at step 3410, where the topology of the new task graph is compared with the topology of the executable graph. In at least one embodiment, the topologies of the new task graph and the executable graph may be compared by verifying that the new task graph and the executable graph have the same node / task arrangement and the same directed edge / dependency arrangement between the nodes / tasks. In at least one embodiment, this can be a complex comparison because two graphs may have the same topology even if one topology has one or more reflections, permutations, etc. between various different nodes relative to the other topology. In at least one embodiment consistent with the task graph 3310 from Figure 33A the task graph in which nodes C and D are swapped is topologically equivalent to task graph 3310.

[0315] However, in at least one embodiment, the comparison of the topologies can be simplified by leveraging the creation order information associated with each node when constructing the new task graph and the creation order information associated with each node when generating the executable graph.

[0316] Figure 35 The figure shows a flowchart according to at least one embodiment. In the context of step 3410, the new task graph and the executable graph correspond to the two graphs being compared. Although the method steps 3510 - 3570 of method 3500 are described in connection with Figures 1 - 31 the systems and techniques of, it should be understood that any system configured to implement steps 3510 - 3570 is consistent with at least one embodiment. In at least one embodiment, one or more of the steps 3510 - 3570 of method 3500 may be at least partially implemented as executable code stored on a non - transitory tangible machine - readable medium, which when run by one or more processors such as Figures 1 - 31When run by one or more processors, such as any processor or computing device, the one or more processors can cause one or more of steps 3510 - 3570 to be implemented.

[0317] In at least one embodiment, method 3500 begins at step 3510, where it is determined whether two graphs being compared have the same number of nodes. In at least one embodiment, when the two graphs have the same number of nodes, the comparison of the two graphs continues with step 3520. In at least one embodiment, when the two graphs do not have the same number of nodes, the two graphs have different topologies, and the result is returned using step 3570.

[0318] In at least one embodiment, at step 3520, the types of each corresponding node of the two graphs are compared. In at least one embodiment, nodes from one graph and nodes from the other graph are corresponding to each other when they are both associated with the same creation or ordinal. In at least one embodiment, once the corresponding nodes are identified, the type of a node is compared with the type of the corresponding node to determine whether those types are the same type. In at least one embodiment, when a node is a kernel node for implementing a particular type of kernel task, the corresponding node must be a kernel node for implementing the same particular type of kernel task for those types to be the same type. In at least one embodiment, when a node corresponds to a memory setup task for setting a one - dimensional memory block of a particular memory type to a fill value, the corresponding node must correspond to a memory setup task for setting a one - dimensional memory block of the same particular memory type to a fill value for those types to be the same type. In at least one embodiment, similar tests can be performed on nodes corresponding to memory setup tasks, memory copy tasks, host tasks, etc. In at least one embodiment, if a node in a graph does not have a corresponding node in the other graph, for example, because no corresponding node with the same associated creation order is found in the other graph, then there are no corresponding nodes of the same type. In at least one embodiment, as soon as it is found that any node and the corresponding node do not have the same type, step 3520 can end without comparing the types of any additional nodes.

[0319] In at least one embodiment, at step 3530, it is determined whether each node and the corresponding node are of the same type. In at least one embodiment, when any node in one graph and the corresponding node in the other graph are not of the same type, the graphs have different topologies, and the result is returned using step 3570. In at least one embodiment, when each node in one graph has a corresponding node of the same type in the other graph and each node in the other graph has a corresponding node of the same type in the graph, the graphs proceed to further comparison with step 3540.

[0320] In at least one embodiment, at step 3540, the dependencies of each corresponding node of these graphs are compared. In at least one embodiment, when the list of dependencies of a node in one graph includes a node having the same associated creation order information as the list of dependencies of the corresponding node in another graph, the node has the same dependencies as the corresponding node. In at least one embodiment, if the node associated with creation order ordinal 15 has a dependency on the nodes associated with creation order ordinals 6 and 11, and the corresponding node having creation ordinal 15 has a dependency on the corresponding nodes associated with creation order ordinals 6 and 11, then the node and the corresponding node have the same dependencies. In at least one embodiment, as soon as it is found that any node and the corresponding node do not have the same dependencies, step 3540 can end without comparing the dependencies of any additional nodes.

[0321] In at least one embodiment, at step 3550, it is determined that each node in one graph and the corresponding node in another graph have the same list of dependencies. In at least one embodiment, when any node in one graph and the corresponding node in another graph do not have the same list of dependencies, the graphs have different topologies, and the result is returned using step 3570. In at least one embodiment, when each node in one graph has a corresponding node in the other graph with the same list of dependencies, the graphs are considered to have the same topology, and the result is returned using step 3560.

[0322] In at least one embodiment, at step 3560, a result indicating that the two graphs have the same topology is returned. In at least one embodiment, an indication that the two graphs have the same topology can be provided as a return value to the function that invokes to implement method 3500. In at least one embodiment, as soon as step 3560 is completed, method 3500 ends.

[0323] In at least one embodiment, at step 3570, a result indicating that the two graphs have different topologies is returned. In at least one embodiment, an indication that the two graphs have different topologies can be provided as a return value to the function that invokes to implement method 3500. In at least one embodiment, as soon as step 3570 is completed, method 3500 ends.

[0324] In at least one embodiment, the steps of method 3500 can be implemented in an order different from that Figure 35 shown. In at least one embodiment, steps 3540 and 3550 can be implemented before steps 3520 and 3530. In at least one embodiment, steps 3520 and 3540 can be implemented concurrently by comparing both the type and dependencies of each node one node at a time.

[0325] Back referenceFigure 34 , in at least one embodiment and at step 3420, determine whether the topologies of the new task graph and the executable graph are the same topology based on the comparison performed during step 3410. In at least one embodiment, the result returned by step 3560 or 3570 can be used to determine whether the topologies of the new task graph and the executable graph are the same or different topologies. In at least one embodiment, when the topologies of the new task graph and the executable graph are the same topology, start applying the parameters of the new task graph to the executable graph at step 3430. In at least one embodiment, when the topologies of the new task graph and the executable graph are different topologies, use the failure returned by step 3460 in applying the new task graph to the executable graph.

[0326] In at least one embodiment, at step 3430, apply the parameters of the new task graph to the executable task. In at least one embodiment, in order to modify the executable graph such that the executable graph can implement the workload of the new task graph, modify the executable graph based on the parameters of the new task. In at least one embodiment, the modifications that can be made to the executable graph may interfere with and / or interrupt one or more of the optimizations performed during step 3220. In at least one embodiment, the modifications that can be made to the executable graph do not interfere with the optimizations performed during step 3220. In at least one embodiment, the topology comparison of step 3410 has confirmed that the tasks and dependencies of the new task graph and the executable graph have corresponding nodes of the same type and the same dependencies. In at least one embodiment, applying the parameters of the new task graph to the executable graph is restricted such that for certain types of tasks, any modification to the node parameters in the executable graph when applying the parameters of the new task graph is restricted to certain types of parameters. In at least one embodiment, for certain types of tasks, the allowed modifications to the executable graph can be restricted to certain types of parameters that do not interfere with the optimizations performed during step 3220.

[0327] In at least one embodiment, a node corresponding to a host task that triggers a callback function on the host CPU in connection with an associated computing resource may include modifiable parameters, the modifiable parameters including an identifier such as a pointer to the callback function on the host CPU and / or one or more of the data that can be accessed by the associated computing device to be passed as one or more arguments to the callback function. In at least one embodiment, it is possible to modify the callback function and / or the argument parameters because these modifications do not modify the function being performed by the associated computing resource to implement the host task or interfere with any optimizations typically applied to the host task during the generation of the executable graph.

[0328] In at least one embodiment, a node corresponding to a memory setup task in which an associated computing resource sets a memory block accessible to the associated computing resource to a specified fill value may have modifiable parameters including one or more of the following: a starting position of the memory block, a size of the memory block, and / or a fill value to be set at a memory location in the memory block. In at least one embodiment, the size of the memory block may correspond to the number of memory words of a one-dimensional memory block, the number of rows and columns of a two-dimensional memory block, and so on. In at least one embodiment, modifying the position, size, and / or fill value parameters does not modify the functionality being implemented by the associated computing resource to perform the memory setup task and / or interfere with any optimizations typically applied to the memory setup task during generation of an executable graph. In at least one embodiment, it may not be possible to modify, for example, the block type of the memory block from a one-dimensional block to a two-dimensional block or vice versa, and / or modify the type of memory being set between memories such as global, local, texture, constant, shared, register, etc., because these types of modifications may involve different implementations of the memory setup task.

[0329] In at least one embodiment, a node corresponding to a memory copy task in which an associated computing resource moves the content of a memory block accessible to the associated computing resource to another memory block accessible to the associated computing resource may have modifiable parameters including one or more of the following: a source position of the memory block, a destination position of the other memory block, and / or a size of the amount of memory to be copied. In at least one embodiment, the size of the amount of memory to be copied may correspond to the number of memory words of a one-dimensional memory block, the number of rows and columns of a two-dimensional memory block, and so on. In at least one embodiment, modifying the source position, destination position, and / or size parameters does not modify the functionality being implemented by the associated computing resource to perform the memory copy task or interfere with any optimizations typically applied to the memory copy task during generation of an executable graph. In at least one embodiment, it may not be possible to modify, for example, the block type of the memory being copied from a one-dimensional block to a two-dimensional block or vice versa, and / or modify the type of memory in either memory block between memories such as global, local, texture, constant, shared, register, etc., because these types of modifications may involve different implementations of the memory copy task.

[0330] In at least one embodiment, a node corresponding to a kernel task that performs a computing task involving associated computing resources may have modifiable parameters, where the modifiable parameters include a pointer to data accessible by the associated computing resources that is a parameter of the computing task and / or the number of threads to be used. In at least one embodiment, modification of these parameters is possible because they do not modify the functions being performed by the associated computing resources to perform the kernel task or interfere with any optimizations typically applied to the kernel task during generation of the executable graph. In at least one embodiment, it may not be possible to modify the starting address of the computing task, the computing task to be performed, the pointer to the next task to be performed, etc., because these modify the underlying computing task and / or may interfere with dependencies between tasks in the executable graph.

[0331] In at least one embodiment, at step 3440, it is determined whether the parameters of the new task graph have been successfully applied to the executable graph in step 3430. In at least one embodiment, when each parameter of the new task graph has been successfully applied to the executable graph, success is returned using step 3450. In at least one embodiment, when any parameter of the new task graph cannot be successfully applied to the executable graph, failure is returned using step 3460.

[0332] In at least one embodiment, at step 3450, a result indicating that the new task graph has been successfully applied to the executable graph is returned. In at least one embodiment, success in applying the new task graph to the executable graph may be provided as a return value to the function called to perform method 3400. In at least one embodiment, upon completion of step 3450, method 3400 ends.

[0333] In at least one embodiment, at step 3460, a result indicating that the new task graph has not been successfully applied to the executable graph is returned. In at least one embodiment, failure in applying the new task graph to the executable graph may be provided as a return value to the function called to perform method 3400. In at least one embodiment, upon completion of step 3460, method 3400 ends.

[0334] In at least one embodiment, the steps of method 3400 may be in a manner consistent with Figure 34Implement in different orders as shown. In at least one embodiment, steps 3410 and 3430 can be implemented concurrently by comparing the topologies of the new task graph and the executable graph and applying the parameters of the new task graph to the executable graph through the nodes of the new task graph and the executable graph in the same pass. In at least one embodiment, any failures detected and / or encountered during the comparison of the topologies of the new task graph and the executable graph and / or in applying the parameters of the new task graph to the executable graph can be reported to a log, a user, a user interface, etc. In at least one embodiment, the failure can be reported via an alert, a warning, an error, etc. In at least one embodiment, the failure can be reported with an explanation of the source of the failure. In at least one embodiment, reporting the failure can help the user change how to construct the task graph to avoid the recurrence of the failure.

[0335] Back reference Figure 32 , in at least one embodiment and at step 3260, determine whether another task graph has been successfully applied to the executable graph. In at least one embodiment, the success of each API and / or library function call that applies another task graph to the executable graph made during step 3250 can be determined based on the status and / or error code returned by each API and / or library function call. In at least one embodiment, the status and / or error code can correspond to the success returned by step 3450 and / or the failure returned by step 3460. In at least one embodiment, when another task graph has been successfully applied to the executable graph, the executable can be started by returning to step 3230 Figure One one or more times, where the executable graph is used to cause one or more computing resources to implement another workload of the other task graph for each start of the executable graph. In at least one embodiment, when another task graph has not been successfully applied to the executable graph, control returns to step 3220, where a new executable graph is generated from the other task graph. In at least one embodiment, the new executable graph can then be started using step 3230 to cause one or more computing resources to implement the workload of the other task graph.

[0336] In at least one embodiment, even though the executable graph obtained by applying another task graph to an executable graph using step 3250 may result in slower performance of another workload compared to the case where the executable graph is generated from another task graph using step 3220, the time savings from avoiding the overhead of generating another executable graph from another task graph using step 3220 may also result in a shorter overall processing time for implementing another workload. In at least one embodiment, when applying another task graph to an executable graph using step 3250 interferes with and / or interrupts one or more optimizations applied to the executable graph when creating the executable graph, the executable graph obtained by applying another task graph to the executable graph using step 3250 may result in slower performance of another workload.

[0337] In summary, the disclosed techniques can be used to allow an executable graph to implement a workload associated with a new task graph. In at least one embodiment, a task graph describing the workload is constructed. Next, an executable graph is generated from the task graph using a process such as instantiation. Next, the executable graph is started one or more times, and each start of the executable graph causes the workload of the task graph to be implemented. Next, another task graph describing another workload is constructed. Instead of generating a new executable graph from the other task graph, one or more API and / or library function calls are then used to attempt to apply the other task graph to the already generated executable graph. Attempting to apply the other task graph to the executable graph includes comparing the topologies of the other task graph and the executable graph to see if those topologies are the same topology. If the topologies are the same topology, then the parameters of the other task graph are applied to the executable graph. If each parameter of the other task graph is successfully applied to the executable graph, then the executable graph can be started so that one or more computing resources implement the other workload of the other task graph without incurring the overhead of generating a new executable graph from the other task graph. If the other task graph cannot be successfully applied to the executable graph, a new executable graph is generated from the other task graph.

[0338] At least one technical advantage of the disclosed embodiments is that the disclosed embodiments can be used to apply a new task graph to an executable graph generated from another task graph without having to generate additional executable graphs and / or know the topology of any of the task graph and / or executable graph, such that the executable graph can be executed to implement another workload. Thus, with the disclosed embodiments, one or more computing resources can be made to implement different workloads without incurring the often significant overhead associated with generating a new executable graph whenever a new task graph is built. In contrast to prior art methods, the disclosed embodiments enable a single executable graph to be used to cause one or more computing resources to implement multiple different workloads, which reduces the overall execution time relative to prior art methods. These technical advantages provide one or more technological advancements over prior art methods.

[0339] 1. In at least one embodiment, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to perform at least one application programming interface (API) call to modify an executable version of a first task graph by applying a non-executable version of a second task graph to the executable version of the first task graph.

[0340] 2. The non-transitory computer-readable medium according to clause 1, wherein applying the non-executable version of the second task graph to the executable version of the first graph comprises: comparing a first topology associated with the executable version of the first task graph and a second topology associated with the non-executable version of the second task graph; determining that the first topology is the same topology as the second topology; and applying one or more parameters of the non-executable version of the second task graph to the executable version of the first task graph.

[0341] 3. The non-transitory computer-readable medium according to clause 1 or clause 2, wherein after applying the one or more parameters of the non-executable version of the second task graph to the executable version of the first task graph, the executable version of the first task graph configures one or more computing resources at startup to implement a workload associated with the non-executable version of the second task graph rather than a workload associated with the non-executable version of the first task graph.

[0342] 4. The non-transitory computer-readable medium according to any one of clauses 1-3, wherein when a first parameter of the non-executable version of the second task graph is applied to the executable version of the first task graph, the first parameter does not interfere with the optimizations incorporated into the executable version of the first task graph when generating the executable version of the first task graph from the non-executable version of the first task graph.

[0343] 5. A non-transitory computer-readable medium according to any one of claims 1-4, wherein applying a first parameter of a non-executable version of a second task graph to a corresponding parameter of an executable version of a first task graph changes the corresponding parameter, the corresponding parameter including: for a host task associated with the executable version of the first task graph, a pointer to a callback function or an argument of the callback function on a host central processing unit; for a memory setup task associated with the executable version of the first task graph, a location of a memory block to be set, a size of the memory block, or a fill value of the memory block; for a memory copy task associated with the executable version of the first task, a location of a source memory block, a destination location where the content of the source memory is to be copied, or a size of the source memory block; or for a kernel task associated with the executable version of the first graph, one or more arguments or a number of threads.

[0344] 6. A non-transitory computer-readable medium according to any one of claims 1-5, wherein comparing a first topology with a second topology is performed by: determining whether each node included in the first topology corresponds to a node having the same task type included in the second topology; and determining whether each node included in the second topology corresponds to a node having the same task type included in the first topology.

[0345] 7. A non-transitory computer-readable medium according to any one of claims 1-6, wherein two nodes correspond to each other when the nodes included in the first topology and the nodes included in the second topology are associated with the same creation order value.

[0346] 8. A non-transitory computer-readable medium according to any one of claims 1-7, wherein comparing the first topology with the second topology is further performed by determining whether each node included in the first topology and the corresponding node included in the second topology have the same dependency list.

[0347] 9. A non-transitory computer-readable medium according to any one of claims 1-8, wherein comparing the first topology with the second topology is further performed by determining whether the first topology and the second topology have the same number of nodes.

[0348] 10. In at least one embodiment, a computer-implemented method for applying a new task graph to an executable graph includes: modifying an executable version of a first task graph by applying a non-executable second task graph to the executable version of the first task graph.

[0349] 11. The computer-implemented method according to claim 10, further comprising: after applying the non-executable second task graph to the executable version of the first task graph, the executable version of the first task graph configures one or more computing resources at startup to implement the workload associated with the non-executable version of the second task graph, rather than the workload associated with the non-executable version of the first task graph.

[0350] 12. The computer-implemented method according to claim 10 or 11, wherein applying the second task graph to the executable version of the first graph comprises: comparing a first topology of the executable version of the first task graph with a second topology of the non-executable version of the second task graph; and when the first topology is the same topology as the second topology, applying the parameters of the second task graph to the executable version of the first task graph.

[0351] 13. The computer-implemented method according to any one of claims 10-12, wherein applying a first parameter of the non-executable second task graph to a corresponding parameter of the executable version of the first task graph changes the corresponding parameter, and the corresponding parameter includes: for a host task associated with the executable version of the first task graph, a pointer to a callback function or an argument of the callback function on a host central processing unit; for a memory setting task associated with the executable version of the first task graph, the location of a memory block to be set, the size of the memory block, or the fill value of the memory block; for a memory copy task associated with the executable version of the first task, the location of a source memory block, the destination location where the content of the source memory is to be copied, or the size of the source memory block; or for a kernel task associated with the executable version of the first graph, one or more arguments or the number of threads.

[0352] 14. The computer-implemented method according to any one of claims 10-13, wherein comparing the first topology with the second topology includes one or more of the following: determining whether the first topology and the second topology have the same number of nodes; determining whether each first node in the first topology has a corresponding second node of the same task type in the second topology; determining whether each second node in the second topology has a corresponding first node of the same task type in the first topology; determining whether each first node in the first topology and the corresponding second node in the second topology have the same dependency list; or determining whether each second node in the second topology and the corresponding first node in the first topology have the same dependency list.

[0353] 15. The computer-implemented method according to any one of claims 10-14, wherein at least one of the first task graph or the second task graph includes a Compute Unified Device Architecture (CUDA) graph or an executable Heterogeneous Interface for Portability (HIP) graph.

[0354] 16. In at least one embodiment, a computer-implemented method for applying a new task graph to an executable graph includes: generating an executable graph from a non-executable first task graph, wherein the executable graph configures one or more computing resources to implement a first workload associated with the non-executable first task graph; and modifying the executable graph by applying one or more parameters of a non-executable second task graph to the executable graph, wherein the modified executable graph configures the one or more computing resources to implement a second workload associated with the non-executable second task graph and different from the first workload.

[0355] 17. The computer-implemented method according to clause 16, wherein applying the non-executable second task graph to the executable graph includes: comparing a first topology of the executable graph with a second topology of the non-executable second task graph; and when the first topology and the second topology are the same topology, applying the parameters of the non-executable second task graph to the executable graph.

[0356] 18. The computer-implemented method according to clause 16 or 17, wherein comparing the first topology with the second topology includes: determining whether a first node in the first topology and a corresponding second node in the second topology have the same type and the same list of dependencies.

[0357] 19. The computer-implemented method according to any one of clauses 16-18, wherein the corresponding second node corresponds to the first node when the first node and the second node are associated with the same creation order value.

[0358] 20. The computer-implemented method according to any one of clauses 16-19, wherein at least one of the first task graph or the second task graph includes a Compute Unified Device Architecture (CUDA) graph or an executable Heterogeneous Interface for Portability (HIP) graph.

[0359] Any and all combinations, in any way, of any claim elements recited in any claim and / or any elements described in this application fall within the intended scope of this embodiment and protection.

[0360] Other variations are within the spirit of this disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative configurations, certain illustrative embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but rather, the intention is to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the disclosure as defined in the appended claims.

[0361] In the context of describing the disclosed embodiments (especially in the context of the following claims), the use of the terms "a", "an", "the", and similar references should be construed to cover both the singular and the plural, unless otherwise specified herein or clearly contradicted by the context and not as a definition of the term. The terms "comprising", "having", "including", and "containing" should be construed as open-ended terms (meaning "including but not limited to"), unless otherwise stated. The term "connected" when unmodified and referring to a physical connection should be construed to mean partially or wholly enclosed within, attached to, or joined together, even if something intervenes. The recitation of a range of values herein is merely intended to be a shorthand method of individually referring to each separate value falling within the range, unless otherwise specified herein, and each separate value is incorporated into the specification as if it were individually recited herein. Unless otherwise stated or contradicted by the context, the use of the term "set" (e.g., "set of items") or "subset" should be construed to include a non-empty set of one or more members. Further, unless otherwise stated or contradicted by the context, the term "subset" of a corresponding set does not necessarily denote a proper subset of the corresponding set, but rather the subset and the corresponding set can be equal.

[0362] Unless otherwise expressly stated or clearly contradicted by the context, conjunctive language such as phrases of the form "at least one of A, B, and C" or "one or more of A, B, and C" is otherwise understood in context and is generally used to present that the items, elements, etc. can be either A or B or C, or any non-empty subset of the set of A and B and C. For example, in an illustrative example of a set with three members, the conjunctive phrases "at least one of A, B, and C" and "one or more of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language is not generally intended to imply that some embodiments require the presence of each of at least one of A, at least one of B, and at least one of C. Further, unless otherwise stated or contradicted by the context, the term "plural" indicates a state of being plural (e.g., "plural items" indicates multiple items). The number of plural items is at least two, but can be more when explicitly or by context so indicated. Further, unless otherwise stated or otherwise clear from the context, the phrase "based on" means "at least partially based on" and not "merely based on".

[0363] The operations of the processes described herein may be implemented in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are implemented under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed collectively on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored, for example, in the form of a computer program including a plurality of instructions executable by one or more processors, on a computer-readable storage medium. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., propagating transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, the code (e.g., executable code or source code) is stored on a collection of one or more non-transitory computer-readable storage media (or other memory storing executable instructions) on which the executable instructions, when executed by one or more processors of a computer system (i.e., as a result of being executed by them), cause the computer system to perform the operations described herein. In at least one embodiment, the collection of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media of the plurality of non-transitory computer-readable storage media lack all of the code, while the plurality of non-transitory computer-readable storage media collectively store all of the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors - for example, the non-transitory computer-readable storage media stores instructions, and a main central processing unit ("CPU") executes some instructions while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of a computer system have separate processors, and different processors execute different subsets of instructions.

[0364] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software that permits the operations to be performed. Further, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system including a plurality of devices that operate differently such that the distributed computer system performs the operations described herein and such that a single device does not perform all of the operations.

[0365] Unless otherwise stated, any and all examples or use of exemplary language (e.g., "such as / for example") provided herein are merely intended to better illustrate embodiments of the present disclosure and do not constitute a limitation on the scope of the present disclosure. The language of the specification should not be construed as indicating that any non-claimed element is essential to the practice of the present disclosure.

[0366] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference in their entirety, to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in full herein.

[0367] In the specification and claims, the terms "coupled" and "connected" may be used with their derivatives. It should be understood that these terms are not necessarily intended as synonyms for each other. Instead, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0368] Unless otherwise specifically stated, it is understood that throughout the specification, terms such as "processing", "computing", "operating", "determining", etc. refer to actions and / or processes of a computer or computing system, or similar electronic computing device, which manipulate and / or transform data represented as physical quantities, such as electrical, in registers and / or memories of the computing system into other data similarly represented as physical quantities in memories, registers, or other such information storage, transmission, or display devices of the computing system.

[0369] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data from registers and / or memories and transforms the electronic data into other electronic data that may be stored in registers and / or memories. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process may involve multiple processes for executing instructions continuously or intermittently, sequentially or in parallel. The terms "system" and "method" may be used interchangeably herein, provided that a system can embody one or more methods and a method can be considered a system.

[0370] In this document, there may be mention of obtaining, acquiring, receiving analog or digital data or inputting such data into a subsystem, computer system, or computer-implemented machine. The process of obtaining, acquiring, receiving, or inputting analog and digital data can be implemented in a variety of ways - for example, by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be implemented by transmitting data from a providing entity to an acquiring entity via a computer network. There may also be mention of providing, outputting, transferring, sending, or presenting analog or digital data. In various different examples, the process of providing, outputting, transferring, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter of a function call, an application programming interface, or a parameter of an interprocess communication mechanism.

[0371] Although the above discussion describes example implementations of the described technology, other architectures can be used to implement the described functionality and are expected to be within the scope of this disclosure. Additionally, although specific distributions of responsibilities were defined above for purposes of discussion, the various different functions and responsibilities may be distributed and divided in different ways depending on the circumstances.

[0372] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.

Claims

1. A non - transitory computer - readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform at least one application programming interface (API) call to modify an executable version of a first task graph by applying one or more parameters of a non - executable version of a second task graph to the executable version of the first task graph, to configure one or more computing resources to implement a workload associated with the non - executable version of the second task graph instead of implementing a workload associated with the non - executable version of the first task graph.

2. The non - transitory computer - readable medium of claim 1, wherein the non - executable version of the second task graph is applied to the executable version of the first task graph by: comparing a first topology associated with the executable version of the first task graph and a second topology associated with the non - executable version of the second task graph; determining that the first topology is the same topology as the second topology; and applying the one or more parameters of the non - executable version of the second task graph to the executable version of the first task graph.

3. The non - transitory computer - readable medium of claim 2, wherein when a first parameter of the non - executable version of the second task graph is applied to the executable version of the first task graph, the first parameter does not interfere with the optimizations incorporated into the executable version of the first task graph when the executable version of the first task graph is generated from the non - executable version of the first task graph.

4. The non - transitory computer - readable medium of claim 2, wherein applying the first parameter of the non - executable version of the second task graph to a corresponding parameter of the executable version of the first task graph changes the corresponding parameter, and the corresponding parameter includes: for a host task associated with the executable version of the first task graph, a pointer to a callback function or an argument of the callback function on a host central processing unit; for a memory - setting task associated with the executable version of the first task graph, the location of a memory block to be set, the size of the memory block, or the fill value of the memory block; for a memory - copy task associated with the executable version of the first task graph, the location of a source memory block, the location of a destination to which the content of the source memory is to be copied, or the size of the source memory block; or for a kernel task associated with the executable version of the first task graph, the number of threads or one or more arguments.

5. The non - transitory computer - readable medium of claim 2, wherein the first topology is compared with the second topology by: determining whether each node included in the first topology corresponds to a node having the same task type included in the second topology; and Determine whether each node included in the second topology corresponds to a node with the same task type included in the first topology.

6. The non-transitory computer-readable medium according to claim 5, wherein when a node included in the first topology and a node included in the second topology are associated with the same creation order value, the node included in the first topology corresponds to the node included in the second topology.

7. The non-transitory computer-readable medium according to claim 5, wherein the first topology is further compared with the second topology by determining whether each node included in the first topology and the corresponding node included in the second topology have the same dependency list.

8. The non-transitory computer-readable medium according to claim 5, wherein the first topology is further compared with the second topology by determining whether the first topology and the second topology have the same number of nodes.

9. A computer-implemented method for applying a new task graph to an executable graph, the method comprising: Modifying the executable version of the first task graph by applying one or more parameters of a non-executable second task graph to the executable version of the first task graph to configure one or more computing resources to implement the workload associated with the non-executable version of the second task graph instead of implementing the workload associated with the non-executable version of the first task graph.

10. The computer-implemented method according to claim 9, wherein applying the second task graph to the executable version of the first task graph comprises: Comparing a first topology of the executable version of the first task graph with a second topology of the non-executable version of the second task graph; And When the first topology is the same topology as the second topology, applying the parameters of the second task graph to the executable version of the first task graph.

11. The computer-implemented method according to claim 10, wherein applying the first parameter of the non-executable second task graph to the corresponding parameter of the executable version of the first task graph changes the corresponding parameter, and the corresponding parameter includes: For a host task associated with the executable version of the first task graph, a pointer to a callback function or an argument of the callback function on a host central processing unit; For a memory setting task associated with the executable version of the first task graph, the location of a memory block to be set, the size of the memory block, or the fill value of the memory block; For a memory copy task associated with the executable version of the first task graph, the location of a source memory block, the location of a destination to which the content of the source memory is to be copied, or the size of the source memory block; Or For a kernel task associated with the executable version of the first task graph, the number of threads or one or more arguments.

12. The computer-implemented method according to claim 10, wherein comparing the first topology with the second topology includes one or more of the following: Determining whether the first topology and the second topology have the same number of nodes; Determining whether each first node in the first topology has a corresponding second node of the same task type in the second topology; Determining whether each second node in the second topology has a corresponding first node of the same task type in the first topology; Determining whether each first node in the first topology and the corresponding second node in the second topology have the same dependency list; Or Determining whether each second node in the second topology and the corresponding first node in the first topology have the same dependency list.

13. The computer-implemented method according to claim 9, wherein at least one of the first task graph or the second task graph includes a Compute Unified Device Architecture (CUDA) graph or an executable Heterogeneous Interface for Portability (HIP) graph.

14. A computer-implemented method for applying a new task graph to an executable graph, the method comprising: Generating an executable graph from a non-executable first task graph, wherein the executable graph configures one or more computing resources to implement a first workload associated with the non-executable first task graph; and Modifying the executable graph by applying one or more parameters of a non-executable second task graph to the executable graph, wherein the modified executable graph configures the one or more computing resources to implement a second workload associated with the non-executable second task graph and different from the first workload.

15. The computer-implemented method according to claim 14, wherein applying the non-executable second task graph to the executable graph includes: Comparing a first topology of the executable graph with a second topology of the non-executable second task graph; And When the first topology and the second topology are the same topology, applying the parameters of the non-executable second task graph to the executable graph.

16. The computer-implemented method according to claim 15, wherein comparing the first topology with the second topology comprises: Determining whether a first node in the first topology and a corresponding second node in the second topology have the same type and the same dependency list.

17. The computer-implemented method according to claim 16, wherein when both the first node and the second node are associated with the same creation order value, the corresponding second node corresponds to the first node.

18. The computer-implemented method according to claim 14, wherein at least one of the first task graph or the second task graph includes a Compute Unified Device Architecture (CUDA) graph or an executable Heterogeneous Interface for Portability (HIP) graph.

Citation Information

Patent Citations

  • Differencing of executable dataflow graphs

    US20180157579A1

  • Parameterized graphs with conditional components

    US7164422B1

  • Electronic device, control method therefor, and non-transitory computer readable recording medium

    WO2018155810A1