Scalable centralized error queues in processing architecture

By designing a scalable centralized error queue in the graphics processor, the problem of inefficient error detection, recording and reporting in the prior art is solved, and the system is high reliability and usability is achieved.

CN120196470APending Publication Date: 2025-06-24INTEL CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411573959.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to implement efficient error detection, recording and reporting in graphics processors, resulting in limitations in system reliability, availability and serviceability.

Method used

A scalable centralized error queue is designed to achieve centralized management and reporting of errors through the error log format and report message format in the graphics processor. The system includes error log format, error log header format, and error log report message format, and is able to record and report various types of errors.

Benefits of technology

Improve the reliability and availability of the graphics processor system. By centrally managing errors, problems can be detected and repaired quickly and the serviceability of the system can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196470A_ABST
    Figure CN120196470A_ABST
Patent Text Reader

Abstract

An apparatus for facilitating processing scalable centralized error queues in an architecture is disclosed. The apparatus comprises: a processor comprising a system interface hosting an error aggregator, where the processor is configured to: hosting at least one centralized error queue in the error aggregator, the at least one centralized error queue to store an error log for errors detected by components of the processor; receiving an error report message from one of the components of the processor, the error report message corresponding to an error detected by the component; and based on the error type of the error, recording the error as an entry in at least one centralized error queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document generally relates to data processing, and more specifically, to a scalable centralized error queue in a processing architecture. Background Art

[0002] Current parallel graphics data processing includes systems and methods developed to perform specific operations on graphics data, such as, for example, linear interpolation, tessellation, rasterization, texture mapping, depth testing, etc. Traditionally, graphics processors have used fixed-function computing units to process graphics data; however, more recently, multiple parts of graphics processors have been made programmable, enabling such processors to support a broader variety of operations for processing vertex data and fragment data.

[0003] To further improve performance, graphics processors typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across different parts of the graphics pipeline. Parallel graphics processors with single instruction, multiple data (SIMD) or single instruction, multiple thread (SIMT) architectures are designed to maximize the amount of parallel processing in the graphics pipeline. In the SIMD architecture, a computer with multiple processing elements attempts to perform the same operation on multiple data points simultaneously. In the SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as frequently as possible to improve processing efficiency.

[0004] Graphics processors are often used in applications in the fields of artificial intelligence (AI) and machine learning (ML). General-purpose graphics processing units (GPGPUs) with matrix acceleration circuitry are often deployed and hosted in data centers. Hardware resilience is a requirement for graphics architectures in the data center market segment. Architectures with reliability, availability, and serviceability (RAS) features are designed to meet resilience goals. Reliability refers to how reliable the operation of the design is. Availability refers to the uptime of the operation that the design can still provide in the presence of errors. Serviceability refers to how easily the design can be serviced to restore it to reliable operation once an error occurs.

[0005] To do this, the design should be able to detect and record errors (to improve reliability), correct these errors if possible (to improve availability), and report to a higher-level system component (such as a driver) when the error is not correctable. The system software can then take appropriate actions to service the error and bring the design back to reliable operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Thus, to understand in detail the manner in which the features of the current embodiments described above are implemented, a more specific description of the embodiments briefly outlined above may be made with reference to the embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings only illustrate typical embodiments and should not be considered as limiting the scope of the embodiments.

[0008] Figure 1 is a block diagram of a processing system.

[0009] Figures 2A - 2D illustrates a computing system and a graphics processor.

[0010] Figures 3A - 3C illustrates a block diagram of an additional graphics processor and computing accelerator architecture.

[0011] Figure 4 is a block diagram of a graphics processing engine of a graphics processor.

[0012] Figures 5A - 5B illustrates thread execution logic including an array of processing elements employed in a graphics processor core.

[0013] Figure 6 illustrates additional execution units.

[0014] Figure 7 is a block diagram illustrating a graphics processor instruction format.

[0015] Figure 8 is a block diagram of an additional graphics processor architecture.

[0016] Figures 9A - 9B illustrates a graphics processor command format and command sequence.

[0017] Figure 10 illustrates an example graphics software architecture for a data processing system.

[0018] Figure 11A is a block diagram illustrating an IP core development system.

[0019] Figure 11B illustrates a cross-sectional side view of an integrated circuit package component.

[0020] Figure 11CThe figure shows a packaged component that includes a hardware logic die connected to multiple units of a substrate (e.g., a base die).

[0021] Figure 11D The figure shows a packaged component that includes interchangeable dies.

[0022] Figure 12 is a block diagram of an example system-on-chip integrated circuit shown in the figure.

[0023] Figures 13A - 13B is a block diagram of an example graphics processing unit for use within a SoC shown in the figure.

[0024] Figure 14 The figure shows an embodiment of a graphics processing unit that includes a scalable centralized error queue according to the implementations herein.

[0025] Figure 15 Depicts an error log format for reporting errors to a scalable centralized error queue according to the implementations herein.

[0026] Figure 16 Depicts an error log header format for an error log report message for a scalable centralized error queue according to the implementations herein.

[0027] Figure 17 Depicts a grouping of an internal error log and a structural error log for an error log report message for a scalable centralized error queue according to the implementations herein.

[0028] Figure 18 Depicts status and control registers used by software as part of reading entries from a scalable centralized error queue according to the implementations herein.

[0029] Figure 19 is a flowchart of an embodiment of a method for reporting errors to a scalable centralized error queue in a graphics architecture.

[0030] Figure 20 is a flowchart of an embodiment of a method for recording errors in a scalable centralized error queue in a graphics architecture.

[0031] Figure 21 is a flowchart of an embodiment of a method for recovering errors in a scalable centralized error queue in a graphics architecture. Detailed Description

[0032] A graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate, for example, graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU can be integrated on the same package or die as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or die). Regardless of the manner in which the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0033] In the following description, numerous specific details are set forth in order to provide a more thorough understanding. However, it will be apparent to one of ordinary skill in the art that embodiments described herein may be practiced without one or more of these specific details. In other instances, well-known features have not been described in order to avoid obscuring the details of the current embodiments. System Overview

[0034] Figure 1 is a block diagram of a processing system 100 according to an embodiment. The system 100 can be used in: a single-processor desktop computer system, a multi-processor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in a mobile device, a handheld device, or an embedded device, such as for use within an Internet-of-things (IoT) device having wired or wireless connectivity to a local area network or a wide area network.

[0035] In one embodiment, system 100 may include, may be coupled with, or may be integrated within: a server-based gaming platform; a game console, including a gaming and media console; a mobile gaming console, a handheld gaming console, or an online gaming console. In some embodiments, system 100 is part of a mobile phone, a smartphone, a tablet computing device, or a mobile Internet-connected device (such as a laptop computer with low internal storage capacity). Processing system 100 may also include, be coupled with, or be integrated within: a wearable device, such as a smartwatch wearable device; smart glasses or clothing that are enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, audio, or tactile outputs to supplement real-world visual, audio, or tactile experiences or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. In some embodiments, processing system 100 includes a television or set-top box device, or is part of a television or set-top box device. In one embodiment, system 100 may include, be coupled with, or be integrated within a self-driving vehicle, such as a bus, a tractor-trailer, a car, a motor or electric cycle, an airplane or a glider (or any combination thereof). The self-driving vehicle may use system 100 to process the environment sensed around the vehicle.

[0036] In some embodiments, one or more processors 102 each include one or more processor cores 107 that are configured to process instructions that, when executed, perform operations for system and user software. In the embodiments herein, a processor may refer to dedicated hardware circuitry for efficiently processing commands / instructions and may be referred to as processor circuitry. In some embodiments, at least one of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computing via Very Long Instruction Word (VLIW). The one or more processor cores 107 may process different instruction sets 109, and the different instruction sets 109 may include instructions for facilitating the emulation of other instruction sets. The processor core 107 may also include other processing devices, such as a Digital Signal Processor (DSP).

[0037] In some embodiments, the processor 102 includes a cache memory 104. Depending on the architecture, the processor 102 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among the various components of the processor 102. In some embodiments, the processor 102 also uses an external cache (e.g., a third-level (L3) cache or a Last Level Cache (LLC)) (not shown), and the external cache may be shared among the processor cores 107 using known cache coherence techniques. A register file 106 may additionally be included in the processor 102 and may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while other registers may be dedicated to the design of the processor 102.

[0038] In some embodiments, one or more processors 102 are coupled to one or more interface buses 110 to transfer communication signals, such as address, data, or control signals, between the processors 102 and other components in the system 100. In one embodiment, the interface bus 110 can be a processor bus, such as a certain version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In one embodiment, the (one or more) processors 102 include an integrated memory controller 116 and a Platform Controller Hub 130. The memory controller 116 facilitates communication between the memory device and other components of the system 100, while the Platform Controller Hub (PCH) 130 provides connections to I / O devices via a local I / O bus.

[0039] The memory device 120 can be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device having suitable performance to act as a process memory. In one embodiment, the memory device 120 can operate as the system memory for the system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. The memory controller 116 is also coupled to an optional external graphics processor 118, which can communicate with one or more graphics processors 108 in the processor 102 to perform graphics operations and media operations. In some embodiments, the graphics operations, media operations, and / or computing operations can be assisted by an accelerator 112, which is a coprocessor that can be configured to execute a set of specialized graphics operations, media operations, or computing operations. For example, in one embodiment, the accelerator 112 is a matrix multiplication accelerator for optimizing machine learning or computing operations. In one embodiment, the accelerator 112 is a ray tracing accelerator, which can be used to perform ray tracing operations in cooperation with the graphics processor 108. In one embodiment, an external accelerator 119 can be used instead of the accelerator 112, or can be used in cooperation with the accelerator 112.

[0040] In some embodiments, the display device 111 may be connected to the processor(s) 102. The display device 111 may be one or more of the following: an internal display device, such as in a mobile electronic device or a laptop computer device; or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 111 may be a head mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.

[0041] In some embodiments, the platform controller hub 130 enables peripheral devices to be connected to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripheral devices include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D Xpoint, etc.). The data storage device 124 may be connected via a storage interface (e.g., SATA) or via a peripheral bus (such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express)). The touch sensor 125 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, 5G, or Long-Term Evolution (LTE) transceiver. The firmware interface 128 enables communication with system firmware and may be, for example, a unified extensible firmware interface (UEFI). The network controller 134 enables a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. In one embodiment, the audio controller 146 is a multi-channel high-definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The platform controller hub 130 may also be connected to one or more Universal Serial Bus (USB) controllers 142 to connect to input devices, such as a keyboard and mouse 143 combination, a camera 144, or other USB input devices.

[0042] It will be appreciated that the illustrated system 100 is exemplary and not restrictive, as other types of data processing systems configured in different ways may also be used. For example, instances of the memory controller 116 and the platform controller hub 130 may be integrated into a discrete external graphics processor, such as the external graphics processor 118. In one embodiment, the platform controller hub 130 and / or the memory controller 116 may be external to one or more of the processors 102. For example, the system 100 may include an external memory controller 116 and a platform controller hub 130, and the platform controller hub 130 may be a memory controller hub and a peripheral controller hub within a system-on-chip configured to communicate with the processor(s) 102.

[0043] For example, a circuit board ("sled") may be used, on which components such as a CPU, memory, and other components are placed, and on which the components (such as a CPU, memory, and other components) are designed to achieve improved thermal performance. In some examples, processing components such as processors are located on the top side of the sled, while nearby memory such as DIMMs is located on the bottom side of the sled. As a result of the enhanced airflow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby improving performance. Additionally, the sled is configured for blind mating of power and data communication cables in a rack, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the sled (such as processors, accelerators, memory, and data storage drives) are configured to be easily upgradable due to their increased spacing from each other. In an illustrative embodiment, the components additionally include hardware authentication features for proving their authenticity.

[0044] A data center may utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omni-Path. The sled may be coupled to a switch via optical fiber, which provides higher bandwidth and lower latency than typical twisted-pair cabling (e.g., Category 5, Category 5e, Category 6, etc.). Due to the high-bandwidth, low-latency interconnect and network architecture, the data center in use may centralize physically dispersed resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network, and / or artificial intelligence accelerators, etc.), and data storage drives, and provide them to computing resources (e.g., processors) on a requested basis, enabling the computing resources to access the centralized resources as if the centralized resources were local.

[0045] A power supply or power source can supply voltage and / or current to system 100 or any component or system described herein. In one example, the power supply includes an AC to DC (alternating current to direct current) adapter for insertion into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source, such as an external AC to DC converter. In one example, the power source or power supply includes wireless charging hardware for charging via a proximity charging field. In one example, the power source can include an internal battery, an AC supply, an action-based power supply, a solar power supply, or a fuel cell source.

[0046] Figures 2A - 2D Illustrates a computing system and a graphics processor provided by the embodiments described herein. Figures 2A - 2D Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0047] Figure 2A is a block diagram of an embodiment of a processor 200 that has one or more processor cores 202A - 202N, an integrated memory controller 214, and an integrated graphics processor 208. Processor 200 can include additional cores, with the additional cores being up to and including additional core 202N represented by the dashed box. Each of processor cores 202A - 202N includes one or more internal cache units 204A - 204N. In some embodiments, each processor core also has access to one or more shared cache units 206. The internal cache units 204A - 204N and the shared cache units 206 represent the cache memory hierarchy within processor 200. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core and one or more levels of shared mid-level cache, such as a second level (L2), third level (L3), fourth level (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence between cache units 206 and 204A - 204N.

[0048] In some embodiments, the processor 200 may further include a set 216 of one or more bus controller units and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI buses or PCI Express buses. The system agent core 210 provides management functions for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).

[0049] In some embodiments, one or more of the processor cores 202A - 202N include support for simultaneous multithreading operations. In such embodiments, the system agent core 210 includes components for coordinating and operating the cores 202A - 202N during multithreaded processing. The system agent core 210 may additionally include a power control unit (PCU) that includes logic and components for regulating the power states of the processor cores 202A - 202N and the graphics processor 208.

[0050] In some embodiments, the processor 200 additionally includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to a set 206 of shared cache units and the system agent core 210, which includes one or more integrated memory controllers 214. In some embodiments, the system agent core 210 further includes a display controller 211 for driving the graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 can also be a separate module coupled to the graphics processor via at least one interconnect, or can be integrated within the graphics processor 208.

[0051] In some embodiments, a ring - based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units can be used, such as point - to - point interconnects, switched interconnects, or other techniques, including those well - known in the art. In some embodiments, the graphics processor 208 is coupled to the ring - based interconnect 212 via an I / O link 213.

[0052] The exemplary I / O link 213 represents at least one of a plurality of various I / O interconnects, including an on - package I / O interconnect that facilitates communication between the various processor components and a high - performance embedded memory module 218, such as an eDRAM module. In some embodiments, each of the processor cores 202A - 202N and the graphics processor 208 can use the embedded memory module 218 as a shared last - level cache.

[0053] In some embodiments, the processor cores 202A - 202N are homogeneous cores that execute the same instruction set architecture. In another embodiment, the processor cores 202A - 202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A - 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of computing power. Additionally, the processor 200 can be implemented on one or more chips or be implemented as a SoC integrated circuit that also has the illustrated components in addition to other components.

[0054] Figure 2B is a block diagram of the hardware logic of the graphics processor core 219 according to some embodiments described herein. Figure 2B Elements in with the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. The graphics processor core 219 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. The graphics processor core 219 is an example of a graphics core slice, and based on the target power envelope and performance envelope, a graphics processor as described herein can include multiple graphics core slices. Each graphics processor core 219 can include a fixed function block 230 that is coupled to a plurality of sub - cores 221A - 221F (also referred to as sub - slices), and the plurality of sub - cores 221A - 221F include blocks of modular general - purpose and fixed - function logic.

[0055] In some embodiments, the fixed function block 230 includes a geometry / fixed function pipeline 231, which can be shared by all sub - cores in the graphics processor core 219, for example, in a lower - performance and / or lower - power graphics processor implementation. In embodiments, the geometry / fixed function pipeline 231 includes a 3D fixed function pipeline (e.g., such as the 3D pipeline 312 in Figure 3A and Figure 4 ), a video front - end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer (e.g., the unified return buffer 418 in Figure 4 ).

[0056] In one embodiment, the fixed function block 230 further includes a graphics SoC interface 232, a graphics microcontroller 233, and a media pipeline 234. The graphics SoC interface 232 provides an interface between the graphics processor core 219 and other processor cores within the system-on-chip integrated circuit. The graphics microcontroller 233 is a programmable sub-processor that can be configured to manage various functions of the graphics processor core 219, including thread dispatch, scheduling, and preemption. The media pipeline 234 (e.g., Figure 3A and Figure 4 media pipeline 316) includes logic for facilitating the decoding, encoding, preprocessing, and / or postprocessing of multimedia data including image data and video data. The media pipeline 234 implements media operations via requests to the computing or sampling logic within the sub-cores 221 - 221F.

[0057] In one embodiment, the SoC interface 232 enables the graphics processor core 219 to communicate with a general-purpose application processor core (e.g., CPU) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache memory, system RAM, and / or embedded on-chip or package-on-chip DRAM. The SoC interface 232 can also enable communication with fixed function devices within the SoC such as a camera imaging pipeline, and enable the use and / or implementation of global memory atomicity, which can be shared between the graphics processor core 219 and the CPU within the SoC. The SoC interface 232 can also implement power management control for the graphics processor core 219 and enable an interface between the clock domain of the graphics core 219 and other clock domains within the SoC. In one embodiment, the SoC interface 232 enables receipt of command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. When a media operation is to be performed, these commands and instructions can be dispatched to the media pipeline 234, or when a graphics processing operation is to be performed, these commands and instructions can be dispatched to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 231, geometry and fixed function pipeline 237).

[0058] The graphics microcontroller 233 can be configured to perform various scheduling tasks and management tasks for the graphics processor core 219. In one embodiment, the graphics microcontroller 233 can schedule graphics and / or compute workloads for the execution unit (EU) arrays 222A - 222F, 224A - 224F within the sub - cores 221A - 221F. In this scheduling model, host software executing on the CPU core of the SoC that includes the graphics processor core 219 can submit a workload to one of the multiple graphics processor doorbells, which invokes a scheduling operation for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, pre - empting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 233 can also facilitate the low - power or idle state of the graphics processor core 219, thereby providing the ability to save and restore registers within the graphics processor core 219 across low - power state transitions independent of the operating system and / or graphics driver software on the system.

[0059] The graphics processor core 219 can have more or fewer sub - cores 221A - 221F than shown, up to N modular sub - cores. For each set of N sub - cores, the graphics processor core 219 can also include shared functional logic 235, shared and / or cache memory 236, a geometry / fixed - function pipeline 237, and additional fixed - function logic 238 for accelerating various graphics and compute processing operations. The shared functional logic 235 can include logic units that are associated with Figure 4 the shared functional logic 420 (e.g., sampler logic, math logic, and / or inter - thread communication logic) and can be shared by each N sub - cores within the graphics processor core 219. The shared and / or cache memory 236 can be the last - level cache for the set of N sub - cores 221A - 221F within the graphics processor core 219 and can also act as shared memory accessible by multiple sub - cores. The geometry / fixed - function pipeline 237 rather than the geometry / fixed - function pipeline 231 can be included within the fixed - function block 230, and the geometry / fixed - function pipeline 237 can include the same or similar logic units.

[0060] In one embodiment, the graphics processor core 219 includes additional fixed function logic 238, which may include various fixed function acceleration logics for use by the graphics processor core 219. In one embodiment, the additional fixed function logic 238 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are two geometry pipelines: the full geometry pipeline within the geometry / fixed function pipelines 237, 231; and a culling pipeline, which is an additional geometry pipeline that may be included within the additional fixed function logic 238. In one embodiment, the culling pipeline is a trimmed-down version of the full geometry pipeline. The full pipeline and the culling pipeline may execute different instances of the same application, each instance having a separate context. Position-only shading can hide the long culling runs of discarded triangles, thus enabling earlier completion of shading in some instances. For example and in one embodiment, the culling pipeline logic within the additional fixed function logic 238 can execute the position shader in parallel with the main application and generally produce results faster than the full pipeline, because the culling pipeline only takes the position attributes of the vertices and only shades the position attributes of the vertices, without performing rasterization and rendering of pixels to the frame buffer. The culling pipeline can use the generated results to calculate the visibility information of all triangles, regardless of whether those triangles are culled. The full pipeline (which may be referred to as a replay pipeline in this instance) can consume this visibility information to skip the culled triangles, thus only shading the visible triangles that are ultimately passed to the rasterization stage.

[0061] In one embodiment, the additional fixed function logic 238 may further include machine learning acceleration logic, such as fixed function matrix multiplication logic, for an implementation that includes optimizations for machine learning training or inference.

[0062] Each graphics sub-core 221A-221F includes a set of execution resources that can be used to perform graphics operations, media operations, and compute operations in response to requests made by the graphics pipeline, media pipeline, or shader program. The graphics sub-cores 221A-221F include: multiple EU arrays 222A-222F, 224A-224F; thread dispatch and inter-thread communication (TD / IC) logic 223A-223F; 3D (e.g., texture) samplers 225A-225F; media samplers 206A-206F; shader processors 227A-227F; and shared local memory (SLM) 228A-228F. The EU arrays 222A-222F, 224A-224F each include a plurality of execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations to service graphics operations, media operations, or compute operations (including graphics programs, media programs, or compute shader programs). The TD / IC logic 223A-223F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. The 3D samplers 225A-225F can read texture or other 3D graphics-related data into memory. The 3D samplers can read texture data in different ways based on the configured sample state and the texture format associated with a given texture. The media samplers 206A-206F can perform similar read operations based on the type and format associated with the media data. In one embodiment, each graphics sub-core 221A-221F can alternatively include unified 3D and media samplers. Threads executing on the execution units within each of the sub-cores 221A-221F can utilize the shared local memory 228A-228F within each sub-core to enable threads executing within a thread group to use a common pool of on-chip memory for execution.

[0063] Figure 2C Illustrated is a graphics processing unit (GPU) 239 that includes a set of dedicated graphics processing resources arranged as multi-core groups 240A-240N. While details are provided for only a single multi-core group 240A, it will be understood that the other multi-core groups 240B-240N can be equipped with the same or similar sets of graphics processing resources.

[0064] The set of register files 242 stores operand values used by cores 243, 244, 245 when executing graphics threads. These register files may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. In one embodiment, the tile registers are implemented as a combined set of vector registers.

[0065] One or more combined level-1 (L1) caches and shared memory units 247 store graphics data locally within each multi-core group 240A, such as texture data, vertex data, pixel data, light data, bounding volume data, etc. One or more texture units 247 may also be used to perform texture operations, such as texture mapping and sampling. A level-2 (L2) cache 253 shared by all multi-core groups 240A - 240N or a subset of multi-core groups 240A - 240N stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 253 can be shared across multiple multi-core groups 240A - 240N. One or more memory controllers 248 couple the GPU 239 to a memory 249, which can be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0066] Input / Output (I / O) circuitry 250 couples the GPU 239 to one or more I / O devices 252, such as digital signal processors (DSPs), network controllers, or user input devices. On-chip interconnects can be used to couple the I / O devices 252 to the GPU 239 and the memory 249. One or more I / O memory management units (IOMMUs) 251 of the I / O circuitry 250 directly couple the I / O devices 252 to the system memory 249. In one embodiment, the IOMMU 251 manages multiple sets of page tables for mapping virtual addresses to physical addresses in the system memory 249. In this embodiment, the I / O devices 252, the (one or more) CPUs 246, and the (one or more) GPUs 239 can share the same virtual address space.

[0067] In one implementation, the IOMMU 251 supports virtualization. In this case, the IOMMU 251 can manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 249). The base address of each of the first set of page tables and the second set of page tables can be stored in a control register and swapped out during a context switch (e.g., such that a new context is provided with access to the relevant page table set). Although not illustrated in Figure 2C each of the cores 243, 244, 245, and / or the multi-core groups 240A - 240N may include translation lookaside buffers (TLBs) for caching guest virtual-to-guest physical translations, guest physical-to-host physical translations, and guest virtual-to-host physical translations.

[0068] In one embodiment, the CPU 246, GPU 239, and I / O device 252 are integrated on a single semiconductor chip and / or chip package. The illustrated memory 249 may be integrated on the same chip or may be coupled to the memory controller 248 via an off-chip interface. In one implementation, the memory 249 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles of the embodiments herein are not limited to this particular implementation.

[0069] In one embodiment, the tensor core 244 includes a plurality of execution units specifically designed to perform matrix operations, which are fundamental computational operations for performing deep learning operations. For example, synchronous matrix multiplication operations can be used for neural network training and inference. The tensor core 244 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). In one embodiment, a neural network implementation extracts features of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.

[0070] In a deep learning implementation, parallel matrix multiplication work can be scheduled for execution on the tensor core 244. The training of neural networks particularly utilizes a large number of matrix dot product operations. To handle the inner product formulation of an N x N x N matrix multiplication, the tensor core 244 may include at least N dot product processing elements. Before the matrix multiplication begins, a complete matrix is loaded into the on-chip registers, and for each of the N loops, at least one column of the second matrix is loaded. For each loop, there are N dot products to be processed.

[0071] Depending on the specific implementation, matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for tensor core 244 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads, which can tolerate quantization down to bytes and nibbles).

[0072] In one embodiment, ray tracing core 245 accelerates ray tracing operations for both real-time ray tracing implementations and non-real-time ray tracing implementations. Specifically, ray tracing core 245 includes a ray traversal / intersection circuit that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays enclosed within the BVH volume and primitives. Ray tracing core 245 may also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, ray tracing core 245 performs traversal and intersection operations in concert with the image denoising techniques described herein, at least part of which may be performed on tensor core 244. For example, in one embodiment, tensor core 244 implements a deep learning neural network to perform denoising of frames generated by ray tracing core 245. However, the (one or more) CPUs 246, graphics core 243, and / or ray tracing core 245 may also implement all or part of the denoising and / or deep learning algorithms.

[0073] In addition, as described above, a distributed approach to denoising can be employed, where GPU 239 is in a computing device coupled to other computing devices via a network or high-speed interconnect. In this embodiment, the interconnected computing devices share neural network learning / training data to improve the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.

[0074] In one embodiment, the ray tracing core 245 processes all BVH traversals and ray-primitive intersections, freeing the graphics core 243 from being overloaded with thousands of instructions per ray. In one embodiment, each ray tracing core 245 includes a first set of specialized circuitry for performing bounding box tests (e.g., for traversal operations) and a second set of specialized circuitry for performing ray-triangle intersection tests (e.g., intersecting the traversed rays). Thus, in one embodiment, the multi-core group 240A can simply initiate a ray query, and the ray tracing core 245 independently performs ray traversal and intersection and returns hit data (e.g., hit, miss, multiple hits, etc.) to the thread context. While the ray tracing core 245 performs traversal and intersection operations, the other cores 243, 244 are freed to perform other graphics or compute work.

[0075] In one embodiment, each ray tracing core 245 includes a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "miss", or "multiple hits" responses, which the intersection unit provides to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 243 and the tensor core 244) are freed to perform other forms of graphics work.

[0076] In a particular embodiment described below, a hybrid rasterization / ray tracing method is used in which work is distributed between the graphics core 243 and the ray tracing core 245.

[0077] In one embodiment, the ray tracing core 245 (and / or other cores 243, 244) includes hardware support for a ray tracing instruction set such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, closest hit shaders, any hit shaders, and miss shaders, which enable assignment of shader and texture sets to each object. Another ray tracing platform that can be supported by the ray tracing core 245, the graphics core 243, and the tensor core 244 is Vulkan 1.1.85. However, note that the basic principles of the embodiments herein are not limited to any particular ray tracing ISA.

[0078] Generally, the respective cores 245, 244, 243 can support a ray tracing instruction set including instructions / functions for the following: ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traverse, and exception. More specifically, one embodiment includes ray tracing instructions for performing the following functions:

[0079] Light Generation - Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment.

[0080] Nearest Hit - The closest hit instruction can be executed to locate the closest intersection of a ray with a primitive within a scene.

[0081] Any Hit - Any hit instruction identifies multiple intersections between a ray and primitives within a scene, potentially identifying a new closest intersection.

[0082] Intersection - The intersection instruction performs a ray-primitive intersection test and outputs the result.

[0083] Primitive Bounding Box Construction - This instruction builds a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure).

[0084] Miss - Indicates that the ray misses all geometries within the scene or a specified region of the scene.

[0085] Visit - Indicates the child volume that the ray will traverse.

[0086] Exception - Includes various types of exception handlers (e.g., called for various error conditions).

[0087] Figure 2D is a block diagram of a general-purpose graphics processing unit (GPGPU) 270 according to an embodiment described herein. The GPGPU 270 can be configured as a graphics processor and / or a computing accelerator. The GPGPU 270 can be interconnected with a host processor (e.g., one or more CPUs 246) and memories 271, 272 via one or more system and / or memory buses. In one embodiment, the memory 271 is a system memory that can be shared with one or more CPUs 246, and the memory 272 is a device memory dedicated to the GPGPU 270. In one embodiment, the components within the GPGPU 270 and the device memory 272 can be mapped to memory addresses that can be accessed by one or more CPUs 246. Access to the memories 271 and 272 can be facilitated via a memory controller 268. In one embodiment, the memory controller 268 includes an internal direct memory access (DMA) controller 269, or can include logic for performing operations that would otherwise be performed by a DMA controller.

[0088] The GPGPU 270 includes multiple cache memories, which include the L2 cache 253, the L1 cache 254, the instruction cache 255, and the shared memory 256. At least a part of the shared memory 256 can also be partitioned as a cache memory. The GPGPU 270 also includes multiple computing units 260A - 260N. Each computing unit 260A - 260N includes a set of vector registers 261, a set of scalar registers 262, a set of vector logic units 263, and a set of scalar logic units 264. The computing units 260A - 260N may also include a local shared memory 265 and a program counter 266. The computing units 260A - 260N can be coupled to a constant cache 267, which can be used to store constant data, that is, data that does not change during the execution of a kernel program or a shader program on the GPGPU 270. In one embodiment, the constant cache 267 is a scalar data cache, and the cached data can be directly fetched into the scalar register 262.

[0089] During operation, one or more CPUs 246 can write commands into registers in the GPGPU 270, or into memory in the GPGPU 270 that has been mapped to an accessible address space. The command processor 257 can read commands from the registers or the memory and determine how those commands will be processed within the GPGPU 270. Subsequently, the thread dispatcher 258 can be used to dispatch threads to the computing units 260A - 260N to execute those commands. Each computing unit 260A - 260N can execute threads independently of other computing units. In addition, each computing unit 260A - 260N can be independently configured for conditional computing and can conditionally output the results of the computing to the memory. When the submitted commands are completed, the command processor 257 can interrupt one or more CPUs 246.

[0090] Figures 3A - 3C A block diagram of an additional graphics processor and computing accelerator architecture provided by the embodiments described herein. Figures 3A - 3C Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0091] Figure 3Ais a block diagram of a graphics processor 300, which can be a discrete graphics processing unit or can be a graphics processor integrated with multiple processing cores or other semiconductor devices, such as but not limited to memory devices or network interfaces. In some embodiments, the graphics processor communicates via a memory-mapped I / O interface to registers on the graphics processor and uses commands placed in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0092] In some embodiments, the graphics processor 300 also includes a display controller 302 for driving display output data to a display device 318. The display controller 302 includes hardware for one or more overlay planes for the display and the composition of multiple layers of video or user interface elements. The display device 318 can be an internal or external display device. In one embodiment, the display device 318 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, which one or more media encoding formats include but are not limited to: Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9), and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG, and Motion JPEG (MJPEG) formats).

[0093] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterizer operations, including, for example, bit-block transfers. However, in one embodiment, one or more components of the graphics processing engine (GPE) 310 are used to perform 2D graphics operations. In some embodiments, the GPE 310 is a computational engine for performing graphics operations, which include three-dimensional (3D) graphics operations and media operations.

[0094] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed-function elements that perform various tasks within the elements and / or generate execution threads to the 3D / media subsystem 315. Although the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 that is dedicated to performing media operations, such as video post-processing and image enhancement.

[0095] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units for performing one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of, or on behalf of, the video codec engine 306. In some embodiments, the media pipeline 316 additionally includes a thread generation unit for generating threads for execution on the 3D / media subsystem 315. The generated threads execute computations for media operations on one or more graphics execution units included in the 3D / media subsystem 315.

[0096] In some embodiments, the 3D / media subsystem 315 includes logic for executing the threads generated by the 3D pipeline 312 and the media pipeline 316. In some embodiments, the pipelines send thread execution requests to the 3D / media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching various requests for available thread execution resources. The execution resources include an array of graphics execution units for processing 3D threads and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes a shared memory for sharing data between threads and for storing output data, which includes registers and addressable memory.

[0097] Figure 3BFIG. illustrates a graphics processor 320 according to an embodiment described herein, the graphics processor 320 having a tiled architecture. In one embodiment, the graphics processor 320 includes a graphics processing engine cluster 322, the graphics processing engine cluster 322 having multiple instances of graphics processing engines 310 within graphics engine tiles 310A - 310D. Each graphics engine tile 310A - 310D may be interconnected via a set of tile interconnects 323A - 323F. Each graphics engine tile 310A - 310D may also be connected to a memory module or memory device 326A - 326D via a memory interconnect 325A - 325D. The memory devices 326A - 326D may use any graphics memory technology. For example, the memory devices 326A - 326D may be Graphics Double Data Rate (GDDR) memory. In one embodiment, the memory devices 326A - 326D are High - Bandwidth Memory (HBM) modules, and these HBM modules may be on - die with their respective graphics engine tiles 310A - 310D. In one embodiment, the memory devices 326A - 326D are stacked memory devices that may be stacked on top of their respective graphics engine tiles 310A - 310D. In one embodiment, each graphics engine tile 310A - 310D and the associated memory 326A - 326D reside on separate dies, and these separate dies are bonded to a base die or a base substrate, as further described in detail in Figure 3A The graphics processor 320 may be configured with a Non - Uniform Memory Access (NUMA) system in which the memory devices 326A - 326D are coupled to the associated graphics engine tiles 310A - 310D. A given memory device may be accessed by a graphics engine tile different from the graphics engine tile directly connected to the memory device. However, when accessing the local tile, the access latency to the memory devices 326A - 326D may be the lowest. In one embodiment, a Cache - Coherent NUMA (ccNUMA) system is enabled, the ccNUMA system using the tile interconnects 323A - 323F to enable communication between cache controllers within the graphics engine tiles 310A - 310D in order to maintain a consistent memory image when more than one cache stores the same memory location. Figures 11B - 11D as further described in detail in

[0098] as further described in detail in

[0099] The graphics processing engine cluster 322 can be connected to an on-chip or on-package fabric interconnect 324. The fabric interconnect 324 can enable communication between the graphics engine dies 310A - 310D and components such as the video codec 306 and one or more copy engines 304. The copy engines 304 can be used to move data out of the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory), move data into the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory), and move data between the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory). The fabric interconnect 324 can also be used to interconnect the graphics engine dies 310A - 310D. The graphics processor 320 can optionally include a display controller 302 for enabling connection to an external display device 318. The graphics processor can also be configured as a graphics or computing accelerator. In the accelerator configuration, the display controller 302 and the display device 318 can be omitted.

[0100] The graphics processor 320 can be connected to a host system via a host interface 328. The host interface 328 can enable communication between the graphics processor 320, system memory, and / or other system components. The host interface 328 can be, for example, a PCI Express bus or another type of host system interface.

[0101] Figure 3C Illustrated is a computing accelerator 330 according to an embodiment described herein. The computing accelerator 330 can include an architectural similarity to the Figure 3B graphics processor 320 and is optimized for computing acceleration. The compute engine cluster 332 can include a set of compute engine dies 340A - 340D that includes execution logic optimized for parallel or vector-based general computing operations. In some embodiments, the compute engine dies 340A - 340D do not include fixed-function graphics processing logic, but in one embodiment, one or more of the compute engine dies 340A - 340D can include logic for performing media acceleration. The compute engine dies 340A - 340D can be connected to the memories 326A - 326D via memory interconnects 325A - 325D. The memories 326A - 326D and the memory interconnects 325A - 325D can be of a similar technology as in the graphics processor 320 or can be a different technology. The graphics compute engine dies 340A - 340D can also be interconnected via a set of die interconnects 323A - 323F and can be connected to and / or interconnected by the fabric interconnect 324. In one embodiment, the computing accelerator 330 includes a large L3 cache 336 that can be configured as a device-wide cache. The computing accelerator 330 can also operate in a manner similar toFigure 3B is connected to the host processor and memory via a host interface 328 in a manner similar to that of the graphics processor 320. Graphics Processing Engine

[0102] Figure 4 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is Figure 3A a certain version of the GPE 310 shown in Figure 3B and may also represent Figure 4 the graphics engine slices 310A - 310D of Figure 3A Elements with the same reference numerals (or names) as elements in any other figure herein can operate or run in any manner similar to the ways described elsewhere herein, but are not limited thereto. For example,

[0103] In some embodiments, the GPE 410 is coupled to or includes a command stream converter 403 that provides a command stream to the 3D pipeline 312 and / or the media pipeline 316. In some embodiments, the command stream converter 403 is coupled to a memory, which may be a system memory, or one or more of an internal cache memory and a shared cache memory. In some embodiments, the command stream converter 403 receives commands from the memory and sends the commands to the 3D pipeline 312 and / or the media pipeline 316. The commands are indications fetched from a ring buffer that stores commands for the 3D pipeline 312 and the media pipeline 316. In one embodiment, the ring buffer may additionally include a batch command buffer that stores batches of multiple commands. The commands for the 3D pipeline 312 may also include references to data stored in the memory, such as, but not limited to, vertex data and geometry data for the 3D pipeline 312 and / or image data and memory objects for the media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process commands and data by performing operations via logic within the respective pipelines or by dispatching one or more execution threads to the graphics core cluster 414. In one embodiment, the graphics core array 414 includes one or more graphics core blocks (e.g., (one or more) graphics cores 415A, (one or more) graphics cores 415B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources that includes: general and graphics-specific execution logic for performing graphics operations and compute operations; and fixed-function texture processing logic and / or machine learning and artificial intelligence acceleration logic.

[0104] In embodiments, the 3D pipeline 312 may include fixed-function and programmable logic for processing one or more shader programs by processing instructions and dispatching execution threads to the graphics core array 414, the one or more shader programs such as, vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core array 414 provides a unified block of execution resources for use in processing these shader programs. The multi-functional execution logic (e.g., execution units) within the (one or more) graphics cores 415A - 415B of the graphics core array 414 includes support for various 3D API shader languages and may execute multiple synchronized execution threads associated with multiple shaders.

[0105] In some embodiments, the graphics core array 414 includes execution logic for performing media functions such as video and / or image processing. In one embodiment, in addition to graphics processing operations, the execution units also include general-purpose logic programmable to perform parallel general-purpose compute operations. The general-purpose logic may perform operations in parallel or in combinationFigure 1 one or more of the processor cores 107 or general logic within cores 202A - 202N as in Figure 2A to perform processing operations.

[0106] Output data generated by threads executing on the graphics core array 414 can output data to memory in a unified return buffer (URB) 418. The URB 418 can store data for multiple threads. In some embodiments, the URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, the URB 418 can additionally be used for synchronization between threads on the graphics core array and fixed function logic within the shared function logic 420.

[0107] In some embodiments, the graphics core array 414 is scalable such that the array includes a variable number of graphics cores, each having a variable number of graphics cores based on the target power and performance levels of the GPE 410. In one embodiment, the execution resources are dynamically scalable such that the execution resources can be enabled or disabled.

[0108] The graphics core array 414 is coupled to the shared function logic 420, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide specialized complementary functions to the graphics core array 414. In embodiments, the shared function logic 420 includes, but is not limited to, sampler 421 logic, math 422 logic, and inter - thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.

[0109] Implement shared functionality at least in cases where the demand for a given specialized function is not sufficient to be included within the graphics core array 414. Instead, a single instantiation of that specialized function is implemented as a stand-alone entity within the shared functionality logic 420 and is shared among the execution resources within the graphics core array 414. The exact set of functions that are shared among and included within the graphics core arrays 414 varies from embodiment to embodiment. In some embodiments, specific shared functions that are widely used by the graphics core arrays 414 within the shared functionality logic 420 may be included within the shared functionality logic 416 within the graphics core arrays 414. In various embodiments, the shared functionality logic 416 within the graphics core arrays 414 may include some or all of the logic within the shared functionality logic 420. In one embodiment, all of the logic elements within the shared functionality logic 420 may be replicated within the shared functionality logic 416 of the graphics core arrays 414. In one embodiment, the shared functionality logic 420 is excluded in favor of the shared functionality logic 416 within the graphics core arrays 414. Execution Unit

[0110] Figures 5A - 5B Illustrated is thread execution logic 500 according to an embodiment described herein, the thread execution logic 500 including an array of processing elements employed within a graphics processor core. Figures 5A - 5B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. Figures 5A - 5B An overview of the thread execution logic 500 is illustrated, the thread execution logic 500 which may represent hardware logic illustrated in each of sub-cores 221A - 221F in Figure 2B Each sub-core 221A - 221F in Figure 5A represents execution units within a general-purpose graphics processor, while Figure 5B represents execution units that may be used within a compute accelerator.

[0111] As in Figure 5AAs illustrated, in some embodiments, the thread execution logic 500 includes a shader processor 502, a thread dispatcher 504, an instruction cache 506, a scalable execution unit array including a plurality of execution units 508A - 508N, a sampler 510, a shared local memory 511, a data cache 512, and a data port 514. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 508A, 508B, 508C, 508D, up to 508N - 1 and 508N) based on the computational requirements of the workload. In one embodiment, the included components are interconnected via an interconnect structure that links to each of the components. In some embodiments, the thread execution logic 500 includes one or more connections to memory (such as system memory or cache memory) through the instruction cache 506, the data port 514, the sampler 510, and one or more of the execution units 508A - 508N. In some embodiments, each execution unit (e.g., 508A) is an independent programmable general - purpose computing unit capable of executing multiple synchronized hardware threads and processing multiple data elements in parallel for each thread. In embodiments, the array of execution units 508A - 508N is scalable to include any number of individual execution units.

[0112] In some embodiments, the execution units 508A - 508N are primarily used to execute shader programs. The shader processor 502 can process various shader programs and can dispatch execution threads associated with the shader programs via the thread dispatcher 504. In one embodiment, the thread dispatcher includes logic for arbitrating requests for threads initiated from the graphics pipeline and the media pipeline and instantiating the requested threads on one or more of the execution units 508A - 508N. For example, the geometry pipeline can dispatch vertex shaders, tessellation shaders, or geometry shaders to the thread execution logic for processing. In some embodiments, the thread dispatcher 504 can also process runtime thread generation requests from executed shader programs.

[0113] In some embodiments, execution units 508A - 508N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct 3D and OpenGL) to be executed with minimal translation. These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). Each of the execution units 508A - 508N is capable of multi - issue single - instruction multiple - data (SIMD) execution, and multi - threading operations enable an efficient execution environment in the face of higher - latency memory accesses. Each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. Execution is multi - issued per clock for pipelines that can perform integer operations, single - precision floating - point operations, and double - precision floating - point operations, that can have SIMD branch capabilities, that can perform logical operations, that can perform transcendental operations, and that can perform other miscellaneous operations. While waiting for data from one of the shared functions in memory or a shared function, dependency logic within execution units 508A - 508N puts the waiting threads to sleep until the requested data has been returned. While the waiting threads are sleeping, the hardware resources can be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution units can perform operations for pixel shaders, fragment shaders, or another type of shader program including a different vertex shader. Embodiments can be applied to use execution that utilizes single - instruction multiple - threads (SIMT), as an alternative to the use of SIMD, or as an addition to the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to combinations of SIMD and SIMT.

[0114] Each of the execution units 508A - 508N operates on an array of data elements. The number of data elements is the "execution size", or the number of channels for the instruction. Execution channels are the logical units for data element access, masking, and flow - control execution within an instruction. The number of channels can be independent of the number of physical arithmetic logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In some embodiments, execution units 508A - 508N support integer and floating - point data types.

[0115] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution unit will process the various elements based on the data size of the elements. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four individual 64-bit packed data elements (Quad-Word (QW) size data elements), eight individual 32-bit packed data elements (Double Word (DW) size data elements), sixteen individual 16-bit packed data elements (Word (W) size data elements), or thirty-two individual 8-bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible.

[0116] In one embodiment, one or more execution units can be combined into fused execution units 509A - 509N, which have thread control logic (507A - 507N) common to the fused EUs. Multiple EUs can be fused into an EU group. Each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary according to the embodiment. Additionally, various SIMD widths can be executed per EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 509A - 509N includes at least two execution units. For example, fused execution unit 509A includes a first EU 508A, a second EU 508B, and thread control logic 507A common to the first EU 508A and the second EU 508B. Thread control logic 507A controls the threads executed on the fused graphics execution unit 509A, allowing each EU within the fused execution units 509A - 509N to execute using a common instruction pointer register.

[0117] One or more internal instruction caches (e.g., 506) are included in the thread execution logic 500 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 512) are included to cache thread data during thread execution. Threads executing on the execution logic 500 can also store explicitly managed data in the shared local memory 511. In some embodiments, a sampler 510 is included to provide texture sampling for 3D operations and media sampling for media operations. In some embodiments, the sampler 510 includes specialized texture or media sampling functions for processing texture data or media data during the sampling process before providing the sampled data to the execution unit.

[0118] It should be noted that there is an error in the original text where it mentions "four individual 54-bit packed data elements" which should likely be "four individual 64-bit packed data elements" for the sake of correct arithmetic when dealing with a 256-bit vector divided into four parts. The translation has been made with this correction in mind.During execution, the graphics pipeline and the media pipeline send thread launch requests to the thread execution logic 500 via thread generation and dispatch logic. Once a group of geometric objects has been processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 502 is called to further compute output information and cause the results to be written to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes the values of various vertex attributes, and the values of the various vertex attributes are interpolated across the rasterized objects. In some embodiments, the pixel processor logic within the shader processor 502 then executes a pixel shader program or a fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 502 dispatches threads to execution units (e.g., 508A) via the thread dispatcher 504. In some embodiments, the shader processor 502 uses texture sampling logic in the sampler 510 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometric data compute pixel color data for each geometric fragment, or discard one or more pixels without further processing.

[0119] In some embodiments, the data port 514 provides a memory access mechanism for the thread execution logic 500 to output processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, the data port 514 includes or is coupled to one or more cache memories (e.g., the data cache 512) to cache data for memory access via the data port.

[0120] In one embodiment, the execution logic 500 may further include a ray tracer 505 that can provide ray tracing acceleration functionality. The ray tracer 505 may support a ray tracing instruction set that includes instructions / functions for ray generation. The ray tracing instruction set may be similar to or different from the ray tracing instruction set supported by Figure 2C the ray tracing core 245 therein.

[0121] Figure 5BFIG. illustrates example internal details of execution unit 508 according to an embodiment. The graphics execution unit 508 may include an instruction fetch unit 537, a general register file array (GRF) 524, an architectural register file array (ARF) 526, a thread arbiter 522, a dispatch unit 530, a branch unit 532, a set of SIMD floating point units (FPUs) 534, and in one embodiment, a set of dedicated integer SIMD ALUs 535. The GRF 524 and the ARF 526 include a set of general register files and architectural register files associated with each synchronous hardware thread that can be active in the graphics execution unit 508. In one embodiment, the per-thread architectural state is maintained in the ARF 526, while data used during thread execution is stored in the GRF 524. The execution state of each thread (including the instruction pointer for each thread) may be saved in thread-specific registers in the ARF 526.

[0122] In one embodiment, the graphics execution unit 508 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and fine-grained Interleaved Multi-Threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronous threads and the number of registers per execution unit, where execution unit resources are divided across the logic for executing multiple synchronous threads. The number of logical threads that can be executed by the graphics execution unit 508 is not limited to the number of hardware threads, and multiple logical threads may be assigned to each hardware thread.

[0123] In one embodiment, the graphics execution unit 508 may issue multiple instructions in concert, and these instructions may each be different instructions. The thread arbiter 522 of the graphics execution unit thread 508 may dispatch the instructions to one of the following for execution: the send unit 530, the branch unit 532, or the (one or more) SIMD FPUs 534. Each execution thread may access 128 general-purpose registers within the GRF 524, where each register may store 32 bytes that can be accessed as a SIMD 8-element vector with 32-bit data elements. In one embodiment, each execution unit thread has access to 4 kilobytes within the GRF 524, but the embodiments are not limited thereto, and more or fewer register resources may be provided in other embodiments. In one embodiment, the graphics execution unit 508 is partitioned into seven hardware threads that can independently perform computational operations, but the number of threads per execution unit may also vary according to the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4 kilobytes, the GRF 524 may store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 524 may store a total of 64 kilobytes. Flexible addressing modes may permit addressing registers together, thereby effectively creating wider registers or representing strided rectangular block data structures.

[0124] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the messaging send unit 530. In one embodiment, branch instructions are dispatched to a dedicated branch unit 532 to facilitate SIMD scatter and eventual gather.

[0125] In one embodiment, the graphics execution unit 508 includes one or more SIMD floating-point units (FPUs) 534 for performing floating-point operations. In one embodiment, the (one or more) FPUs 534 also support integer computations. In one embodiment, the (one or more) FPUs 534 may SIMD execute up to a maximum number M of 32-bit floating-point (or integer) operations, or SIMD execute up to 2M 16-bit integer or 16-bit floating-point operations. In one embodiment, at least one of the (one or more) FPUs provides extended mathematical capabilities that support high-throughput transcendental mathematical functions and double-precision 64-bit floating-point. In some embodiments, a set 535 of 8-bit integer SIMD ALUs also exists and may be specifically optimized to perform operations associated with machine learning computations.

[0126] In one embodiment, an array of multiple instances of the graphics execution unit 508 may be instantiated in a graphics sub-core grouping (e.g., sub-slice). For scalability, the product architect may select the exact number of execution units per sub-core grouping. In one embodiment, the execution unit 508 may execute instructions across multiple execution channels. In a further embodiment, each thread executed on the graphics execution unit 508 is executed on a different channel.

[0127] Figure 6 FIG. illustrates an additional execution unit 600 according to an embodiment. The execution unit 600 may be a compute-optimized execution unit for use in, for example, Figure 3C compute engine slices 340A - 340D, but is not limited thereto. Variants of the execution unit 600 may also be used in Figure 3B graphics engine slices 310A - 310D. In one embodiment, the execution unit 600 includes a thread control unit 601, a thread state unit 602, an instruction fetch / prefetch unit 603, and an instruction decoding unit 604 (also referred to herein as a decoder). The execution unit 600 additionally includes a register file 606 that stores registers that may be assigned to hardware threads within the execution unit. The execution unit 600 additionally includes a dispatch unit 607 and a branch unit 608. In one embodiment, the dispatch unit 607 and the branch unit 608 can operate in a manner similar to Figure 5B the dispatch unit 530 and the branch unit 532 of the graphics execution unit 508.

[0128] The execution unit 600 further includes a computing unit 610, and the computing unit 610 includes a plurality of different types of functional units. In one embodiment, the computing unit 610 includes an ALU unit 611, and the ALU unit 611 includes an array of arithmetic logic units. The ALU unit 611 can be configured to perform 64-bit, 32-bit, and 16-bit integer and floating-point operations. Integer and floating-point operations can be performed simultaneously. The computing unit 610 may further include a systolic array 612 and a math unit 613. The systolic array 612 includes a network of data processing units that is W wide and D deep, and it can be used to perform vector or other data parallel operations in a systolic manner. In one embodiment, the systolic array 612 can be configured to perform matrix operations (such as matrix dot product operations). In one embodiment, the systolic array 612 supports 16-bit floating-point operations, as well as 8-bit and 4-bit integer operations. In one embodiment, the systolic array 612 can be configured to accelerate machine learning operations. In such embodiments, the systolic array 612 can be configured with support for the bfloat 16-bit floating-point format. In one embodiment, the math unit 613 may be included to perform a specific subset of math operations in an efficient and lower-power manner than the ALU unit 611. The math unit 613 may include a variant of the math logic (e.g., the math logic 422 of the shared function logic 420 in Figure 4 the shared function logic 420) that can be found in the shared function logic of the graphics processing engine provided by other embodiments. In one embodiment, the math unit 613 can be configured to perform 32-bit and 64-bit floating-point operations.

[0129] The thread control unit 601 includes logic for controlling the execution of threads within the execution unit. The thread control unit 601 may include thread arbitration logic for starting, stopping, and preempting the execution of threads within the execution unit 600. The thread state unit 602 can be used to store the thread state of the threads assigned to execute on the execution unit 600. Storing the thread state within the execution unit 600 enables those threads to be quickly preempted when they become blocked or idle. The instruction fetch / prefetch unit 603 can fetch instructions from the instruction cache of a higher-level execution logic (e.g., the instruction cache 506 in Figure 5A ). The instruction fetch / prefetch unit 603 can also issue a prefetch request for instructions to be loaded into the instruction cache based on an analysis of the current execution thread. The instruction decoding unit 604 can be used to decode the instructions to be executed by the computing unit. In one embodiment, the instruction decoding unit 604 can be used as a secondary decoder to decode complex instructions into constituent micro-operations.

[0130] Execution unit 600 additionally includes register file 606, which can be used by hardware threads executing on execution unit 600. The registers in register file 606 can be partitioned across the logic for multiple synchronous threads within compute units 610 that execute the computations of execution unit 600. The number of logical threads that can be executed by graphics execution unit 600 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. Based on the number of supported hardware threads, the size of register file 606 can vary across embodiments. In one embodiment, register renaming can be used to dynamically allocate registers to hardware threads.

[0131] Figure 7 is a block diagram of an illustrative graphics processor instruction format 700 according to some embodiments. In one or more embodiments, a graphics processor execution unit supports an instruction set with instructions in multiple formats. The solid boxes illustrate components that are typically included in execution unit instructions, while the dashed boxes include optional or components that are only included in a subset of the instructions. In some embodiments, the described and illustrated instruction format 700 is a macro-instruction because they are the instructions supplied to the execution unit, as opposed to micro-operations that result from instruction decoding once the instructions are processed.

[0132] In some embodiments, the graphics processor execution unit natively supports instructions in a 128-bit instruction format 710. Based on the selected instructions, instruction options, and number of operands, a 64-bit compact instruction format 730 can be used for some instructions. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The native instructions available in the 64-bit format 730 vary across embodiments. In some embodiments, a set of index values in index field 713 is used to partially compress the instructions. The execution unit hardware references a set of compression tables based on the index values and uses the compression table outputs to reconstruct the native instructions in 128-bit instruction format 710. Instructions of other sizes and formats can be used.

[0133] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 710, the execution size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.

[0134] Some execution unit instructions have up to three operands, including two source operands src0 720, src1 722, and one destination 718. In some embodiments, the execution unit supports dual-destination instructions, where one of the destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction may be an immediate (e.g., hard-coded) value passed with the instruction.

[0135] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that, for example, specifies whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are directly provided by bits in the instruction.

[0136] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that specifies the addressing mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment of the instruction. Some embodiments support access modes including a 16-byte alignment access mode and a 1-byte alignment access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in the first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in the second mode, the instruction may use 16-byte-aligned addressing for all source and destination operands.

[0137] In one embodiment, the addressing mode portion of the access / addressing mode field 726 determines whether the instruction is to use direct addressing or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.

[0138] In some embodiments, the instructions are grouped based on the opcode 712 bit field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of the opcode. The exact opcode grouping shown is only an example. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 share the five most significant bits (MSB), where the move (mov) instruction takes the form of 0000xxxxb, and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jump) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 748 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel math group 748 performs arithmetic operations in parallel across data channels. The vector math group 750 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In one embodiment, the illustrated opcode decoding 740 can be used to determine which part of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by the systolic array. Other instructions (such as ray tracing instructions (not shown)) can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic. Graphics Pipeline

[0139] Figure 8 is a block diagram of another embodiment of the graphics processor 800. Figure 8 Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0140] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a render output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued through the ring interconnect 802 to the graphics processor 800. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. The command stream converter 803 interprets commands from the ring interconnect 802 and supplies instructions to the various components of the geometry pipeline 820 or the media pipeline 830.

[0141] In some embodiments, the command stream converter 803 directs the operation of the vertex fetcher 805, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 803. In some embodiments, the vertex fetcher 805 provides vertex data to the vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, the vertex fetcher 805 and the vertex shader 807 execute vertex processing instructions by dispatching execution threads to the execution units 852A - 852B via the thread dispatcher 831.

[0142] In some embodiments, the execution units 852A - 852B are an array of vector processors having instruction sets for performing graphics operations and media operations. In some embodiments, the execution units 852A - 852B have attached L1 caches 851 dedicated to each array or shared between the arrays. The caches can be configured as data caches, instruction caches, or a single cache partitioned to contain data and instructions in different partitions.

[0143] In some embodiments, the geometry pipeline 820 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some embodiments, the programmable hull shader 811 configures the tessellation operation. The programmable domain shader 817 provides backend evaluation of the tessellation output. The tessellator 813 operates under the direction of the hull shader 811 and includes dedicated logic for generating a set of detailed geometric objects based on a rough geometric model that is provided as input to the geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation components (e.g., the hull shader 811, the tessellator 813, and the domain shader 817) can be bypassed. The tessellation components can operate based on data received from the vertex shader 807.

[0144] In some embodiments, the complete geometric object can be processed by the geometry shader 819 via one or more threads dispatched to execution units 852A - 852B, or can proceed directly to the clipper 829. In some embodiments, the geometry shader operates on the entire geometric object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 is programmable by a geometry shader program to perform geometric tessellation in cases where the tessellation unit is disabled.

[0145] Before rasterization, the clipper 829 processes vertex data. The clipper 829 can be a fixed - function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer and depth test component 873 in the render output pipeline 870 dispatch the pixel shader to convert the geometric object into a per - pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, the application can bypass the rasterizer and depth test component 873 and access the un - rasterized vertex data via the outflow unit 823.

[0146] The graphics processor 800 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some embodiments, the execution units 852A - 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via a data port 856 to perform memory accesses and communicate with the render output pipeline components of the processor. In some embodiments, the sampler 854, caches 851, 858, and execution units 852A - 852B each have separate memory access paths. In one embodiment, the texture cache 858 can also be configured as a sampler cache.

[0147] In some embodiments, the rendering output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects to associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. In some embodiments, associated render cache 878 and depth cache 879 are also available. Pixel operation component 877 performs pixel-based operations on the data, but in some instances, pixel operations associated with 2D operations (e.g., bit blit image transfer with blending) are performed by 2D engine 841 or, at display time, by display controller 843 using an overlay display plane instead. In some embodiments, shared L3 cache 875 is available to all graphics components, allowing data to be shared without using the main system memory.

[0148] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front end 834 processes the media commands before sending them to media engine 837. In some embodiments, media engine 837 includes a thread generation function for generating threads to be dispatched via thread dispatcher 831 to thread execution logic 850.

[0149] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and can be coupled to the graphics processor via ring interconnect 802, or some other interconnect bus or fabric. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 contains dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which can be a system-integrated display device such as in a laptop computer or an external display device attached via a display device connector.

[0150] In some embodiments, the geometry pipeline 820 and the media pipeline 830 may be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). In some embodiments, driver software for the graphics processor converts API calls that are specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. In some embodiments, support may also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the pipelines of the future APIs to the pipelines of the graphics processor. Graphics Pipeline Programming

[0151] Figure 9A is a block diagram illustrating a graphics processor command format 900 according to some embodiments. Figure 9B is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 9A The solid boxes in generally represent components that are typically included in a graphics command, while the dashed boxes include optional or components that are only included in a subset of graphics commands. Figure 9A An example graphics processor command format 900 includes data fields for a client 902 that identifies the command, a command operation code (opcode) 904, and data 906. A sub-opcode 905 and a command size 908 are also included in some commands.

[0152] In some embodiments, client 902 designates the client unit of the graphics device that processes command data. In some embodiments, the graphics processor command parser examines the client field of each command to adjust further processing of the command and routes the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 904 and sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses the information in the data field 906 to execute the command. For some commands, an explicit command size 908 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands in the command based on the command opcode. In some embodiments, the commands are aligned by multiples of a double word. Other command formats may be used.

[0153] Figure 9B The flowchart example in FIG. 910 shows a graphics processor command sequence. In some embodiments, the software or firmware of a data processing system characterized by an embodiment of the graphics processor uses a version of the shown command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only, as embodiments are not limited to these specific commands or this command sequence. Additionally, commands may be issued in the command sequence as a batch of commands such that the graphics processor will process the command sequence in at least a partially concurrent manner.

[0154] In some embodiments, the graphics processor command sequence 910 can start with a pipeline flush clear command 912 to cause any active graphics pipeline to complete the current outstanding commands for the pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate concurrently. Executing the pipeline flush clear causes the active graphics pipeline to complete any outstanding commands. In response to the pipeline flush clear, the command parser for the graphics processor will pause command processing until the active drawing engine completes the outstanding operations and the associated read cache is invalidated. Optionally, any data marked "dirty" in the render cache may be flushed to memory. In some embodiments, the pipeline flush clear command 912 can be used for pipeline synchronization or can be used before placing the graphics processor in a low power state.

[0155] In some embodiments, the pipeline select command 913 is used when a command sequence explicitly switches between pipelines using a graphics processor. In some embodiments, the pipeline select command 913 is utilized once in the execution context before issuing pipeline commands, unless the context is issuing commands for both pipelines. In some embodiments, the pipeline dump clear command 912 is utilized immediately before the pipeline switch via the pipeline select command 913.

[0156] In some embodiments, the pipeline control command 914 configures the graphics pipeline for operation and programs the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state for the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands.

[0157] In some embodiments, the return buffer status command 916 is used to configure the set of return buffers for a corresponding pipeline for writing data. Some pipeline operations utilize the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer status 916 includes the size and number of return buffers to select for a set of pipeline operations.

[0158] The remaining commands in the command sequence differ based on the active pipeline for operation. Based on the pipeline determination 920, the command sequence is customized for the 3D pipeline 922 starting in the 3D pipeline state 930 or the media pipeline 924 starting at the media pipeline state 940.

[0159] Commands for configuring the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the particular 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass certain pipeline elements if they will not be used.

[0160] In some embodiments, the 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 commands are used to perform vertex operations on the 3D primitives via the vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.

[0161] In some embodiments, the 3D pipeline 922 is triggered via the execution of 934 commands or events. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a "go" or "kick" command in a command sequence. In some embodiments, command execution uses pipeline synchronization commands to trigger a command sequence dump clearance through the graphics pipeline. The 3D pipeline will perform geometric processing on the 3D primitives. Once the operation is complete, the resulting geometric object is rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands may also be included for controlling pixel coloring and pixel backend operations.

[0162] In some embodiments, when performing media operations, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the specific uses and ways of programming the media pipeline 924 depend on the media or compute operations to be performed. During media decoding, specific media decoding operations may be migrated to the media pipeline. In some embodiments, the media pipeline may also be bypassed, and media decoding may be performed entirely or partially using the resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0163] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. A set of commands for configuring the media pipeline state 940 is dispatched or placed into the command sequence before the media object commands 942. In some embodiments, the commands for the media pipeline state 940 include data for configuring the media pipeline elements that will be used to process the media objects. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands for the media pipeline state 940 also support the use of one or more pointers to "indirect" state elements that point to a batch of state settings.

[0164] In some embodiments, the media object command 942 supplies a pointer to a media object for processing by the media pipeline. The media object includes a memory buffer that contains video data to be processed. In some embodiments, all media pipeline states should be valid before issuing the media object command 942. Once the pipeline states are configured and the media object command 942 is queued, the media pipeline 924 is triggered via an execute command 944 or an equivalent execution event (e.g., a register write). Subsequently, the output from the media pipeline 924 can be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations. Graphics Software Architecture

[0165] Figure 10 FIG. illustrates an example graphics software architecture for a data processing system 1000 according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.

[0166] In some embodiments, the 3D graphics application 1010 includes one or more shader programs that include shader instructions 1012. The shader language instructions may be in a high-level shader language such as the High-Level Shader Language (HLSL) of Direct3D, the OpenGL Shader Language (GLSL), and so on. The application also includes executable instructions 1014 in machine language suitable for execution by the general-purpose processor cores 1034. The application also includes a graphics object 1016 defined by vertex data.

[0167] In some embodiments, the operating system 1020 is from Microsoft Corporation An operating system, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. The operating system 1020 can support a graphics API 1022, such as, the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation can be just-in-time (JIT) compilation or application executable shader pre-compilation. In some embodiments, during the compilation of the 3D graphics application 1010, high-level shaders are compiled into low-level shaders. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.

[0168] In some embodiments, the user-mode graphics driver 1026 includes a backend shader compiler 1027 to compile the shader instructions 1012 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses the operating system kernel-mode function 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions. IP Core Implementation Method

[0169] One or more aspects of at least one embodiment can be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium can include instructions that represent various logics within the processor. In some embodiments, the machine-readable medium is also referred to herein as a computer-readable medium or a non-transitory computer-readable medium. When read by a machine, the instructions can cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit, and these reusable units can be stored on a tangible, machine-readable medium as a hardware model that describes the organization of the integrated circuit. The hardware model can be supplied to various customers or manufacturing facilities that load the hardware model on a manufacturing machine for manufacturing an integrated circuit. The integrated circuit can be manufactured such that the circuit performs the operations described in association with any of the embodiments described herein.

[0170] Figure 11A FIG. is a block diagram of an IP core development system 1100 that can be used to fabricate an integrated circuit to perform operations according to an embodiment. The IP core development system 1100 can be used to generate a modular, reusable design that can be incorporated into a larger design or used to build an entire integrated circuit (e.g., a SOC integrated circuit). A design facility 1130 can generate a software simulation 1110 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 can include a functional simulation, a behavioral simulation, and / or a timing simulation. Subsequently, a register transfer level (RTL) design 1115 can be created or synthesized from the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic performed using the modeled digital signals). In addition to the RTL design 1115, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0171] The RTL design 1115 or an equivalent can be further synthesized by the design facility into a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted via a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The manufacturing facility 1165 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0172] Figure 11BFIG. 0 is a cross-sectional side view of an integrated circuit package assembly 1170 in accordance with some embodiments described herein. The integrated circuit package assembly 1170 illustrates an implementation of one or more processor or accelerator devices as described herein. The package assembly 1170 includes a plurality of hardware logic units 1172, 1174 coupled to a substrate 1180. The logic 1172, 1174 may be implemented at least in part in configurable logic or fixed-function logic hardware and may include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. Each logic unit 1172, 1174 may be implemented within a semiconductor die and is coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic 1172, 1174 and the substrate 1180 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 1172, 1174. In some embodiments, the substrate 1180 is an epoxy-based laminated substrate. In other embodiments, the substrate 1180 may include other suitable types of substrates. The package assembly 1170 may be connected to other electrical devices via package interconnects 1183. The package interconnects 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.

[0173] In some embodiments, the logic units 1172, 1174 are electrically coupled to a bridge 1182 that is configured to route electrical signals between the logic 1172 and the logic 1174. The bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. The bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide a chip-to-chip connection between the logic 1172 and the logic 1174.

[0174] Although two logic units 1172, 1174 and a bridge 1182 are illustrated, the embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges, as the bridge 1182 may be excluded when the logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges may be connected together in other possible configurations including three-dimensional configurations.

[0175] Figure 11CIllustrated is a packaged component 1190 that includes hardware logic die that connect to multiple units of a substrate 1180 (e.g., a base die). Graphics processing units, parallel processors, and / or compute accelerators as described herein can be composed of various silicon die fabricated separately. In this context, a die is an integrated circuit that is at least partially packaged and includes different logic units that can be assembled with other die into a larger package. Die with various collections of different IP core logic can be assembled into a single device. Additionally, die can be integrated into a base die or base die using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP cores can be fabricated using different process technologies and composed during fabrication, which avoids the complexity of converging multiple IPs into the same manufacturing process, especially for large SoCs with several styles of IP. Allowing the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Additionally, decomposed IP is more easily modified to be independently power gated, and components not in use for a given workload can be turned off, thereby reducing overall power consumption.

[0176] The hardware logic die can include dedicated hardware logic die 1172, logic or I / O die 1174, and / or memory die 1175. The hardware logic die 1172 and the logic or I / O die 1174 can be implemented at least partially in configurable logic or fixed function logic hardware and can include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory die 1175 can be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory.

[0177] Each die can be fabricated as a separate semiconductor die and coupled to the substrate 1180 via an interconnect fabric 1173. The interconnect fabric 1173 can be configured to route electrical signals between the various die and logic within the substrate 1180. The interconnect fabric 1173 can include interconnects such as, but not limited to, bumps or posts. In some embodiments, the interconnect fabric 1173 can be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic, I / O, and memory die.

[0178] In some embodiments, substrate 1180 is an epoxy-based laminated substrate. In other embodiments, substrate 1180 may include other suitable types of substrates. Package assembly 1190 may be connected to other electrical devices via package interconnect 1183. Package interconnect 1183 may be coupled to the surface of substrate 1180 to route electrical signals to other electrical devices, such as a motherboard, other chip sets, or a multi-chip module.

[0179] In some embodiments, logic or I / O die 1174 and memory die 1175 may be electrically coupled via bridge 1187, which is configured to route electrical signals between logic or I / O die 1174 and memory die 1175. Bridge 1187 may be a dense interconnect fabric that provides routing for electrical signals. Bridge 1187 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide die-to-die connections between logic or I / O die 1174 and memory die 1175. Bridge 1187 may also be referred to as a silicon bridge or an interconnect bridge. For example, in some embodiments, bridge 1187 is an Embedded Multi-die Interconnect Bridge (EMIB). In some embodiments, bridge 1187 may simply be a direct connection from one die to another die.

[0180] Substrate 1180 may include hardware components for I / O 1191, cache memory 1192, and other hardware logic 1193. Structure 1185 may be embedded in substrate 1180 to enable communication between various logic dies within substrate 1180 and logic 1191, 1193. In one embodiment, I / O 1191, structure 1185, cache, bridge, and other hardware logic 1193 may be integrated into a base die stacked on top of substrate 1180. Structure 1185 may be a on-chip network interconnect or another form of packet-switched fabric that exchanges data packets between components of the package assembly.

[0181] In various embodiments, the packaged component 1190 may include fewer or greater numbers of components and dies interconnected by a structure 1185 or one or more bridges 1187. The dies within the packaged component 1190 can be arranged in a 3D arrangement or a 2.5D arrangement. Generally, the bridge fabric 1187 can be used to facilitate point-to-point interconnections such as between logic or I / O dies and memory dies. The structure 1185 can be used to interconnect various logic and / or I / O dies (e.g., dies 1172, 1174, 1191, 1193) with other logic and / or I / O dies. In one embodiment, a cache memory 1192 within the substrate can act as a global cache for the packaged component 1190, act as part of a distributed global cache, or act as a dedicated cache for the structure 1185.

[0182] Figure 11D FIG. illustrates a packaged component 1194 including interchangeable dies 1195 according to an embodiment. The interchangeable dies 1195 can be assembled into standardized slots on one or more base dies 1196, 1198. The base dies 1196, 1198 can be coupled via a bridge interconnect 1197, which can be similar to other bridge interconnects described herein and can be, for example, an EMIB. Memory dies can also be connected to logic or I / O dies via a bridge interconnect. The I / O and logic dies can communicate via an interconnect structure. Each of the base dies can support one or more slots in a standardized format for one of logic or I / O or memory / cache.

[0183] In one embodiment, SRAM and power delivery circuitry can be fabricated into one or more of the base dies 1196, 1198, which can be fabricated using a different process technology relative to the interchangeable die 1195, with the interchangeable die 1195 stacked on top of the base die. For example, the base dies 1196, 1198 can be fabricated using a larger process technology while the interchangeable die can be fabricated using a smaller process technology. One or more of the interchangeable dies 1195 can be memory (e.g., DRAM) dies. Different memory densities can be selected for the packaged component 1194 based on the power and / or performance for the product using the packaged component 1194. Additionally, logic dies with different numbers of types of functional units can be selected based on the power and / or performance for the product at the time of assembly. Further, dies containing different types of IP logic cores can be inserted into the interchangeable die slots, enabling a hybrid processor design that can mix and match IP blocks of different technologies. Example System-on-Chip Integrated Circuit

[0184] Figures 12 - 13BFIG. illustrates an example integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0185] Figure 12 is a block diagram illustrating an example system-on-chip integrated circuit 1200 that can be fabricated using one or more IP cores according to an embodiment. The example integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210 and may additionally include an image processor 1215 and / or a video processor 1220, either of the image processor 1215 and the video processor 1220 can be a modular IP core from the same design facility or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235 and I 2 S / I 2 C controller 1240. In addition, the integrated circuit may include a display device 1245 that is coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display interface 1255. Storage may be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface may be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 1270.

[0186] Figures 13A - 13B is a block diagram illustrating an example graphics processor for use within a SoC according to embodiments described herein. Figure 13A illustrates an example graphics processor 1310 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. Figure 13B illustrates an additional example graphics processor 1340 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. Figure 13A The graphics processor 1310 is an example of a low-power graphics processor core. Figure 13B The graphics processor 1340 is an example of a higher-performance graphics processor core. Each of the graphics processors 1310, 1340 can be Figure 12 a variant of the graphics processor 1210.

[0187] As Figure 13A shown, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A - 1315N (e.g., 1315A, 1315B, 1315C, 1315D, up to 1315N - 1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic such that the vertex processor 1305 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1315A - 1315N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. The vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The (one or more) fragment processors 1315A - 1315N use the primitive data and vertex data generated by the vertex processor 1305 to produce a frame buffer that is displayed on a display device. In one embodiment, the (one or more) fragment processors 1315A - 1315N are optimized to execute fragment shader programs as provided in the OpenGL API, and these fragment shader programs can be used to perform operations similar to pixel shader programs as provided in the Direct 3D API.

[0188] The graphics processor 1310 additionally includes one or more memory management units (MMUs) 1320A - 1320B, (one or more) caches 1325A - 1325B, and (one or more) circuit interconnects 1330A - 1330B. The one or more MMUs 1320A - 1320B provide virtual - to - physical address mapping for the graphics processor 1310 (including for the vertex processor 1305 and / or the (one or more) fragment processors 1315A - 1315N), and this virtual - to - physical address mapping can reference vertex data or image / texture data stored in memory in addition to vertex data or image / texture data stored in the one or more caches 1325A - 1325B. In one embodiment, the one or more MMUs 1320A - 1320B can be synchronized with other MMUs within the system such that each processor 1205 - 1220 can participate in a shared or unified virtual memory system, and the other MMUs within the system include one or more MMUs associated with Figure 12 one or more application processors 1205, image processors 1215, and / or video processors 1220. According to an embodiment, the one or more circuit interconnects 1330A - 1330B enable the graphics processor 1310 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0189] As Figure 13B shown, the graphics processor 1340 includesFigure 13A One or more MMUs 1320A - 1320B, caches 1325A - 1325B, and circuit interconnects 1330A - 1330B of the graphics processor 1310. The graphics processor 1340 includes one or more shader cores 1355A - 1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N - 1 and 1355N), which provide a unified shader core architecture, where a single core or any type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary depending on the embodiment and implementation. Additionally, the graphics processor 1340 includes an inter-core task manager 1345, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1355A - 1355N and a tiling unit 1358 for accelerating tiling operations for tile-based rendering, in which rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or to optimize the use of internal caches.

[0190] In some embodiments, as described herein, processing resources represent processing elements (e.g., GPGPU cores, ray tracing cores, tensor cores, execution resources, execution units (EUs), stream processors, streaming multiprocessors (SMs), graphics multiprocessors) associated with a graphics processor structure in a graphics processor or GPU (e.g., parallel processing units, graphics processing engines, multi-core groups, computing units, computing units of the next graphics core). For example, a processing resource can be: a GPGPU core of a graphics multiprocessor, or one of the tensor / ray tracing cores; a ray tracing core, tensor core, or GPGPU core of a graphics multiprocessor; an execution resource of a graphics multiprocessor; one of the GFX cores, tensor cores, or ray tracing cores of a multi-core group; one of the vector logic units or scalar logic units of a computing unit; an execution unit with an EU array or an EU array; an execution unit of execution logic; and / or an execution unit. A processing resource can also be, for example, a graphics processing engine, a processing cluster, a GPGPU, a GPGPU, a graphics processing engine, a graphics processing engine cluster, and / or an execution resource within a graphics processing engine. A processing resource can also be a graphics processor, a graphics processor, and / or a processing resource within a graphics processor. Scalable Centralized Error Queue in Processing Architecture

[0191] Parallel computing is a type of computing in which the execution of many computations or processes is performed simultaneously. Parallel computing can take various forms, including but not limited to SIMD or SIMT. SIMD describes a computer with multiple processing elements that perform the same operation on multiple data points simultaneously. In one example, the Figures 5A - 5B refers to SIMD and its implementation in general-purpose processors in terms of EUs, FPUs, and ALUs. In a common SIMD machine, data is packed into registers, each of which contains an array of lanes. Instructions operate on the data found in lane n of one register and the data found in the same lane of another register. SIMD machines are advantageous in areas where a single instruction sequence can be applied simultaneously to a large amount of data. For example, in one embodiment, a graphics processing unit (e.g., GPGPU, GPU, etc.) can be used to perform SIMD vector operations using a compute shader program.

[0192] Embodiments can also be applied to use execution via single-instruction, multiple-threading (SIMT) as an alternative to, or in addition to, the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to a combination of SIMD and SIMT. The following description is discussed in terms of SIMD machines. However, the embodiments herein are not limited to applications in the SIMD context and can also be applied to other parallel computing paradigms, such as, for example, SIMT. For ease of discussion and explanation, the following description generally focuses on SIMD implementations. However, the embodiments can be similarly applied to SIMT machines without modifying the described techniques and methods. Regarding SIMT machines, a similar pattern as discussed below can be followed to provide instructions to a systolic array and execute the instructions on a SIMT machine. Other types of parallel computer machines can also utilize the embodiments herein.

[0193] Embodiments can implement a GPGPU with a matrix acceleration circuitry system. Such a matrix acceleration circuitry system can be used to accelerate machine learning (ML) operations. A GPGPU with a matrix acceleration circuitry system is often deployed and hosted in a data center. Hardware resilience is a requirement for graphics architectures in the data center market segment. Architectures with reliability, availability, and serviceability (RAS) features are designed to meet resilience goals. Reliability refers to how reliable the operation of the design is. Availability refers to the uptime of the operation that the design can still provide in the presence of errors. Serviceability refers to how easily the design can be serviced to resume reliable operation once an error occurs.

[0194] To do this, the design should be able to detect and record errors (to improve reliability), correct these errors if possible (to improve availability), and report to a higher-level system component (such as a driver) when the errors are uncorrectable. The system software can then take appropriate actions to service the errors and bring the design back to reliable operation.

[0195] This proposal introduces a scalable centralized error queue for RAS. A solution for centralizing fine-grained logs in a centralized unit of a graphics processor is provided herein. By centralizing the error queue in the implementation herein, the register space allocation for recording errors in the (one or more) leaf units can be avoided. In addition, since the central queue is common across sources and the updates are not exploited by the queue when a new leaf source is added, scalability is achieved. In some implementations, the scalable centralized error queue can also avoid the burden of implementing "sticky" registers within the leaf-level units, and these scalable centralized error queues survive across thermal resets.

[0196] In the implementation herein, a fine-grained, scalable error recording scheme is defined, in which units write their error information to a system graphics interface (SGI) unit (also more generally referred to herein as a system interface (e.g., when hosted on a non-graphics processing device)) through a message channel, and the SGI unit hosts at least one (one or more) centralized queue that is agnostic to the error source. The SGI can host a single centralized queue for each error type, or multiple centralized queues, including a centralized queue for correctable errors, a centralized queue for uncorrectable local errors (non-fatal errors, for reference only), and a centralized queue for uncorrectable global errors (fatal errors that require a reset operation).

[0197] As previously mentioned, the centralized queue can also implement the ability to survive from a thermal reset, which is referred to herein as "sticky". The term sticky (e.g., sticky queue) as used herein can refer to the ability of a component (such as a queue) to maintain the stored data in the queue (i.e., survive from a thermal reset) after a thermal reset of a graphics system. Thus, the centralized error logging scheme described herein is a scalable solution that provides both stickiness and scalability.

[0198] Figure 14FIG. illustrates an embodiment of a graphics processor 1400, according to an implementation herein, that includes a scalable centralized error queue. In one implementation, the graphics process 1400 may include an SGI 1410, a graphics engine 1450, a security engine 1460, and a display engine 1470. Note, however, that the basic principles of the present disclosure are not limited to this implementation. For example, while the illustrated embodiment uses a system graphics interface 1410, the techniques described herein are equally applicable to non-graphics interfaces.

[0199] The system graphics interface 1410 compiles error data received from each source within an error aggregator 1430. Sources of errors may include units of the graphics processor 1400, such as the graphics engine 1450, the security engine 1460, the display engine 1470, and other non-graphics component(s) 1480. Each of the error sources may include error routing circuitry 1455, 1465, 1475, 1485 responsible for detecting errors and reporting the errors to the SGI 1410.

[0200] Errors may be classified into different types of errors. In an implementation herein, the types of errors may include correctable errors, uncorrectable fatal errors, or uncorrectable non-fatal errors. Correctable errors are those errors that can be corrected by hardware and do not require intervention provided by software for continued operation. For example, a storage device with an ECC SECDED (single error correction and double error detection) scheme is capable of automatically correcting single-bit errors. Errors on a storage device with parity protection and a replay / rewind mechanism are also considered correctable errors.

[0201] Uncorrectable errors are errors that cannot be corrected by hardware. These errors should be reported to software for error recovery. Structural errors (such as double-bit errors on an ECC-protected storage device), single-bit errors on a parity-protected storage device, and functional protocol errors (such as a failure to complete a write operation) are examples of uncorrectable errors. Recovery from such errors depends on the impact of the error. Thus, uncorrectable errors are further classified into contained uncorrectable errors and non-contained uncorrectable errors. Contained uncorrectable errors are those errors that can be recovered without affecting other contexts that may be executing on the processor. In some implementations, errors may be classified as informational errors, which are those errors that do not utilize any action or recovery.

[0202] Contained uncorrectable errors (also known as local errors, local uncorrectable errors, or non-fatal errors) are errors that are detected internally by the agent, or whose source can be traced back to a transaction specific to the agent that only affects the context running on that agent. As such, contained uncorrectable errors can be recovered without affecting other contexts. Such errors are called contained errors (or local errors) because they are local to the context / engine. An example of a contained error would be an ECC error in the GRF.

[0203] Non-contained uncorrectable errors (also known as global errors, global uncorrectable errors, or fatal errors) are errors that cannot be traced back to only affect a single context. As such, these errors affect all contexts on the GPU and have a larger radius of impact. For recovery, these types of errors minimally utilize GPU resets. Once a non-contained error is signaled to the system, the system processes the appropriate recovery flow.

[0204] In one implementation, the error reporting circuitry 1455, 1465, 1475, 1485 of each distributed component of the graphics processor 1400 can report errors using a standard error log format for a specific error type (e.g., uncorrectable error log format or correctable error log format), as further discussed below with respect to Figure 15 The error routing circuitry 1455, 1465, 1475, 1485 can then transmit the compiled error data to the SGI 1410 via an on-chip fabric (e.g., such as an IO system fabric (IOSF)).

[0205] The error classifier 1440 of the SGI 1410 can identify the error type from the error report message and route the error report message to the corresponding centralized error log queue maintained by the error aggregator 1430. The centralized error log queues maintained by the error aggregator 1430 can include an uncorrectable centralized error queue 1432, an informational (info) centralized error queue 1434, and a correctable centralized error queue 1436.

[0206] The centralized error log queues 1432, 1434, 1436 are fine-grained error logs that can be used to determine the cause of the error, for debugging, for telemetry, and for silicon health assessment in production. In some embodiments, the centralized error log queues 1432, 1434, 1436 are configured to be sticky across thermal resets (i.e., retain their values across resets). Since a global error causes a thermal reset of the graphics processor 1400, the error information corresponding to a global uncorrectable error would be lost if it were not sticky.

[0207] In one implementation, the uncorrectable global centralized error queue 1432 can be implemented as a queue with a depth UNCORR_ERR_Q_SIZE parameter for uncorrectable errors, where each entry is UNCORR_ERR_Q_WIDTH double-word (double-word, 32 bits) wide. The UNCORR_ERR_Q_SIZE parameter should be exposed to the software. In one example, this parameter can take the value 4. In another example, the UNCORR_ERR_Q_WIDTH parameter can be 16. Additionally, uncorr_sticky_log_write_offset (uncorrectable_sticky_log_write_offset) should be defined for the unit to write to this queue 1432. uncorr_sticky_log_read_offset (uncorrectable_sticky_log_read_offset) should also be defined for the software to read from the queue 1432, as further discussed with respect to Figure 18 Further discussion.

[0208] In one implementation, the informational centralized error queue 1434 can be implemented as a queue with a depth UNCORR_ERR_INFO_Q_SIZE (uncorrectable_error_informational_Q_size) parameter for informational-only errors, where each entry is UNCORR_ERR_INFO_Q_WIDTH (uncorrectable_error_informational_Q_width) double-word wide. This queue 1434 is used to record error conditions that must interrupt the software. The UNCORR_ERR_INFO_Q_SIZE parameter should be exposed to the software. In one example, the value of this parameter can be 4. In another example, the width of the UNCORR_ERR_INFO_Q_WIDTH parameter can be 16 double-words. info_sticky_log_write_offset (informational_sticky_log_write_offset) should be defined for the unit to write to this queue 1434. info_sticky_log_read_offset (informational_sticky_log_read_offset) should be defined for the SW to read from the queue 1434, as further discussed with respect to Figure 18 Further discussion.

[0209] In one implementation, the correctable centralized error queue 1436 can be implemented as a queue with a depth CORR_ERR_Q_SIZE (Correctable Error Q Size) parameter for correctable errors, where the width of each entry is CORR_ERR_Q_WIDTH (Correctable Error Q Width) doublewords. In one example, CORR_ERR_Q_SIZE can be 16. In another example, the width of CORR_ERR_Q_WIDTH can be 4 doublewords. corr_sticky_log_write_offset (Correctable Sticky Log Write Offset) should be defined for the unit to write to this queue 1436. corr_sticky_log_read_offset (Correctable Sticky Log Read Offset) should be defined for the software to read from the queue 1436, as further discussed with respect to Figure 18 discussed further below.

[0210] Additionally, in the illustrated embodiment, the system graphics interface 1410 uses the error interrupt circuitry 1420 for managing interrupts related to errors received at the SGI 1410. In this example, a message signaled interrupt (MSI) can be generated to convey the error. The MSI can call the graphics driver, which can then utilize the error control and status register 1438 to access the error data that should be parsed and read. The driver can then read the appropriate log register.

[0211] As previously mentioned, in the implementations herein, the error log format to be used by the reporting unit depends on whether the detected error is a non-correctable error or a correctable error. Figure 15 Depicted is an error log format for reporting errors to a scalable centralized error queue according to implementations herein. Figure 15 The error log format shown is only one example of an error log format, and other formats can be utilized by the implementations herein. As shown, Figure 15 depicted are a non-correctable error log format 1500 (covering both non-correctable non-inclusive errors and non-correctable inclusive errors) and a correctable error log format 1510.

[0212] In one embodiment, if an uncorrectable error (including an (informational) error or a non-inclusive error) is detected, the distributed unit reporting the error (also referred to herein as a "leaf unit") shall write the granular error information into the appropriate centralized error queue in the SGI 1410. In one example, this is achieved by performing a 16-doubleword write to uncorr / info_sticky_log_write_offset in the SGI 1410. Example bit definitions are shown in the uncorrectable error log format 1500 of Figure 15 For any doublewords that are not applicable, the reporting unit may write 0 for these doublewords and set the appropriate valid bits in the error log header to 0. In some embodiments, a quadword write may be implemented, where the high 32 bits carry the ID of the unit performing the write and the low 32 bits carry the actual data written by that unit. This allows multiple leaf units to simultaneously write multi-cycle messages to different entries in the same error log queue. The ID of the leaf unit in the 64b message allows the queue to de-interleave data from each source.

[0214] Using the ID of the unit as an identifier, the SGI 1410 can store the incoming doublewords into the appropriate entries in the appropriate centralized error queues 1432, 1434, 1436. The SGI 1410 can continue to fill the entries until, for example, 16 doublewords are received from the unit. If further doublewords are received from the same unit, they are treated as a new error log and new entries in the centralized error queues 1432, 1434, 1436 are used.

[0215]

[0216] Figure 16 In one example, the SGI 1410 can store four such detailed sets of error logs in each of the centralized error queues 1432, 1434, 1436, which retain their values across a thermal reset. The first four such errors are logged. In one implementation, any errors received after that are not logged by the SGI 1410 until the queues 1432, 1434, 1436 are cleared by software. Figure 16

[0217] Depicted is the error log header 1600 format for an error log reporting message for a scalable centralized error queue according to an implementation herein. Figure 16The error log header format shown is merely an example of an error log header format, and other formats may be utilized by the implementations herein. As depicted in the uncorrectable error log format 1500 and the correctable error log format 1510, the error log header may be the zeroth (0) double word of the log. For example, the error log header provides information to software to identify the source unit (e.g., leaf source unit) where the error was detected, the error severity, the error log format used, and the valid bits for the granularity error log.

[0217] As shown in the example error log header 1600, the low bits may contain the error source ID (ErrSrcID), which is encoded for software to identify the unit type and instance number of the leaf source of the error. In some implementations, a set of ErrSrcIDs may be defined for the graphics processor 1400 for this purpose. Each unit should use its appropriate ErrSrcID when writing to the error log. In one implementation, it may be assumed that each unit participating in the error reporting flow may have a register to identify its instance number in the overall system.

[0218] InternalErrLog has bits for identifying the internal source of the error (both correctable and uncorrectable errors). The structural error type is recorded per FabricErrLog (structural error log) definition.

[0219] Figure 17 Depicted is a grouping including the internal error log 1700 and the structural error log 1710 for an error log reporting message for a scalable centralized error queue according to implementations herein. Figure 17 The log format shown is merely an example of a log format, and other formats may be utilized by the implementations herein.

[0220] In one implementation, if an internal error is detected, the ErrLogHeader.IntErrLog_Valid (error log header. internal error log_valid) bit in the error log header (e.g., Figure 16 the error log header 1600) should be set to 1. Further details are then recorded using the InternalErrLog (internal error log) double word shown in the internal error log 1700. The internal error log 1700 is used to record the address of the transmitter where an internal error (such as uncorr ECC) was detected. This may be a pointer to the array where the error occurred. This allows software to accumulate telemetry information about such errors over time and make appropriate repair decisions.

[0221] In one implementation, if a structural error is detected, the error log header (e.g.,Figure 16 The ErrLogHeader.FabricErrLog_Valid bit in the error log header 1600 is set to 1. Further details are recorded using the FabricErrLog doubleword shown in the structured error log 1710. In the case of a structured error, the header of the failed packet is recorded in the structured error log 1710 for the original requester and the final responder units. In one implementation, if (one or more) Packet_log entries are recorded, the ErrLogHeadr.PktLog_Valid bit in the error log header 1600 should be set to 1. In this case, the first uncorrectable error pointer points to the bit position in the FabricErrLog doubleword to indicate the first uncorrectable error detected. For fault_code errors, this pointer can point to bit 9.

[0222] Regarding correctable errors, these are errors that the hardware can recover without the need for software intervention. However, the software should still be informed about the number of such correctable errors and the types of errors seen, and the software should also clear the logs for these errors. Similar to the uncorrectable error log format 1500, a correctable error log format 1510 is also defined to store correctable error logs.

[0223] Although the software should be informed about correctable errors, interrupting the software at every correctable error may affect performance. Thus, a hardware counter with a programmable threshold and the ability to mask interrupts can be implemented by the SGI 1410. In some embodiments, the correctable error count can be sticky across a thermal reset. Consider a situation where a certain number of correctable errors are encountered and the count is incremented, but the count is less than the programmed threshold. If a subsequent uncorrectable control error occurs, the unit should undergo a reset and the correctable error count information is lost since it has not been escalated to the software. Thus, making the count sticky resolves such situations.

[0224] To obtain the errors stored in the centralized error log queues 1432, 1434, 1436, the software should read out these entries from the SGI 1410 by successive reads of the same offset (e.g., uncorr_sticky_log_read_offset / info_sticky_log_read_offset / corr_sticky_log_read_offset).Figure 18 Depicts status and control registers used by software as part of reading entries from a scalable centralized error queue, according to implementations herein. Figure 18 Depicts centralized log read control register 1800, centralized log status register 1810, and centralized log control register 1820. Figure 18 The (one or more) status and control register formats shown are provided only as examples, and other formats may be utilized by implementations herein.

[0225] The centralized log status register 1810 can be utilized by software to determine the number of the entry to be read out. Then, software can program the centralized log read control register 1800 with the number of the entry it wants to read. Software can read any number of doublewords of that entry by successive reads of uncorr_sticky_log_read_offset / info_sticky_log_read_offset / corr_sticky_log_read_offset. This enables software to read all the information of an entry (e.g., during a debug scenario), or enables software to read a small subset of the information of an entry for telemetry purposes (e.g., during a production scenario). Software can parse error information based on the format and unit information provided in the zeroth doubleword of each entry.

[0226] For containment errors, software should read and clear the local error interrupt upon receiving the local error interrupt or upon engine reset during engine restart. For fatal errors, software should read and clear the fatal error after reset as part of the normal error collection routine.

[0227] The centralized log control register 1820 provides control and status registers for software to determine the status of the centralized log and clear them.

[0228] Figure 19 Is a flowchart illustrating an embodiment of a method 1900 for reporting errors to a scalable centralized error queue in a graphics architecture. Method 1900 may be executed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof. For the sake of simplicity and clarity of presentation, the processes of method 1900 are illustrated in a linear order; however, any number of them are contemplated to be executed in parallel, asynchronously, or in a different order. Further, for simplicity, clarity, and ease of understanding, reference is made to Figures 1 - 18Many of the components and processes described may not be repeated or discussed below. In one implementation, a processing unit (such as, Figure 14 the graphics processor 1400) may execute method 1900.

[0229] Method 1900 begins at processing block 1910, where the processor may detect an error in a component of the distributed error reporting hierarchy of the graphics architecture. Then, at block 1920, the processor may classify the error as a certain error type, where the error type includes one or more of the following: uncorrectable (non-inclusive) error, uncorrectable inclusive error, informational error, or correctable error.

[0230] Subsequently, at block 1930, the processor may generate an error report message for the detected error. In one implementation, the format of the error report message is consistent with the error type of the error. Finally, at block 1940, the processor may send the error report message to an error aggregator of the system graphics interface of the graphics architecture. In one implementation, the error aggregator stores information about the error from the error report message in a centralized error queue corresponding to the error type of the error.

[0231] Figure 20 is a flowchart illustrating an embodiment of method 2000 for recording errors in a scalable centralized error queue in a graphics architecture. Method 2000 may be executed by processing logic that may include hardware (e.g., circuitry, special purpose logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof. For purposes of presentation simplicity and clarity, the processes of method 2000 are illustrated in a linear order; however, any number of them are contemplated to be executed in parallel, asynchronously, or in a different order. Further, for simplicity, clarity, and ease of understanding, reference is made to Figures 1 - 19 Many of the components and processes described may not be repeated or discussed below. In one implementation, a processing unit (such as, Figure 14 the graphics processor 1400) may execute method 2000.

[0232] Method 2000 begins at processing block 2010, where the processor may receive an error report message from a component of the distributed error reporting hierarchy of the graphics architecture. Subsequently, at block 2020, the processor may identify the error type of the error from the error report message. In one implementation, the error type includes one or more of the following: uncorrectable non-inclusive error, uncorrectable inclusive error, informational error, or correctable error.

[0233] Subsequently, at block 2030, the processor may record the error in an entry of a centralized error queue corresponding to the error type. In one implementation, the entry is populated with the information included in the error report message corresponding to the error. Finally, at block 2040, the processor may update the control and status registers corresponding to the centralized error queue to reflect the status of the centralized error queue.

[0234] Figure 21 is a flowchart illustrating an embodiment of a method 2100 for recovering errors in a scalable centralized error queue in a graphics architecture. Method 2100 may be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof. For purposes of presentation simplicity and clarity, the processes of method 2100 are illustrated in a linear order; however, any number of them may be performed in parallel, asynchronously, or in a different order. Further, for simplicity, clarity, and ease of understanding, many of the components and processes described with reference to Figures 1 - 20 may not be repeated or discussed hereinafter. In one implementation, a processing unit (such as, Figure 14 the graphics processor 1400) may perform method 2100.

[0235] Method 2100 begins at processing block 2110, where the processor may enter an error recovery process to identify errors aggregated by an error aggregator component reported to the system graphics interface of the graphics architecture. Then, at block 2120, the processor may determine the number of entries in each of the plurality of centralized error queues maintained by the error aggregator. In one implementation, the plurality of centralized error queues includes an uncorrectable non-inclusive error queue, an uncorrectable inclusive error queue, and a correctable error queue.

[0236] Subsequently, at block 2130, the processor may program a control register with the entry number of one of the plurality of centralized error queues to be read. At block 2140, the processor may perform sequential reads of an entry having the entry number in one of the plurality of centralized error queues. In one implementation, the error information of the entry is parsed based on the format and unit information provided in the zeroth doubleword of the entry. Finally, at block 2150, after reading the entry, the processor may program the log control register bits to cause the entry in one of the plurality of centralized error queues to be cleared.

[0237] The following examples relate to further embodiments. Example 1 is an apparatus for facilitating a scalable centralized error queue in a processing architecture. The apparatus of Example 1 includes a processor that includes a system interface hosting an error aggregator, wherein the processor is configured to: host at least one centralized error queue in the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processor; receive an error report message from a component among the components of the processor, the error report message corresponding to an error detected by the component; and record the error as an entry in the at least one centralized error queue based on the error type of the error.

[0238] In Example 2, the subject matter of Example 1 may optionally include, wherein the processor is further configured to identify the error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable inclusion error, or an uncorrectable non-inclusion error; and wherein the at least one centralized error queue includes one or more of the following to store errors of the corresponding error types in the error types: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue. In Example 3, the subject matter of any one of Examples 1-2 may optionally include, wherein in response to the error type including one of an uncorrectable inclusion error or an uncorrectable non-inclusion error, the error report message is in an uncorrectable error log format.

[0239] In Example 4, the subject matter of any one of Examples 1-3 may optionally include, wherein in response to the error type including a correctable error, the error report message is in a correctable error log format. In Example 5, the subject matter of any one of Examples 1-4 may optionally include, wherein the processor is further configured to update at least one of the following to reflect the updated state of the at least one centralized error queue in which the error is recorded: a centralized log read control register, a centralized log status register, or a centralized log control register.

[0240] In Example 6, the subject matter of any one of Examples 1-5 may optionally include, wherein the driver is configured to read entries of the at least one centralized error queue and clear entries of the at least one centralized error queue using the centralized log read control register, the centralized log status register, and the centralized log control register. In Example 7, the subject matter of any one of Examples 1-6 may optionally include, wherein the at least one centralized error queue maintains entries across a thermal reset of the processor.

[0241] In Example 8, the subject matter of any one of Examples 1-7 may optionally include, wherein at least one centralized error queue is written to utilize a quad-word write, the quad-word write having a high 32 bits and a low 32 bits, the high 32 bits carrying an identifier (ID) of a unit that performs the quad-word write, and the low 32 bits carrying data of the quad-word write. In Example 9, the subject matter of any one of Examples 1-8 may optionally include, wherein the processor includes a graphics processing unit (GPU).

[0242] Example 10 is a method for facilitating a scalable centralized error queue in a processing architecture. The method of Example 10 may include: hosting, by a processing device, at least one centralized error queue in an error aggregator, the processing device including a system interface that hosts the error aggregator, the at least one centralized error queue being for storing error logs for errors detected by components of the processing device; receiving, by the processing device, an error report message from a component among the components of the processing device, the error report message corresponding to an error detected by the component; and recording, by the processing device, the error as an entry in the at least one centralized error queue based on an error type of the error.

[0243] In Example 11, the subject matter of Example 10 may optionally further include identifying, from the error report message, an error type of the error, the error type of the error including one of the following: a correctable error, an uncorrectable inclusion error, or an uncorrectable non-inclusion error; and wherein the at least one centralized error queue includes one or more of the following to store errors of the corresponding error type in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue. In Example 12, the subject matter of Examples 10-11 may optionally include, wherein, in response to the error type including one of an uncorrectable inclusion error or an uncorrectable non-inclusion error, the error report message is in an uncorrectable error log format.

[0244] In Example 13, the subject matter of Examples 10-12 may optionally include, wherein, in response to the error type including a correctable error, the error report message is in a correctable error log format. In Example 14, the subject matter of Examples 10-13 may optionally further include: updating at least one of the following to reflect an updated state of the at least one centralized error queue in which the error is recorded: a centralized log read control register, a centralized log status register, or a centralized log control register.

[0245] In Example 15, the subject matter of Examples 10-14 may optionally include, wherein the driver is used to read entries of at least one centralized error queue and clear entries of at least one centralized error queue by using a centralized log read control register, a centralized log status register, and a centralized log control register.

[0246] Example 16 is a non-transitory computer-readable storage medium for facilitating a scalable centralized error queue in a processing architecture. The non-transitory computer-readable storage medium of Example 16 has instructions stored thereon that, when executed by one or more processors, cause the processors to: host at least one centralized error queue in an error aggregator by a processing device of the one or more processors, the processing device including a system interface hosting the error aggregator, the at least one centralized error queue being used to store error logs for errors detected by components of the processing device; receive an error report message from a component among the components of the processing device, the error report message corresponding to an error detected by the component; and record the error as an entry in the at least one centralized error queue based on an error type of the error.

[0247] In Example 17, the subject matter of Example 16 may optionally include, wherein the one or more processors are further used to: identify an error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable inclusion error, or an uncorrectable non-inclusion error, and wherein the at least one centralized error queue includes one or more of the following to store errors of the corresponding error types in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue. In Example 18, the subject matter of Examples 16-17 may optionally include, wherein in response to the error type including one of an uncorrectable inclusion error or an uncorrectable non-inclusion error, the error report message is in an uncorrectable error log format; and wherein in response to the error type including a correctable error, the error report message is in a correctable error log format.

[0248] In Example 19, the subject matter of Examples 16-18 may optionally include, wherein the one or more processors are further used to update at least one of the following to reflect an updated state of the at least one centralized error queue in which the error is recorded: a centralized log read control register, a centralized log status register, or a centralized log control register. In Example 20, the subject matter of Examples 16-19 may optionally include, wherein the driver is used to read entries of at least one centralized error queue and clear entries of at least one centralized error queue by using a centralized log read control register, a centralized log status register, and a centralized log control register.

[0249] Example 21 is a system for facilitating a scalable centralized error queue in a processing architecture. The system of Example 21 can optionally include: a memory; and a processor communicatively coupled to the memory, and the processor includes a system interface hosting an error aggregator, wherein the processor is configured to: host at least one centralized error queue in the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processor; receive an error report message from a component among the components of the processor, the error report message corresponding to an error detected by the component; and record the error as an entry in the at least one centralized error queue based on an error type of the error.

[0250] In Example 22, the subject matter of Example 21 can optionally include, wherein the processor is further configured to identify an error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable inclusion error, or an uncorrectable non-inclusion error; and wherein the at least one centralized error queue includes one or more of the following to store errors of the corresponding error type in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue. In Example 23, the subject matter of any of Examples 21-22 can optionally include, wherein in response to the error type including one of an uncorrectable inclusion error or an uncorrectable non-inclusion error, the error report message is in an uncorrectable error log format.

[0251] In Example 24, the subject matter of any of Examples 21-23 can optionally include, wherein in response to the error type including a correctable error, the error report message is in a correctable error log format. In Example 25, the subject matter of any of Examples 21-24 can optionally include, wherein the processor is further configured to update at least one of the following to reflect an updated state of the at least one centralized error queue in which the error is recorded: a centralized log read control register, a centralized log status register, or a centralized log control register.

[0252] In Example 26, the subject matter of any of Examples 21-25 can optionally include, wherein the driver is configured to read entries of the at least one centralized error queue and clear entries of the at least one centralized error queue using the centralized log read control register, the centralized log status register, and the centralized log control register. In Example 27, the subject matter of any of Examples 21-26 can optionally include, wherein the at least one centralized error queue maintains entries across a thermal reset of the processor.

[0253] In Example 28, the subject matter of any of Examples 21-27 may optionally include where at least one centralized error queue is written to utilize a quad-word write that has a high 32 bits and a low 32 bits, the high 32 bits carrying an identifier (ID) of the unit that performs the quad-word write, and the low 32 bits carrying the data of the quad-word write. In Example 29, the subject matter of any of Examples 21-28 may optionally include where the processor includes a graphics processing unit (GPU).

[0254] Example 30 is a device for facilitating a scalable centralized error queue in a processing architecture, the device including: means for hosting at least one centralized error queue in an error aggregator using a processing device, the processing device including a system interface that hosts the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processing device; means for receiving an error report message from a component among the components of the processing device, the error report message corresponding to an error detected by the component; and means for recording the error as an entry in the at least one centralized error queue based on the error type of the error. In Example 31, the subject matter of Example 30 may optionally include a device further configured to perform the method of any of Examples 11 to 15.

[0255] Example 32 is at least one machine-readable medium including a plurality of instructions that, in response to being executed on a computing device, cause the computing device to perform the method of any of Examples 10-15. Example 33 is a means for facilitating a scalable centralized error queue in a processing architecture, the means being configured to perform the method of any of Examples 10 to 15. Example 34 is a device for facilitating a scalable centralized error queue in a processing architecture, the device including means for performing the method of any of Examples 10 to 15. Details in the examples may be used anywhere in one or more embodiments.

[0256] The foregoing specification and drawings are to be regarded in an illustrative rather than a restrictive sense. Those skilled in the art will understand that various modifications and changes can be made to the embodiments described herein without departing from the broader spirit and scope of the features as set forth in the appended claims.

Claims

1. A device, comprising: A processor comprising a system interface hosting an error aggregator, wherein the processor is configured to: hosting at least one centralized error queue in the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processor; receiving an error report message from a component of the components of the processor, the error report message corresponding to an error detected by the component; and Based on an error type of the error, the error is recorded as an entry in the at least one centralized error queue.

2. The device according to claim 1, wherein: The processor is further used to identify an error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable contained error, or an uncorrectable non-contained error; and wherein the at least one centralized error queue includes one or more of the following to store errors of corresponding error types in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue.

3. The device according to any one of claims 1 to 2, wherein: In response to the error type comprising one of the uncorrectable contained error or the uncorrectable non-contained error, the error report message is in an uncorrectable error log format.

4. The device according to any one of claims 1 to 3, wherein: In response to the error type including the correctable error, the error report message is in a correctable error log format.

5. The device according to any one of claims 1 to 4, wherein: The processor is further configured to update at least one of the following to reflect an updated status of the at least one centralized error queue that logged the error: a centralized log read control register, a centralized log status register, or a centralized log control register.

6. The device according to any one of claims 1 to 5, wherein: The driver is used to read the entries of the at least one centralized error queue and clear the entries of the at least one centralized error queue by using the centralized log reading control register, the centralized log status register and the centralized log control register.

7. The device according to any one of claims 1 to 6, wherein: The at least one centralized error queue maintains entries across warm resets of the processor.

8. The device according to any one of claims 1 to 7, wherein: The at least one centralized error queue is written to utilize quad word writes having upper 32 bits carrying an identifier ID of a unit performing the quad word write and lower 32 bits carrying data of the quad word write.

9. The device according to any one of claims 1 to 8, wherein: The processor includes a graphics processing unit GPU.

10. A method comprising: hosting, by a processing device, at least one centralized error queue in an error aggregator, the processing device including a system interface hosting the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processing device; receiving, by the processing device, an error report message from a component of the components of the processing device, the error report message corresponding to an error detected by the component; as well as The error is recorded as an entry in the at least one centralized error queue by the processing device based on an error type of the error.

11. The method of claim 10, further comprising identifying an error type of the error from the error report message, the error type of the error comprising one of: a correctable error, an uncorrectable contained error, or an uncorrectable non-contained error; and wherein, The at least one centralized error queue includes one or more of the following items to store errors of corresponding error types: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue.

12. The method according to any one of claims 10 to 11, wherein: In response to the error type comprising one of the uncorrectable contained error or the uncorrectable non-contained error, the error report message is in an uncorrectable error log format.

13. The method according to any one of claims 10 to 12, wherein: In response to the error type including the correctable error, the error report message is in a correctable error log format.

14. The method of any one of claims 10 to 13, further comprising: At least one of the following is updated to reflect the updated status of the at least one centralized error queue that logged the error: a centralized log read control register, a centralized log status register, or a centralized log control register.

15. The method according to any one of claims 10 to 14, wherein: The driver is used to read the entries of the at least one centralized error queue and clear the entries of the at least one centralized error queue by using the centralized log reading control register, the centralized log status register and the centralized log control register.

16. A non-transitory computer readable medium having instructions stored thereon, which instructions, when executed by one or more processors, cause the one or more processors to: hosting at least one centralized error queue in an error aggregator by a processing device of the one or more processors, the processing device comprising a system interface hosting the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processing device; receiving, by the processing device, an error report message from a component of the components of the processing device, the error report message corresponding to an error detected by the component; as well as The error is recorded as an entry in the at least one centralized error queue by the processing device based on an error type of the error.

17. The non-transitory computer readable medium of claim 16, wherein: The one or more processors are further used to identify an error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable contained error, or an uncorrectable non-contained error; and wherein the at least one centralized error queue includes one or more of the following to store errors of corresponding error types in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue.

18. The non-transitory computer readable medium of any one of claims 16-17, wherein: In response to the error type comprising one of the uncorrectable contained error or the uncorrectable non-contained error, the error report message is in an uncorrectable error log format; and wherein, in response to the error type comprising the correctable error, the error report message is in a correctable error log format.

19. The non-transitory computer readable medium of any one of claims 16-18, wherein: The one or more processors are further configured to update at least one of the following to reflect an updated status of the at least one centralized error queue that logged the error: a centralized log read control register, a centralized log status register, or a centralized log control register.

20. The non-transitory computer readable medium of any one of claims 16-19, wherein: The driver is used to read the entries of the at least one centralized error queue and clear the entries of the at least one centralized error queue by using the centralized log reading control register, the centralized log status register and the centralized log control register.

21. A system for facilitating a scalable centralized error queue in a processing architecture, the system comprising: Memory; as well as a processor communicatively coupled to the memory and comprising a system interface hosting an error aggregator, wherein the processor is configured to: hosting at least one centralized error queue in the error aggregator, the at least one centralized error queue for storing error logs for errors detected by components of the processor; receiving an error report message from a component of the components of the processor, the error report message corresponding to an error detected by the component; and Based on an error type of the error, the error is recorded as an entry in the at least one centralized error queue.

22. The system of claim 21, wherein: The processor is further used to identify an error type of the error from the error report message, the error type of the error including one of the following: a correctable error, an uncorrectable contained error, or an uncorrectable non-contained error; and wherein the at least one centralized error queue includes one or more of the following to store errors of corresponding error types in the error type: a correctable centralized error queue, an uncorrectable local centralized error queue, or an uncorrectable global centralized error queue.

23. The system of any one of claims 21-22, wherein: In response to the error type comprising one of the uncorrectable contained error or the uncorrectable non-contained error, the error report message is in an uncorrectable error log format.

24. The system of any one of claims 21-23, wherein: In response to the error type including the correctable error, the error report message is in a correctable error log format.

25. An apparatus for facilitating a scalable centralized error queue in a processing architecture, the apparatus comprising means for performing the method of any of claims 10-15.

Citation Information

Cited By

  • Message tracking circuit and system

    CN121501585A