Correctable error address filtering in processing architecture

By introducing correctable error address filtering technology into the graphics processor, error detection, recording and correction problems in the prior art are solved, and the reliability and usability of the system are improved.

CN120196464APending Publication Date: 2025-06-24INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411576045.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When existing graphics processors handle errors, it is difficult to effectively detect, record and correct errors, affecting the reliability, availability and serviceability of the system.

Method used

A correctable error address filtering technology is designed. By introducing a dedicated processor unit into the graphics processor, error addresses can be detected and recorded and corrected when necessary, ensuring that the system can recover quickly when an error occurs.

Benefits of technology

Improves the reliability and availability of graphics processors, reduces system crashes and data corruption caused by errors, and enhances system serviceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196464A_ABST
    Figure CN120196464A_ABST
Patent Text Reader

Abstract

An apparatus for facilitating correctable erroneous address filtering in a processing architecture is disclosed. The apparatus includes a processor including error routing hardware circuitry to: receive data associated with an error detected in a source hardware component, the source hardware component hosting the error routing circuitry; determining, based on the data, that the error is classified as a correctable error; comparing the address of the correctable error with an entry maintained by correctable error address filtering circuitry of the error routing circuitry; in response to the address of the correctable error matching the entry of the correctable error address filter circuitry, masking the correctable error to prevent reporting of the correctable error to error aggregation hardware circuitry of the processor; and reporting the correctable error to the error aggregation hardware circuitry in response to the address of the correctable error not matching the entry of the correctable error address filter circuitry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document generally relates to data processing, and more specifically, to correctable error address filtering in a processing architecture. Background Art

[0002] Current parallel graphics data processing includes systems and methods developed to perform specific operations on graphics data, such as, for example, linear interpolation, tessellation, rasterization, texture mapping, depth testing, etc. Traditionally, graphics processors have used fixed-function computing units to process graphics data; however, recently, multiple parts of the graphics processor have been made programmable, enabling such processors to support a wider variety of operations for processing vertex data and fragment data.

[0003] To further improve performance, graphics processors typically implement processing techniques such as pipelining, which attempt to process as much graphics data as possible in parallel across different parts of the graphics pipeline. Parallel graphics processors with single instruction, multiple data (SIMD) or single instruction, multiple thread (SIMT) architectures are designed to maximize the amount of parallel processing in the graphics pipeline. In the SIMD architecture, a computer with multiple processing elements attempts to perform the same operation on multiple data points simultaneously. In the SIMT architecture, groups of parallel threads attempt to execute program instructions synchronously together as frequently as possible to improve processing efficiency.

[0004] Graphics processors are often used in applications in the fields of artificial intelligence (AI) and machine learning (ML). General-purpose graphics processing units (GPGPUs) with matrix acceleration circuitry are often deployed and hosted in data centers. Hardware resilience is a requirement for graphics architectures in the data center market segment. Architectures with Reliability, Availability, and Serviceability (RAS) features are designed to meet resilience goals. Reliability refers to how reliable the operation of the design is. Availability refers to the uptime of the operation that the design can still provide in the presence of errors. Serviceability refers to how easily the design can be serviced to restore it to a reliable operation once an error occurs in the design.

[0005] To do this, the design should be able to detect and record errors (to improve reliability), correct these errors if possible (to improve usability), and report to a higher-level system component (such as a driver) when the error is uncorrectable. The system software can then take appropriate actions to service the error and bring the design back into reliable operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Accordingly, to understand in detail the manner in which the features of the current embodiments described above can be obtained, a more particular description of the embodiments briefly summarized above may be referred to, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings only illustrate typical embodiments and should not be considered as limiting the scope of the embodiments.

[0008] Figure 1 is a block diagram of a processing system.

[0009] Figures 2A - 2D illustrates a computing system and a graphics processor.

[0010] Figures 3A - 3C illustrates a block diagram of an additional graphics processor and a computing accelerator architecture.

[0011] Figure 4 is a block diagram of a graphics processing engine of a graphics processor.

[0012] Figures 5A - 5B illustrates thread execution logic including an array of processing elements employed in a graphics processor core.

[0013] Figure 6 illustrates additional execution units.

[0014] Figure 7 is a block diagram illustrating a graphics processor instruction format.

[0015] Figure 8 is a block diagram of an additional graphics processor architecture.

[0016] Figures 9A - 9B illustrates a graphics processor command format and a command sequence.

[0017] Figure 10 illustrates an example graphics software architecture for a data processing system.

[0018] Figure 11A is a block diagram illustrating an IP core development system.

[0019] Figure 11B illustrates a cross-sectional side view of an integrated circuit package component.

[0020] Figure 11CThe figure shows a packaged component that includes a hardware logic die connected to multiple cells of a substrate (e.g., a base die).

[0021] Figure 11D The figure shows a packaged component that includes interchangeable dies.

[0022] Figure 12 is a block diagram of an example system-on-chip integrated circuit shown in the figure.

[0023] Figures 13A - 13B is a block diagram of an example graphics processing unit for use within a SoC shown in the figure.

[0024] Figure 14 The figure shows an embodiment that includes a processor for providing correctable error address filtering according to the implementations herein.

[0025] Figure 15 is a block diagram of a detailed view of a processor that implements correctable error address filtering according to the implementations herein.

[0026] Figure 16 is a flowchart of an embodiment of a method for processing correctable error address filtering in an architecture.

[0027] Figure 17 is a flowchart of an embodiment of a method for clearing entries in correctable error address filtering in a processing architecture. DETAILED DESCRIPTION

[0028] A graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate, for example, graphics operations, machine learning operations, pattern analysis operations, and / or various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). Alternatively, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of the manner in which the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0029] In the following description, numerous specific details are set forth to provide a more thorough understanding. However, it will be apparent to one of ordinary skill in the art that embodiments described herein may be practiced without one or more of these specific details. In other instances, well-known features have not been described so as not to obscure the details of the current embodiments. System Overview

[0030] Figure 1 FIG. 6 is a block diagram of a processing system 100 according to an embodiment. The system 100 can be used in: a single-processor desktop computer system, a multi-processor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in a mobile device, a handheld device, or an embedded device, such as for use within an Internet-of-things (IoT) device having wired or wireless connectivity to a local area network or a wide area network.

[0031] In one embodiment, the system 100 can include, be coupled to, or be integrated within: a server-based gaming platform; a game console, including a game and media console; a mobile game console, a handheld game console, or an online game console. In some embodiments, the system 100 is part of a mobile phone, a smartphone, a tablet computing device, or a mobile Internet-connected device (such as a laptop computer having a low internal storage capacity). The processing system 100 can also include, be coupled to, or be integrated within: a wearable device, such as a smartwatch wearable device; smart glasses or clothing that are enhanced with augmented reality (AR) or virtual reality (VR) features to provide visual, audio, or tactile output to supplement a real-world visual, audio, or tactile experience or otherwise provide text, audio, graphics, video, holographic images or video, or tactile feedback; other augmented reality (AR) devices; or other virtual reality (VR) devices. In some embodiments, the processing system 100 includes a television or a set-top box device, or is part of a television or a set-top box device. In one embodiment, the system 100 can include, be coupled to, or be integrated within a self-driving vehicle, such as a bus, a tractor-trailer, a car, a motorcycle or a power cycle, an airplane or a glider (or any combination thereof). The self-driving vehicle can use the system 100 to process the environment sensed around the vehicle.

[0032] In some embodiments, each of one or more processors 102 includes one or more processor cores 107 for processing instructions that, when executed, perform operations for system and user software. In the embodiments herein, a processor may refer to dedicated hardware circuitry for efficiently processing commands / instructions and may be referred to as processor circuitry. In some embodiments, at least one of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computing via Very Long Instruction Word (VLIW). The one or more processor cores 107 may process different instruction sets 109, and the different instruction sets 109 may include instructions for facilitating emulation of other instruction sets. The processor core 107 may also include other processing devices, such as a Digital Signal Processor (DSP).

[0033] In some embodiments, the processor 102 includes a cache memory 104. Depending on the architecture, the processor 102 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among the various components of the processor 102. In some embodiments, the processor 102 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), and the external cache may be shared among the processor cores 107 using known cache coherence techniques. A register file 106 may additionally be included in the processor 102 and may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while other registers may be dedicated to the design of the processor 102.

[0034] In some embodiments, one or more processors 102 are coupled to one or more interface buses 110 to transfer communication signals, such as address, data, or control signals, between the processors 102 and other components in the system 100. In one embodiment, the interface bus 110 can be a processor bus, such as a version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and can include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In one embodiment, the (one or more) processors 102 include an integrated memory controller 116 and a platform controller hub 130. The memory controller 116 facilitates communication between the memory device and other components of the system 100, while the platform controller hub (PCH) 130 provides a connection to I / O devices via a local I / O bus.

[0035] The memory device 120 can be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device having suitable performance to act as process memory. In one embodiment, the memory device 120 can operate as the system memory for the system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. The memory controller 116 is also coupled to an optional external graphics processor 118, which can communicate with one or more graphics processors 108 in the processor 102 to perform graphics operations and media operations. In some embodiments, the graphics operations, media operations, and / or computing operations can be assisted by an accelerator 112, which is a coprocessor that can be configured to perform a set of specialized graphics operations, media operations, or computing operations. For example, in one embodiment, the accelerator 112 is a matrix multiplication accelerator for optimizing machine learning or computing operations. In one embodiment, the accelerator 112 is a ray tracing accelerator, which can be used to perform ray tracing operations in cooperation with the graphics processor 108. In one embodiment, an external accelerator 119 can be used instead of the accelerator 112, or can be used in cooperation with the accelerator 112.

[0036] In some embodiments, the display device 111 may be connected to the processor(s) 102. The display device 111 may be one or more of the following: an internal display device, such as in a mobile electronic device or a laptop computer device; or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 111 may be a head mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.

[0037] In some embodiments, the platform controller hub 130 enables peripheral devices to be connected to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripheral devices include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., non-volatile memory, volatile memory, hard disk drive, flash memory, NAND, 3D NAND, 3D Xpoint, etc.). The data storage device 124 may be connected via a storage interface (e.g., SATA) or via a peripheral bus (such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express)). The touch sensor 125 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, 5G, or Long-Term Evolution (LTE) transceiver. The firmware interface 128 enables communication with the system firmware and may be, for example, a unified extensible firmware interface (UEFI). The network controller 134 enables a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. In one embodiment, the audio controller 146 is a multi-channel high-definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The platform controller hub 130 may also be connected to one or more Universal Serial Bus (USB) controllers 142 to connect to input devices, such as a keyboard and mouse 143 combination, a camera 144, or other USB input devices.

[0038] It will be appreciated that the illustrated system 100 is an example and non-limiting, as other types of data processing systems configured in different ways may also be used. For example, instances of the memory controller 116 and the platform controller hub 130 may be integrated into a discrete external graphics processor, such as the external graphics processor 118. In one embodiment, the platform controller hub 130 and / or the memory controller 116 may be external to one or more processors 102. For example, the system 100 may include an external memory controller 116 and a platform controller hub 130, which may be a memory controller hub and a peripheral controller hub within a system chipset configured to communicate with the processor(s) 102.

[0039] For example, a circuit board ("sled") may be used, on which components such as a CPU, memory, and other components are placed, and on which the components (such as a CPU, memory, and other components) are designed to achieve improved thermal performance. In some examples, processing components such as processors are located on the top side of the sled, while nearby memory such as DIMMs is located on the bottom side of the sled. As a result of the enhanced airflow provided by this design, the components can operate at higher frequencies and power levels than in a typical system, thereby improving performance. In addition, the sled is configured for blind mating of power and data communication cables in a rack, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Similarly, the individual components located on the sled (such as processors, accelerators, memory, and data storage drives) are configured to be easily upgraded due to their increased spacing from each other. In an illustrative embodiment, the components additionally include hardware authentication features for attesting to their authenticity.

[0040] A data center may utilize a single network architecture ("fabric") that supports multiple other network architectures, including Ethernet and Omni-Path. The sled may be coupled to a switch via optical fiber, which provides higher bandwidth and lower latency than typical twisted pair cabling (e.g., Category 5, 5e, 6, etc.). Due to the high-bandwidth, low-latency interconnect and network architecture, the data center in use may centralize physically dispersed resources such as memory, accelerators (e.g., GPUs, graphics accelerators, FPGAs, ASICs, neural network, and / or artificial intelligence accelerators, etc.), and data storage drives, and provide them to computing resources (e.g., processors) on an as-requested basis, enabling the computing resources to access the centralized resources as if the centralized resources were local.

[0041] A power supply or power source can supply voltage and / or current to system 100 or any component or system described herein. In one example, the power supply includes an AC to DC (alternating current to direct current) adapter for insertion into a wall outlet. Such AC power can be a renewable energy (e.g., solar) power source. In one example, the power source includes a DC power source, such as an external AC to DC converter. In one example, the power source or power supply includes wireless charging hardware for charging via a proximity charging field. In one example, the power source can include an internal battery, an AC supply, an action-based power supply, a solar power supply, or a fuel cell source.

[0042] Figures 2A - 2D Illustrates a computing system and a graphics processor provided by the embodiments described herein. Figures 2A - 2D Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0043] Figure 2A Is a block diagram of an embodiment of a processor 200 that has one or more processor cores 202A - 202N, an integrated memory controller 214, and an integrated graphics processor 208. The processor 200 can include additional cores, where the additional cores are up to and including the additional core 202N represented by the dashed box and include the additional core 202N represented by the dashed box. Each of the processor cores 202A - 202N includes one or more internal cache units 204A - 204N. In some embodiments, each processor core also has access to one or more shared cache units 206. The internal cache units 204A - 204N and the shared cache units 206 represent the cache memory hierarchy within the processor 200. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core and one or more levels of shared mid-level cache, such as a second level (L2), third level (L3), fourth level (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence between the cache units 206 and 204A - 204N.

[0044] In some embodiments, the processor 200 may further include a set 216 of one or more bus controller units and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI buses or PCI Express buses. The system agent core 210 provides management functions for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).

[0045] In some embodiments, one or more of the processor cores 202A - 202N include support for simultaneous multithreading operations. In such embodiments, the system agent core 210 includes components for coordinating and operating the cores 202A - 202N during multithreaded processing. The system agent core 210 may additionally include a power control unit (PCU) that includes logic and components for regulating the power states of the processor cores 202A - 202N and the graphics processor 208.

[0046] In some embodiments, the processor 200 additionally includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to a set of shared cache units 206 and the system agent core 210, which includes one or more integrated memory controllers 214. In some embodiments, the system agent core 210 further includes a display controller 211 for driving the graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 may also be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 208.

[0047] In some embodiments, a ring - based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units may be used, such as point - to - point interconnects, switched interconnects, or other techniques, including those well - known in the art. In some embodiments, the graphics processor 208 is coupled to the ring interconnect 212 via an I / O link 213.

[0048] An example I / O link 213 represents at least one of a plurality of various I / O interconnects, including an on - package I / O interconnect that facilitates communication between the various processor components and a high - performance embedded memory module 218, such as an eDRAM module. In some embodiments, each of the processor cores 202A - 202N and the graphics processor 208 may use the embedded memory module 218 as a shared last - level cache.

[0049] In some embodiments, the processor cores 202A - 202N are homogeneous cores that execute the same instruction set architecture. In another embodiment, the processor cores 202A - 202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A - 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of computing power. Additionally, the processor 200 can be implemented on one or more chips or as a SoC integrated circuit that also has the illustrated components in addition to other components.

[0050] Figure 2B is a block diagram of the hardware logic of the graphics processor core 219 according to some embodiments described herein. Figure 2B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. The graphics processor core 219 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. The graphics processor core 219 is an example of a graphics core slice, and based on the target power envelope and performance envelope, a graphics processor as described herein can include multiple graphics core slices. Each graphics processor core 219 can include a fixed function block 230 that is coupled to a plurality of sub - cores 221A - 221F (also referred to as sub - slices), and the plurality of sub - cores 221A - 221F include blocks of modular general - purpose and fixed - function logic.

[0051] In some embodiments, the fixed function block 230 includes a geometry / fixed function pipeline 231 that can be shared by all sub - cores in the graphics processor core 219, for example, in a lower - performance and / or lower - power graphics processor implementation. In embodiments, the geometry / fixed function pipeline 231 includes a 3D fixed function pipeline (e.g., such as the 3D pipeline 312 described below in Figure 3A and Figure 4 , a video front - end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer (e.g., the unified return buffer 418 described below in Figure 4 ).

[0052] In one embodiment, the fixed function block 230 further includes a graphics SoC interface 232, a graphics microcontroller 233, and a media pipeline 234. The graphics SoC interface 232 provides an interface between the graphics processor core 219 and other processor cores within the system-on-chip integrated circuit. The graphics microcontroller 233 is a programmable sub-processor that can be configured to manage various functions of the graphics processor core 219, including thread dispatch, scheduling, and preemption. The media pipeline 234 (e.g., Figure 3A and Figure 4 media pipeline 316) includes logic for facilitating the decoding, encoding, preprocessing, and / or postprocessing of multimedia data including image data and video data. The media pipeline 234 implements media operations via requests to the computation or sampling logic within the sub-cores 221-221F.

[0053] In one embodiment, the SoC interface 232 enables the graphics processor core 219 to communicate with a general-purpose application processor core (e.g., CPU) and / or other components within the SoC, and the other components include memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or embedded on-chip or package-on-chip DRAM. The SoC interface 232 can also enable communication with fixed function devices such as a camera imaging pipeline within the SoC, and enable the use and / or implementation of global memory atomicity, which can be shared between the graphics processor core 219 and the CPU within the SoC. The SoC interface 232 can also implement power management control for the graphics processor core 219, and enable an interface between the clock domain of the graphics core 219 and other clock domains within the SoC. In one embodiment, the SoC interface 232 enables receipt of command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. When a media operation is to be performed, these commands and instructions can be dispatched to the media pipeline 234, or when a graphics processing operation is to be performed, these commands and instructions can be dispatched to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 231, geometry and fixed function pipeline 237).

[0054] The graphics microcontroller 233 can be configured to perform various scheduling tasks and management tasks for the graphics processor core 219. In one embodiment, the graphics microcontroller 233 can perform graphics and / or compute workload scheduling for the execution unit (EU) arrays 222A-222F, 224A-224F within the sub-cores 221A-221F. In this scheduling model, host software executing on the CPU core of the SoC that includes the graphics processor core 219 can submit a workload to one of the multiple graphics processor doorbells, which invokes a scheduling operation for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 233 can also facilitate the low-power or idle state of the graphics processor core 219, thereby providing the ability to save and restore registers within the graphics processor core 219 across low-power state transitions independent of the operating system and / or graphics driver software on the system.

[0055] The graphics processor core 219 can have more or fewer sub-cores 221A-221F than shown, up to N modular sub-cores. For each set of N sub-cores, the graphics processor core 219 can also include shared functional logic 235, shared and / or cache memory 236, a geometry / fixed-function pipeline 237, and additional fixed-function logic 238 for accelerating various graphics and compute processing operations. The shared functional logic 235 can include logic units that are associated with Figure 4 the shared functional logic 420 (e.g., sampler logic, math logic, and / or inter-thread communication logic) and can be shared by each N sub-cores within the graphics processor core 219. The shared and / or cache memory 236 can be the last-level cache for the set of N sub-cores 221A-221F within the graphics processor core 219 and can also act as shared memory that can be accessed by multiple sub-cores. The geometry / fixed-function pipeline 237, rather than the geometry / fixed-function pipeline 231, can be included within the fixed-function block 230, and the geometry / fixed-function pipeline 237 can include the same or similar logic units.

[0056] In one embodiment, the graphics processor core 219 includes additional fixed function logic 238, which may include various fixed function acceleration logics for use by the graphics processor core 219. In one embodiment, the additional fixed function logic 238 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are two geometry pipelines: the full geometry pipeline within the geometry / fixed function pipelines 237, 231; and the culling pipeline, which is an additional geometry pipeline that may be included within the additional fixed function logic 238. In one embodiment, the culling pipeline is a trimmed down version of the full geometry pipeline. The full pipeline and the culling pipeline may execute different instances of the same application, each instance having a separate context. Position-only shading can hide the long culling runs of discarded triangles, thus enabling earlier completion of shading in some instances. For example and in one embodiment, the culling pipeline logic within the additional fixed function logic 238 can execute the position shader in parallel with the main application and generally generate results faster than the full pipeline, because the culling pipeline only takes the position attributes of the vertices and only shades the position attributes of the vertices, without performing rasterization and rendering of pixels to the frame buffer. The culling pipeline can use the generated results to calculate the visibility information of all triangles, regardless of whether those triangles are culled. The full pipeline (which may be referred to as the replay pipeline in this instance) can consume this visibility information to skip the culled triangles, thus only shading the visible triangles that are ultimately passed to the rasterization stage.

[0057] In one embodiment, the additional fixed function logic 238 may further include machine learning acceleration logic, such as fixed function matrix multiplication logic, for an implementation that includes optimizations for machine learning training or inference.

[0058] Each graphics sub-core 221A-221F includes a set of execution resources that can be used to perform graphics operations, media operations, and compute operations in response to requests made by a graphics pipeline, a media pipeline, or a shader program. The graphics sub-cores 221A-221F include: multiple EU arrays 222A-222F, 224A-224F; thread dispatch and inter-thread communication (TD / IC) logic 223A-223F; 3D (e.g., texture) samplers 225A-225F; media samplers 206A-206F; shader processors 227A-227F; and shared local memory (SLM) 228A-228F. The EU arrays 222A-222F, 224A-224F each include a plurality of execution units that are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations to service graphics operations, media operations, or compute operations (including graphics programs, media programs, or compute shader programs). The TD / IC logic 223A-223F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. The 3D samplers 225A-225F can read texture or other 3D graphics-related data into memory. The 3D samplers can read texture data in different ways based on the configured sample state and the texture format associated with a given texture. The media samplers 206A-206F can perform similar read operations based on the type and format associated with the media data. In one embodiment, each graphics sub-core 221A-221F may alternatively include a unified 3D and media sampler. Threads executing on the execution units within each of the sub-cores 221A-221F can utilize the shared local memory 228A-228F within each sub-core to enable threads executing within a thread group to use a common pool of on-chip memory to execute.

[0059] Figure 2C Illustrated is a graphics processing unit (GPU) 239 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 240A-240N. While details are provided for only a single multi-core group 240A, it will be appreciated that the other multi-core groups 240B-240N may be equipped with the same or similar collection of graphics processing resources.

[0060] As illustrated, the multi-core group 240A may include a group of graphics cores 243, a group of tensor cores 244, and a group of ray tracing cores 245. The scheduler / dispatch unit 241 schedules and dispatches graphics threads for execution on the respective cores 243, 244, 245. The set of register files 242 stores operand values used by the cores 243, 244, 245 when executing graphics threads. These register files may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. In one embodiment, the tile registers are implemented as a combined set of vector registers.

[0061] One or more combined level-1 (L1) caches and shared memory units 247 locally store graphics data within each multi-core group 240A, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. One or more texture units 247 may also be used to perform texture operations, such as texture mapping and sampling. The level-2 (L2) cache 253 shared by all multi-core groups 240A - 240N or a subset of multi-core groups 240A - 240N stores graphics data and / or instructions for multiple concurrent graphics threads. As illustrated, the L2 cache 253 may be shared across multiple multi-core groups 240A - 240N. One or more memory controllers 248 couple the GPU 239 to the memory 249, which may be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0062] The input / output (I / O) circuit 250 couples the GPU 239 to one or more I / O devices 252, such as a digital signal processor (DSP), a network controller, or a user input device. On-chip interconnects may be used to couple the I / O devices 252 to the GPU 239 and the memory 249. One or more I / O memory management units (IOMMUs) 251 of the I / O circuit 250 directly couple the I / O devices 252 to the system memory 249. In one embodiment, the IOMMU 251 manages a set of multiple page tables for mapping virtual addresses to physical addresses in the system memory 249. In this embodiment, the I / O devices 252, the (one or more) CPUs 246, and the (one or more) GPUs 239 may share the same virtual address space.

[0063] In one implementation, the IOMMU 251 supports virtualization. In this case, the IOMMU 251 can manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 249). The base address of each of the first set of page tables and the second set of page tables can be stored in a control register and swapped out during a context switch (e.g., such that a new context is provided access to the relevant page table set). Although not illustrated in Figure 2C , each of the cores 243, 244, 245 and / or the multi-core groups 240A - 240N may include translation lookaside buffers (TLBs) for caching guest virtual to guest physical translations, guest physical to host physical translations, and guest virtual to host physical translations.

[0064] In one embodiment, the CPU 246, GPU 239, and I / O device 252 are integrated on a single semiconductor chip and / or chip package. The illustrated memory 249 may be integrated on the same chip or may be coupled to the memory controller 248 via an off-chip interface. In one implementation, the memory 249 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles of the embodiments herein are not limited to this particular implementation.

[0065] In one embodiment, the tensor core 244 includes a plurality of execution units specifically designed to perform matrix operations, which are the basic computational operations for performing deep learning operations. For example, synchronous matrix multiplication operations can be used for neural network training and inference. The tensor core 244 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). In one embodiment, the neural network implementation extracts features of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.

[0066] In a deep learning implementation, schedulable parallel matrix multiplication work can be used for execution on the tensor core 244. The training of neural networks especially utilizes a large number of matrix dot product operations. To handle the inner product formulation of an N x N x N matrix multiplication, the tensor core 244 may include at least N dot product processing elements. Before the matrix multiplication begins, a complete matrix is loaded into the on-chip registers, and for each of the N loops, at least one column of the second matrix is loaded. For each loop, there are N dot products to be processed.

[0067] Depending on the particular implementation, matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for tensor core 244 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads, which can tolerate quantization down to bytes and nibbles).

[0068] In one embodiment, ray tracing core 245 accelerates ray tracing operations for both real-time and non-real-time ray tracing implementations. Specifically, ray tracing core 245 includes a ray traversal / intersection circuit that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays and primitives enclosed within the BVH volume. Ray tracing core 245 may also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, ray tracing core 245 performs traversal and intersection operations in concert with the image denoising techniques described herein, at least part of which may be performed on tensor core 244. For example, in one embodiment, tensor core 244 implements a deep learning neural network to perform denoising of frames generated by ray tracing core 245. However, the (one or more) CPUs 246, graphics core 243, and / or ray tracing core 245 may also implement all or part of the denoising and / or deep learning algorithms.

[0069] In addition, as described above, a distributed approach to denoising can be employed, where GPU 239 is in a computing device coupled to other computing devices via a network or high-speed interconnect. In this embodiment, the interconnected computing devices share neural network learning / training data to improve the speed at which the overall system learns to perform denoising for different types of image frames and / or different graphics applications.

[0070] In one embodiment, the ray tracing core 245 processes all BVH traversals and ray-primitive intersections, freeing the graphics core 243 from being overloaded with thousands of instructions per ray. In one embodiment, each ray tracing core 245 includes a first set of specialized circuitry for performing bounding box tests (e.g., for traversal operations) and a second set of specialized circuitry for performing ray-triangle intersection tests (e.g., intersecting the traversed rays). Thus, in one embodiment, the multi-core group 240A can simply initiate a ray probe, and the ray tracing core 245 independently performs ray traversal and intersection and returns hit data (e.g., hit, miss, multiple hits, etc.) back to the thread context. While the ray tracing core 245 performs traversal and intersection operations, the other cores 243, 244 are freed up to perform other graphics or compute work.

[0071] In one embodiment, each ray tracing core 245 includes a traversal unit for performing BVH test operations and an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "miss", or "multiple hits" responses, which the intersection unit provides to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 243 and the tensor core 244) are freed up to perform other forms of graphics work.

[0072] In a particular embodiment described below, a hybrid rasterization / ray tracing method is used in which work is distributed between the graphics core 243 and the ray tracing core 245.

[0073] In one embodiment, the ray tracing core 245 (and / or other cores 243, 244) includes hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, closest hit shaders, any hit shaders, and miss shaders, which enable assignment of shader and texture sets to each object. Another ray tracing platform that can be supported by the ray tracing core 245, the graphics core 243, and the tensor core 244 is Vulkan 1.1.85. However, note that the basic principles of the embodiments herein are not limited to any particular ray tracing ISA.

[0074] Generally, the individual cores 245, 244, 243 can support a ray tracing instruction set that includes instructions / functions for the following: ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traverse, and exception. More specifically, one embodiment includes ray tracing instructions for performing the following functions:

[0075] Light Generation - Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment.

[0076] Nearest Hit - The closest hit instruction can be executed to locate the closest intersection of a ray with a primitive within a scene.

[0077] Any Hit - Any hit instruction identifies multiple intersections between a ray and a primitive within a scene, potentially identifying a new closest intersection.

[0078] Intersection - The intersection instruction performs a ray-primitive intersection test and outputs the result.

[0079] Primitive Bounding Box Construction - This instruction builds a bounding box around a given primitive or set of primitives (e.g., when building a new BVH or other acceleration data structure).

[0080] Miss - Indicates that the ray misses all geometries within the scene or a specified region of the scene.

[0081] Visit - Indicates the child volume that the ray will traverse.

[0082] Exception - Includes various types of exception handlers (e.g., called for various error conditions).

[0083] Figure 2D is a block diagram of a general-purpose graphics processing unit (GPGPU) 270 according to an embodiment described herein. The GPGPU 270 can be configured as a graphics processor and / or a computing accelerator. The GPGPU 270 can be interconnected with a host processor (e.g., one or more CPUs 246) and memories 271, 272 via one or more system and / or memory buses. In one embodiment, the memory 271 is a system memory that can be shared with one or more CPUs 246, and the memory 272 is a device memory dedicated to the GPGPU 270. In one embodiment, the components within the GPGPU 270 and the device memory 272 can be mapped to memory addresses accessible by one or more CPUs 246. Access to the memories 271 and 272 can be facilitated via a memory controller 268. In one embodiment, the memory controller 268 includes an internal direct memory access (DMA) controller 269, or can include logic for performing operations otherwise performed by the DMA controller.

[0084] The GPGPU 270 includes multiple cache memories, which include the L2 cache 253, the L1 cache 254, the instruction cache 255, and the shared memory 256. At least part of the shared memory 256 can also be partitioned as a cache memory. The GPGPU 270 also includes multiple computing units 260A - 260N. Each computing unit 260A - 260N includes a set of vector registers 261, a set of scalar registers 262, a set of vector logic units 263, and a set of scalar logic units 264. The computing units 260A - 260N may also include a local shared memory 265 and a program counter 266. The computing units 260A - 260N can be coupled to a constant cache 267, which can be used to store constant data, i.e., data that does not change during the execution of a kernel program or a shader program on the GPGPU 270. In one embodiment, the constant cache 267 is a scalar data cache, and the cached data can be directly fetched into the scalar register 262.

[0085] During operation, one or more CPUs 246 can write commands into registers in the GPGPU 270 or into memory in the GPGPU 270 that has been mapped to an accessible address space. The command processor 257 can read the commands from the registers or the memory and determine how to process those commands within the GPGPU 270. Subsequently, a thread dispatcher 258 can be used to dispatch threads to the computing units 260A - 260N to execute those commands. Each computing unit 260A - 260N can execute threads independently of other computing units. In addition, each computing unit 260A - 260N can be independently configured for conditional computing and can conditionally output the results of the computation to memory. When the submitted commands are completed, the command processor 257 can interrupt one or more CPUs 246.

[0086] Figures 3A - 3C A block diagram illustrating an additional graphics processor and computing accelerator architecture provided by the embodiments described herein. Figures 3A - 3C Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0087] Figure 3Ais a block diagram of a graphics processor 300, which can be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores or other semiconductor devices, such as but not limited to memory devices or network interfaces. In some embodiments, the graphics processor communicates via a memory-mapped I / O interface to registers on the graphics processor and uses commands placed in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0088] In some embodiments, the graphics processor 300 also includes a display controller 302 for driving display output data to a display device 318. The display controller 302 includes hardware for one or more overlay planes for the display and the composition of multiple layers of video or user interface elements. The display device 318 can be an internal or external display device. In one embodiment, the display device 318 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, the one or more media encoding formats including but not limited to: Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC, H.265 / HEVC, Alliance for Open Media (AOMedia) VP8, VP9), and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG) format).

[0089] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterizer operations, including, for example, bit-block transfers. However, in one embodiment, one or more components of the graphics processing engine (GPE) 310 are used to perform 2D graphics operations. In some embodiments, the GPE 310 is a computing engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.

[0090] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed-function elements that perform various tasks within the elements and / or generate execution threads to the 3D / media subsystem 315. Although the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 that is dedicated to performing media operations such as video post-processing and image enhancement.

[0091] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units for performing one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of, or on behalf of, the video codec engine 306. In some embodiments, the media pipeline 316 additionally includes a thread generation unit for generating threads for execution on the 3D / media subsystem 315. The generated threads execute computations for media operations on one or more graphics execution units included in the 3D / media subsystem 315.

[0092] In some embodiments, the 3D / media subsystem 315 includes logic for executing the threads generated by the 3D pipeline 312 and the media pipeline 316. In some embodiments, the pipeline sends thread execution requests to the 3D / media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching various requests for available thread execution resources. The execution resources include an array of graphics execution units for processing 3D threads and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes a shared memory for sharing data between threads and for storing output data, which includes registers and addressable memory.

[0093] Figure 3BFIG. illustrates a graphics processor 320 according to an embodiment described herein, the graphics processor 320 having a tiled architecture. In one embodiment, the graphics processor 320 includes a graphics processing engine cluster 322, the graphics processing engine cluster 322 having multiple instances of the graphics processor engine 310 within the graphics engine tiles 310A-310D. Each graphics engine tile 310A-310D can be interconnected via a set of tile interconnects 323A-323F. Each graphics engine tile 310A-310D can also be connected to a memory module or memory device 326A-326D via a memory interconnect 325A-325D. The memory devices 326A-326D can use any graphics memory technology. For example, the memory devices 326A-326D can be graphics double data rate (GDDR) memories. In one embodiment, the memory devices 326A-326D are high-bandwidth memory (HBM) modules, and these HBM modules can be on-die with their respective graphics engine tiles 310A-310D. In one embodiment, the memory devices 326A-326D are stacked memory devices that can be stacked on top of their respective graphics engine tiles 310A-310D. In one embodiment, each graphics engine tile 310A-310D and the associated memory 326A-326D reside on separate dielets, and these separate dielets are bonded to a base die or a base substrate, as further described in detail in Figure 3A Each graphics engine tile 310A-310D can be interconnected via a set of tile interconnects 323A-323F. Each graphics engine tile 310A-310D can also be connected to a memory module or memory device 326A-326D via a memory interconnect 325A-325D. The memory devices 326A-326D can use any graphics memory technology. For example, the memory devices 326A-326D can be graphics double data rate (GDDR) memories. In one embodiment, the memory devices 326A-326D are high-bandwidth memory (HBM) modules, and these HBM modules can be on-die with their respective graphics engine tiles 310A-310D. In one embodiment, the memory devices 326A-326D are stacked memory devices that can be stacked on top of their respective graphics engine tiles 310A-310D. In one embodiment, each graphics engine tile 310A-310D and the associated memory 326A-326D reside on separate dielets, and these separate dielets are bonded to a base die or a base substrate, as further described in detail in Figures 11B - 11D In one embodiment, the memory devices 326A-326D are high-bandwidth memory (HBM) modules, and these HBM modules can be on-die with their respective graphics engine tiles 310A-310D. In one embodiment, the memory devices 326A-326D are stacked memory devices that can be stacked on top of their respective graphics engine tiles 310A-310D. In one embodiment, each graphics engine tile 310A-310D and the associated memory 326A-326D reside on separate dielets, and these separate dielets are bonded to a base die or a base substrate, as further described in detail in

[0094] The graphics processor 320 can be configured with a non-uniform memory access (NUMA) system, in which the memory devices 326A-326D are coupled to the associated graphics engine tiles 310A-310D. A given memory device can be accessed by a graphics engine tile different from the graphics engine tile directly connected to the memory device. However, when accessing the local tile, the access latency to the memory devices 326A-326D can be minimized. In one embodiment, a cache coherent NUMA (ccNUMA) system is enabled, and the ccNUMA system uses the tile interconnects 323A-323F to enable communication between the cache controllers within the graphics engine tiles 310A-310D so as to maintain a consistent memory image when more than one cache stores the same memory location.

[0095] The graphics processing engine cluster 322 can be connected to an on-chip or on-package fabric interconnect 324. The fabric interconnect 324 can enable communication between the graphics engine chips 310A - 310D and components such as the video codec 306 and one or more copy engines 304. The copy engine 304 can be used to move data out of the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory), move data into the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory), and move data between the memory devices 326A - 326D and memory external to the graphics processor (e.g., system memory). The fabric interconnect 324 can also be used to interconnect the graphics engine chips 310A - 310D. The graphics processor 320 can optionally include a display controller 302 for enabling connection to an external display device 318. The graphics processor can also be configured as a graphics or computing accelerator. In the accelerator configuration, the display controller 302 and the display device 318 can be omitted.

[0096] The graphics processor 320 can be connected to a host system via a host interface 328. The host interface 328 can enable communication between the graphics processor 320, the system memory, and / or other system components. The host interface 328 can be, for example, a PCI Express bus or another type of host system interface.

[0097] Figure 3C Illustrated is a computing accelerator 330 according to an embodiment described herein. The computing accelerator 330 can include an architectural similarity to the Figure 3B graphics processor 320 and is optimized for computing acceleration. The compute engine cluster 332 can include a set of compute engine chips 340A - 340D, the set of compute engine chips 340A - 340D including execution logic optimized for parallel or vector-based general computing operations. In some embodiments, the compute engine chips 340A - 340D do not include fixed-function graphics processing logic, but in one embodiment, one or more of the compute engine chips 340A - 340D can include logic for performing media acceleration. The compute engine chips 340A - 340D can be connected to the memories 326A - 326D via memory interconnects 325A - 325D. The memories 326A - 326D and the memory interconnects 325A - 325D can be of a similar technology as in the graphics processor 320 or can be a different technology. The graphics compute engine chips 340A - 340D can also be interconnected via a set of chip interconnects 323A - 323F and can be connected to and / or interconnected by the fabric interconnect 324. In one embodiment, the computing accelerator 330 includes a large L3 cache 336 that can be configured as a device-wide cache. The computing accelerator 330 can also operate in a manner similar toFigure 3B is connected to the host processor and memory via a host interface 328 in a manner similar to that of the graphics processor 320. Graphics Processing Engine

[0098] Figure 4 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is Figure 3A a certain version of the GPE 310 shown in Figure 3B and may also represent Figure 4 the graphics engine slices 310A - 310D of Figure 3A Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. For example,

[0099] In some embodiments, GPE 410 is coupled with or includes a command stream converter 403 that provides a command stream to 3D pipeline 312 and / or media pipeline 316. In some embodiments, command stream converter 403 is coupled with a memory, which may be system memory, or one or more of internal cache memory and shared cache memory. In some embodiments, command stream converter 403 receives commands from the memory and sends these commands to 3D pipeline 312 and / or media pipeline 316. These commands are instructions fetched from a ring buffer that stores commands for 3D pipeline 312 and media pipeline 316. In one embodiment, the ring buffer may additionally include a batch command buffer that stores batches of multiple commands. Commands for 3D pipeline 312 may also include references to data stored in the memory, such as but not limited to vertex data and geometry data for 3D pipeline 312 and / or image data and memory objects for media pipeline 316. 3D pipeline 312 and media pipeline 316 process commands and data by performing operations via logic within the respective pipelines or by dispatching one or more execution threads to graphics core array 414. In one embodiment, graphics core array 414 includes one or more blocks of graphics cores (e.g., (one or more) graphics cores 415A, (one or more) graphics cores 415B), with each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources that includes: general and graphics-specific execution logic for performing graphics operations and computational operations; and fixed-function texture processing logic and / or machine learning and artificial intelligence acceleration logic.

[0100] In embodiments, 3D pipeline 312 may include fixed-function and programmable logic for processing one or more shader programs by processing instructions and dispatching execution threads to graphics core array 414, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. Graphics core array 414 provides a unified block of execution resources for use in processing these shader programs. The multi-functional execution logic (e.g., execution units) within (one or more) graphics cores 415A - 415B of graphics core array 414 includes support for various 3D API shader languages and can execute multiple synchronized execution threads associated with multiple shaders.

[0101] In some embodiments, graphics core array 414 includes execution logic for performing media functions such as video and / or image processing. In one embodiment, in addition to graphics processing operations, the execution units also include general-purpose logic programmable to perform parallel general-purpose computational operations. The general-purpose logic may perform operations in parallel or in combinationFigure 1 one or more of the processor cores 107 or general logic within cores 202A - 202N as Figure 2A executes processing operations.

[0102] Output data generated by threads executing on the graphics core array 414 can output the data to memory in a unified return buffer (URB) 418. The URB 418 can store data for multiple threads. In some embodiments, the URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, the URB 418 can additionally be used for synchronization between threads on the graphics core array and fixed function logic within the shared function logic 420.

[0103] In some embodiments, the graphics core array 414 is scalable such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance levels of the GPE 410. In one embodiment, the execution resources are dynamically scalable such that the execution resources can be enabled or disabled.

[0104] The graphics core array 414 is coupled to shared function logic 420, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide specialized complementary functions to the graphics core array 414. In embodiments, the shared function logic 420 includes, but is not limited to, sampler 421 logic, math 422 logic, and inter - thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.

[0105] Implement shared functionality at least in cases where the demand for a given specialized function is not sufficient to be included within the graphics core array 414. Instead, a single instantiation of that specialized function is implemented as a stand-alone entity within the shared function logic 420 and is shared among the execution resources within the graphics core array 414. The exact set of functions that are shared among and included within the graphics core arrays 414 varies from embodiment to embodiment. In some embodiments, specific shared functions that are widely used by the graphics core arrays 414 within the shared function logic 420 may be included within the shared function logic 416 within the graphics core arrays 414. In various embodiments, the shared function logic 416 within the graphics core arrays 414 may include some or all of the logic within the shared function logic 420. In one embodiment, all of the logic elements within the shared function logic 420 may be replicated within the shared function logic 416 of the graphics core arrays 414. In one embodiment, the shared function logic 420 is excluded in favor of the shared function logic 416 within the graphics core arrays 414. Execution Unit

[0106] Figures 5A - 5B Illustrated is thread execution logic 500 according to an embodiment described herein, the thread execution logic 500 including an array of processing elements employed within a graphics processor core. Figures 5A - 5B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto. Figures 5A - 5B Illustrated is an overview of thread execution logic 500, the thread execution logic 500 which may represent the hardware logic illustrated in each of sub-cores 221A - 221F in Figure 2B Each of sub-cores 221A - 221F in Figure 5A represents execution units within a general-purpose graphics processor, while Figure 5B represents execution units that may be used within a computing accelerator.

[0107] As in Figure 5AAs illustrated, in some embodiments, the thread execution logic 500 includes a shader processor 502, a thread dispatcher 504, an instruction cache 506, a scalable execution unit array including a plurality of execution units 508A - 508N, a sampler 510, a shared local memory 511, a data cache 512, and a data port 514. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 508A, 508B, 508C, 508D, up to 508N - 1 and 508N) based on the computational requirements of the workload. In one embodiment, the included components are interconnected via an interconnect structure that links to each of the components. In some embodiments, the thread execution logic 500 includes one or more connections to memory (such as system memory or cache memory) via the instruction cache 506, the data port 514, the sampler 510, and one or more of the execution units 508A - 508N. In some embodiments, each execution unit (e.g., 508A) is an independent programmable general - purpose computing unit capable of executing multiple synchronous hardware threads and processing multiple data elements in parallel for each thread. In embodiments, the array of execution units 508A - 508N is scalable to include any number of individual execution units.

[0108] In some embodiments, the execution units 508A - 508N are primarily used to execute shader programs. The shader processor 502 can process various shader programs and can dispatch execution threads associated with the shader programs via the thread dispatcher 504. In one embodiment, the thread dispatcher includes logic for arbitrating requests to initiate threads from the graphics pipeline and the media pipeline and instantiating the requested threads on one or more of the execution units 508A - 508N. For example, the geometry pipeline can dispatch a vertex shader, a tessellation shader, or a geometry shader to the thread execution logic for processing. In some embodiments, the thread dispatcher 504 can also process runtime thread generation requests from executed shader programs.

[0109] In some embodiments, execution units 508A - 508N support instruction sets that include native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct 3D and OpenGL) to be executed with minimal translation. These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). Each of the execution units 508A - 508N is capable of multi - issue single instruction multiple data (SIMD) execution, and multithreaded operation enables an efficient execution environment in the face of higher - latency memory accesses. Each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. Execution is multi - issued per clock for pipelines that can perform integer operations, single - precision floating - point operations, and double - precision floating - point operations, that can have SIMD branching capabilities, that can perform logical operations, that can perform transcendental operations, and that can perform other miscellaneous operations. When waiting for data from one of the shared functions in memory or a shared function, the dependency logic within execution units 508A - 508N puts the waiting thread to sleep until the requested data has been returned. While the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution unit can perform operations for a pixel shader, fragment shader, or another type of shader program that includes a different vertex shader. Embodiments can be applied to use execution that utilizes single instruction multiple threads (SIMT), as an alternative to the use of SIMD, or as an addition to the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to a combination of SIMD and SIMT.

[0110] Each of the execution units 508A - 508N operates on an array of data elements. The number of data elements is the "execution size", or the number of channels for an instruction. Execution channels are the logical units for data element access, masking, and flow control execution within an instruction. The number of channels can be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) for a particular graphics processor. In some embodiments, execution units 508A - 508N support integer and floating - point data types.

[0111] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution unit will process the various elements based on the data size of the elements. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 64-bit packed data elements (Quad-Word (QW) size data elements), eight separate 32-bit packed data elements (Double Word (DW) size data elements), sixteen separate 16-bit packed data elements (Word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible.

[0112] In one embodiment, one or more execution units can be combined into fused execution units 509A - 509N, which have thread control logic (507A - 507N) common to the fused EUs. Multiple EUs can be fused into EU groups. Each EU in a fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in a fused EU group can vary according to the embodiment. Additionally, various SIMD widths can be executed on a per-EU basis, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 509A - 509N includes at least two execution units. For example, fused execution unit 509A includes a first EU 508A, a second EU 508B, and thread control logic 507A common to the first EU 508A and the second EU 508B. Thread control logic 507A controls the threads executed on the fused graphics execution unit 509A, allowing each EU within the fused execution units 509A - 509N to execute using a common instruction pointer register.

[0113] One or more internal instruction caches (e.g., 506) are included in the thread execution logic 500 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 512) are included to cache thread data during thread execution. Threads executed on the execution logic 500 can also store explicitly managed data in the shared local memory 511. In some embodiments, a sampler 510 is included to provide texture sampling for 3D operations and media sampling for media operations. In some embodiments, the sampler 510 includes specialized texture or media sampling functions to process texture data or media data during the sampling process before providing the sampled data to the execution units.

[0114] It should be noted that there is an error in the original text where it says "four separate 54-bit packed data elements" which should probably be "four separate 64-bit packed data elements" for consistency with common SIMD data element sizes. This has been corrected in the translation.During execution, the graphics pipeline and the media pipeline send thread launch requests to the thread execution logic 500 via thread generation and dispatch logic. Once a group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 502 is invoked to further compute output information and cause the results to be written to an output surface (e.g., a color buffer, a depth buffer, a stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes values for vertex attributes that are to be interpolated across the rasterized object. In some embodiments, the pixel processor logic within the shader processor 502 then executes a pixel shader program or a fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 502 dispatches threads to execution units (e.g., 508A) via a thread dispatcher 504. In some embodiments, the shader processor 502 uses texture sampling logic in a sampler 510 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometric data compute pixel color data for each geometric fragment or discard one or more pixels without further processing.

[0115] In some embodiments, a data port 514 provides a memory access mechanism for the thread execution logic 500 to output processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, the data port 514 includes or is coupled to one or more cache memories (e.g., a data cache 512) to cache data for memory access via the data port.

[0116] In one embodiment, the execution logic 500 may further include a ray tracer 505 that can provide ray tracing acceleration functionality. The ray tracer 505 may support a ray tracing instruction set that includes instructions / functions for ray generation. The ray tracing instruction set may be similar to or different from the ray tracing instruction set supported by Figure 2C the ray tracing core 245 in

[0117] Figure 5BFIG. illustrates example internal details of execution unit 508 according to an embodiment. The graphics execution unit 508 may include an instruction fetch unit 537, a general register file array (GRF) 524, an architectural register file array (ARF) 526, a thread arbiter 522, a dispatch unit 530, a branch unit 532, a set of SIMD floating point units (FPUs) 534, and in one embodiment, a set of dedicated integer SIMD ALUs 535. The GRF 524 and ARF 526 include a set of general register files and architectural register files associated with each synchronous hardware thread that may be active in the graphics execution unit 508. In one embodiment, the per-thread architectural state is maintained in the ARF 526, while data used during thread execution is stored in the GRF 524. The execution state of each thread, including the instruction pointer for each thread, may be saved in thread-specific registers in the ARF 526.

[0118] In one embodiment, the graphics execution unit 508 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and fine-grained Interleaved Multi-Threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronous threads and the number of registers per execution unit, where execution unit resources are divided across the logic for executing multiple synchronous threads. The number of logical threads that can be executed by the graphics execution unit 508 is not limited to the number of hardware threads, and multiple logical threads may be assigned to each hardware thread.

[0119] In one embodiment, the graphics execution unit 508 may issue multiple instructions in parallel, and these instructions may each be different instructions. The thread arbiter 522 of the graphics execution unit thread 508 may dispatch the instructions to one of the following for execution: the send unit 530, the branch unit 532, or the (one or more) SIMD FPUs 534. Each execution thread may access 128 general-purpose registers within the GRF 524, where each register may store 32 bytes that can be accessed as a SIMD 8-element vector with 32-bit data elements. In one embodiment, each execution unit thread has access to 4 kilobytes within the GRF 524, but the embodiments are not limited thereto, and more or fewer register resources may be provided in other embodiments. In one embodiment, the graphics execution unit 508 is partitioned into seven hardware threads that can independently perform computational operations, but the number of threads per execution unit may also vary according to the embodiment. For example, in one embodiment, up to 16 hardware threads are supported. In an embodiment where seven threads can access 4 kilobytes, the GRF 524 may store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 524 may store a total of 64 kilobytes. Flexible addressing modes may permit addressing of registers together, effectively creating wider registers or representing strided rectangular block data structures.

[0120] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the messaging send unit 530. In one embodiment, branch instructions are dispatched to the dedicated branch unit 532 to facilitate SIMD scatter and eventual gather.

[0121] In one embodiment, the graphics execution unit 508 includes one or more SIMD floating-point units (FPUs) 534 for performing floating-point operations. In one embodiment, the (one or more) FPUs 534 also support integer computations. In one embodiment, the (one or more) FPUs 534 may perform up to a maximum number M of 32-bit floating-point (or integer) operations in SIMD, or up to 2M 16-bit integer or 16-bit floating-point operations in SIMD. In one embodiment, at least one of the (one or more) FPUs provides extended mathematical capabilities that support high-throughput transcendental mathematical functions and double-precision 64-bit floating point. In some embodiments, a set 535 of 8-bit integer SIMD ALUs also exists and may be specifically optimized to perform operations associated with machine learning computations.

[0122] In one embodiment, an array of multiple instances of the graphics execution unit 508 may be instantiated in graphics sub-core groupings (e.g., sub-slices). For scalability, the product architect may choose the exact number of execution units per sub-core grouping. In one embodiment, the execution unit 508 may execute instructions across multiple execution channels. In a further embodiment, each thread executed on the graphics execution unit 508 is executed on a different channel.

[0123] Figure 6 FIG. illustrates an additional execution unit 600 according to an embodiment. The execution unit 600 may be a compute-optimized execution unit for use in compute engine slices 340A - 340D in, for example, Figure 3C but is not limited thereto. Variants of the execution unit 600 may also be used in Figure 3B graphics engine slices 310A - 310D in. In one embodiment, the execution unit 600 includes a thread control unit 601, a thread state unit 602, an instruction fetch / prefetch unit 603, and an instruction decoding unit 604 (also referred to herein as a decoder). The execution unit 600 additionally includes a register file 606 that stores registers that may be assigned to hardware threads within the execution unit. The execution unit 600 additionally includes a dispatch unit 607 and a branch unit 608. In one embodiment, the dispatch unit 607 and the branch unit 608 can operate in a manner similar to Figure 5B the dispatch unit 530 and the branch unit 532 of the graphics execution unit 508 of

[0124] The execution unit 600 further includes a computing unit 610, and the computing unit 610 includes a plurality of different types of functional units. In one embodiment, the computing unit 610 includes an ALU unit 611, and the ALU unit 611 includes an array of arithmetic logic units. The ALU unit 611 can be configured to perform 64-bit, 32-bit, and 16-bit integer and floating-point operations. Integer and floating-point operations can be performed simultaneously. The computing unit 610 may further include a systolic array 612 and a math unit 613. The systolic array 612 includes a network of data processing units that is W wide and D deep, which can be used to perform vector or other data parallel operations in a systolic manner. In one embodiment, the systolic array 612 can be configured to perform matrix operations (such as matrix dot product operations). In one embodiment, the systolic array 612 supports 16-bit floating-point operations, as well as 8-bit and 4-bit integer operations. In one embodiment, the systolic array 612 can be configured to accelerate machine learning operations. In such embodiments, the systolic array 612 can be configured with support for the bfloat 16-bit floating-point format. In one embodiment, the math unit 613 may be included to perform a specific subset of math operations in an efficient and lower-power manner than the ALU unit 611. The math unit 613 may include a variant of the math logic (e.g., Figure 4 the math logic 422 of the shared functional logic 420 in

[0125] The thread control unit 601 includes logic for controlling the execution of threads within the execution unit. The thread control unit 601 may include thread arbitration logic for starting, stopping, and pre-empting the execution of threads within the execution unit 600. The thread status unit 602 can be used to store the thread status for the threads assigned to execute on the execution unit 600. Storing the thread status within the execution unit 600 enables those threads to be quickly pre-empted when they become blocked or idle. The instruction fetch / prefetch unit 603 can fetch instructions from the instruction cache of a higher-level execution logic (e.g., the instruction cache 506 as in Figure 5A ). The instruction fetch / prefetch unit 603 can also issue a prefetch request for instructions to be loaded into the instruction cache based on an analysis of the current execution thread. The instruction decoding unit 604 can be used to decode the instructions to be executed by the computing unit. In one embodiment, the instruction decoding unit 604 can be used as a secondary decoder to decode complex instructions into constituent micro-operations.

[0126] The execution unit 600 additionally includes a register file 606 that can be used by hardware threads executing on the execution unit 600. The registers in the register file 606 can be partitioned across the logic for multiple synchronous threads within the compute units 610 that execute the execution unit 600. The number of logical threads that can be executed by the graphics execution unit 600 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. Based on the number of supported hardware threads, the size of the register file 606 can vary across embodiments. In one embodiment, register renaming can be used to dynamically allocate registers to hardware threads.

[0127] Figure 7 FIG. is a block diagram illustrating a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, a graphics processor execution unit supports an instruction set of instructions in multiple formats. The solid boxes illustrate components that are typically included in execution unit instructions, while the dashed boxes include optional or components that are only included in a subset of the instructions. In some embodiments, the instruction formats 700 described and illustrated are macro-instructions because they are the instructions supplied to the execution unit, as opposed to micro-operations that result from instruction decoding once the instruction has been processed.

[0128] In some embodiments, the graphics processor execution unit natively supports instructions in a 128-bit instruction format 710. Based on the selected instructions, instruction options, and number of operands, a 64-bit compact instruction format 730 can be used for some instructions. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The native instructions available in the 64-bit format 730 vary across embodiments. In some embodiments, a set of index values in an index field 713 is used to partially compress the instruction. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in the 128-bit instruction format 710. Instructions of other sizes and formats can be used.

[0129] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 710, the execution size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.

[0130] Some execution unit instructions have up to three operands, including two source operands src0 720, src1 722, and one destination 718. In some embodiments, the execution unit supports dual-destination instructions, where one of the destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction may be an immediate (e.g., hard-coded) value passed with the instruction.

[0131] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that, for example, specifies whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are provided directly by bits in the instruction.

[0132] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that specifies the addressing mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment of the instruction. Some embodiments support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands.

[0133] In one embodiment, the addressing mode portion of the access / addressing mode field 726 determines whether the instruction is to use direct addressing or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register address of one or more operands. When using indirect register addressing mode, the register address of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.

[0134] In some embodiments, the instructions are grouped based on the opcode 712 bit field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of the opcode. The exact opcode grouping shown is merely an example. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares the five most significant bits (MSB), where the move (mov) instruction takes the form of 0000xxxxb and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jump) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel math group 748 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel math group 748 performs arithmetic operations in parallel across data channels. The vector math group 750 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In one embodiment, the illustrated opcode decoding 740 can be used to determine which part of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by the systolic array. Other instructions (such as ray tracing instructions (not shown)) can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic. Graphics Pipeline

[0135] Figure 8 is a block diagram of another embodiment of the graphics processor 800. Figure 8 Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.

[0136] In some embodiments, graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a render output pipeline 870. In some embodiments, graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued through ring interconnect 802 to graphics processor 800. In some embodiments, ring interconnect 802 couples graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Command stream converter 803 interprets commands from ring interconnect 802 and supplies instructions to the various components of geometry pipeline 820 or media pipeline 830.

[0137] In some embodiments, command stream converter 803 directs the operation of vertex fetcher 805, which reads vertex data from memory and executes vertex processing commands provided by command stream converter 803. In some embodiments, vertex fetcher 805 provides vertex data to vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, vertex fetcher 805 and vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A - 852B via thread dispatcher 831.

[0138] In some embodiments, execution units 852A - 852B are an array of vector processors having instruction sets for performing graphics operations and media operations. In some embodiments, execution units 852A - 852B may have attached L1 caches 851 dedicated to each array or shared between the arrays. The caches may be configured as data caches, instruction caches, or a single cache partitioned to contain data and instructions in different partitions.

[0139] In some embodiments, geometry pipeline 820 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some embodiments, programmable hull shader 811 configures the tessellation operation. Programmable domain shader 817 provides backend evaluation of the tessellation output. Tessellator 813 operates under the direction of hull shader 811 and includes specialized logic for generating a detailed set of geometric objects based on a rough geometric model that is provided as input to geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation component (e.g., hull shader 811, tessellator 813, and domain shader 817) can be bypassed. The tessellation component may operate based on data received from vertex shader 807.

[0140] In some embodiments, a complete geometric object may be processed by the geometry shader 819 via one or more threads dispatched to execution units 852A - 852B, or may proceed directly to the clipper 829. In some embodiments, the geometry shader operates on the entire geometric object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 is programmable by a geometry shader program to perform geometric tessellation in cases where the tessellation unit is disabled.

[0141] Before rasterization, the clipper 829 processes vertex data. The clipper 829 may be a fixed - function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer and depth test component 873 in the render output pipeline 870 dispatch the pixel shader to convert the geometric object into a per - pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, the application may bypass the rasterizer and depth test component 873 and access the un - rasterized vertex data via the egress unit 823.

[0142] The graphics processor 800 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some embodiments, the execution units 852A - 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via a data port 856 to perform memory access and communicate with the render output pipeline components of the processor. In some embodiments, the sampler 854, caches 851, 858, and execution units 852A - 852B each have separate memory access paths. In one embodiment, the texture cache 858 may also be configured as a sampler cache.

[0143] In some embodiments, the rendering output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects to associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. In some embodiments, associated render cache 878 and depth cache 879 are also available. Pixel operation component 877 performs pixel-based operations on the data, but in some instances, pixel operations associated with 2D operations (e.g., bit-block transfer with blending) are performed by 2D engine 841 or, when displaying, by display controller 843 using an overlay display plane instead. In some embodiments, shared L3 cache 875 is available to all graphics components, allowing data to be shared without using main system memory.

[0144] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front end 834 processes the media commands before sending them to media engine 837. In some embodiments, media engine 837 includes a thread generation function for generating threads to be dispatched to thread execution logic 850 via thread dispatcher 831.

[0145] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and is coupled to the graphics processor via a ring interconnect 802, or some other interconnect bus or fabric. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 contains dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system-integrated display device such as in a laptop computer or an external display device attached via a display device connector.

[0146] In some embodiments, the geometry pipeline 820 and the media pipeline 830 can be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). In some embodiments, the driver software for the graphics processor converts API calls dedicated to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. In some embodiments, support can also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, combinations of these libraries can be supported. Support can also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the pipeline of the future API to the pipeline of the graphics processor. Graphics Pipeline Programming

[0147] Figure 9A is a block diagram illustrating a graphics processor command format 900 according to some embodiments. Figure 9B is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 9A The solid box diagrams in generally include components that are typically included in a graphics command, while the dashed lines include components that are optional or are only included in a subset of the graphics commands. Figure 9A An example graphics processor command format 900 includes data fields for a client 902 that identifies the command, a command operation code (opcode) 904, and data 906. A sub-opcode 905 and a command size 908 are also included in some commands.

[0148] In some embodiments, client 902 designates a client unit of the graphics device that processes command data. In some embodiments, the graphics processor command parser examines the client field of each command to adjust further processing of the command and routes the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 904 and sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses the information in the data field 906 to execute the command. For some commands, an explicit command size 908 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands in the command based on the command opcode. In some embodiments, commands are aligned by multiples of doublewords. Other command formats may be used.

[0149] Figure 9B The flowchart example in FIG. shows a graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of the graphics processor uses a version of the shown command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only, as embodiments are not limited to these particular commands or this command sequence. Additionally, commands may be issued in a batch in the command sequence such that the graphics processor will process the command sequence in at least a partially concurrent manner.

[0150] In some embodiments, the graphics processor command sequence 910 can begin with a pipeline flush clear command 912 to cause any active graphics pipeline to complete the current outstanding commands for the pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate concurrently. Executing the pipeline flush clear causes the active graphics pipeline to complete any outstanding commands. In response to the pipeline flush clear, the command parser for the graphics processor will pause command processing until the active drawing engine has completed the outstanding operations and the associated read cache has been invalidated. Optionally, any data marked "dirty" in the render cache may be flushed to memory. In some embodiments, the pipeline flush clear command 912 can be used for pipeline synchronization or can be used before putting the graphics processor in a low power state.

[0151] In some embodiments, the pipeline select command 913 is used when a command sequence explicitly switches between pipelines using a graphics processor. In some embodiments, the pipeline select command 913 is utilized only once in an execution context before issuing pipeline commands, unless the context is to issue commands for both pipelines. In some embodiments, the pipeline dump clear command 912 is utilized immediately before a pipeline switch via the pipeline select command 913.

[0152] In some embodiments, the pipeline control command 914 is configured to operate the graphics pipeline and to program the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state for the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands.

[0153] In some embodiments, the return buffer status command 916 is used to configure a set of return buffers for a corresponding pipeline for writing data. Some pipeline operations utilize the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and to perform cross-thread communication. In some embodiments, the return buffer status 916 includes selecting the size and number of return buffers to be used for a set of pipeline operations.

[0154] The remaining commands in the command sequence differ based on the active pipeline for the operation. Based on the pipeline determination 920, the command sequence is customized for the 3D pipeline 922 starting with the 3D pipeline state 930, or for the media pipeline 924 starting at the media pipeline state 940.

[0155] Commands for configuring the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the particular 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass certain pipeline elements in cases where those elements will not be used.

[0156] In some embodiments, the 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 commands are used to perform vertex operations on the 3D primitives via the vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.

[0157] In some embodiments, the 3D pipeline 922 is triggered via the execute 934 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a "go" or "kick" command in the command sequence. In some embodiments, command execution uses pipeline synchronization commands to trigger a command sequence dump clear through the graphics pipeline. The 3D pipeline will perform geometric processing on the 3D primitives. Once the operation is complete, the resulting geometric object is rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands may also be included for controlling pixel coloring and pixel backend operations.

[0158] In some embodiments, when performing media operations, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the specific uses and ways of programming the media pipeline 924 depend on the media or computing operations to be performed. During media decoding, specific media decoding operations may be migrated to the media pipeline. In some embodiments, the media pipeline may also be bypassed, and media decoding may be performed entirely or partially using the resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0159] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. The set of commands for configuring the media pipeline state 940 is dispatched or placed into the command sequence before the media object commands 942. In some embodiments, the commands for the media pipeline state 940 include data for configuring the media pipeline elements that will be used to process the media object. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands for the media pipeline state 940 also support the use of one or more pointers to "indirect" state elements that point to a batch of state settings.

[0160] In some embodiments, the media object command 942 supplies a pointer to a media object for processing by a media pipeline. The media object includes a memory buffer that contains video data to be processed. In some embodiments, all media pipeline states should be valid before the media object command 942 is issued. Once the pipeline state is configured and the media object command 942 is queued, the media pipeline 924 is triggered via an execute command 944 or an equivalent execution event (e.g., a register write). Subsequently, the output from the media pipeline 924 can be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations. Graphics Software Architecture

[0161] Figure 10 FIG. illustrates an example graphics software architecture for a data processing system 1000 according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.

[0162] In some embodiments, the 3D graphics application 1010 includes one or more shader programs that include shader instructions 1012. The shader language instructions may be in a high-level shader language, such as the High-Level Shader Language (HLSL) of Direct3D, the OpenGL Shader Language (GLSL), and so on. The application also includes executable instructions 1014 in machine language suitable for execution by the general-purpose processor cores 1034. The application also includes a graphics object 1016 defined by vertex data.

[0163] In some embodiments, the operating system 1020 is from Microsoft Corporation An operating system, an exclusive UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. The operating system 1020 can support a graphics API 1022, such as, Direct3D API, OpenGL API, or Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation can be just-in-time (JIT) compilation or application executable shader pre-compilation. In some embodiments, during the compilation of the 3D graphics application 1010, high-level shaders are compiled into low-level shaders. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the Standard Portable Intermediate Representation (SPIR) used by the Vulkan API.

[0164] In some embodiments, the user-mode graphics driver 1026 includes a backend shader compiler 1027 to compile the shader instructions 1012 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses the operating system kernel-mode function 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions. IP Core Implementation

[0165] One or more aspects of at least one embodiment can be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium can include instructions representing various logics within the processor. In some embodiments, the machine-readable medium is also referred to herein as a computer-readable medium or a non-transitory computer-readable medium. When read by a machine, the instructions can cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit, and these reusable units can be stored as a hardware model describing the organization of the integrated circuit on a tangible, machine-readable medium. The hardware model can be supplied to each customer or manufacturing facility that loads the hardware model on a manufacturing machine for fabricating the integrated circuit. The integrated circuit can be fabricated such that the circuit performs the operations described in association with any of the embodiments described herein.

[0166] Figure 11A FIG. is a block diagram of an IP core development system 1100 that can be used to manufacture integrated circuits to perform operations according to an embodiment. The IP core development system 1100 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build an entire integrated circuit (e.g., a system-on-a-chip integrated circuit). A design facility 1130 can generate a software simulation 1110 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 can include functional simulation, behavioral simulation, and / or timing simulation. Subsequently, a register transfer level (RTL) design 1115 can be created or synthesized from the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic performed using the modeled digital signals). In addition to the RTL design 1115, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0167] The RTL design 1115 or an equivalent can be further synthesized by the design facility into a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted via a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The manufacturing facility 1165 can then manufacture an integrated circuit that is at least partially based on the IP core design. The manufactured integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0168] Figure 11BFIG. is a cross-sectional side view of an integrated circuit package component 1170 in accordance with some embodiments described herein. The integrated circuit package component 1170 illustrates an implementation of one or more processor or accelerator devices as described herein. The package component 1170 includes a plurality of hardware logic units 1172, 1174 coupled to a substrate 1180. The logic 1172, 1174 may be implemented at least partially in configurable logic or fixed-function logic hardware and may include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, or other accelerator devices described herein. Each logic unit 1172, 1174 may be implemented within a semiconductor die and is coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic 1172, 1174 and the substrate 1180 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 1172, 1174. In some embodiments, the substrate 1180 is an epoxy-based laminated substrate. In other embodiments, the substrate 1180 may include other suitable types of substrates. The package component 1170 may be connected to other electrical devices via package interconnects 1183. The package interconnects 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or a multi-chip module.

[0169] In some embodiments, the logic units 1172, 1174 are electrically coupled to a bridge 1182 that is configured to route electrical signals between the logic 1172 and the logic 1174. The bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. The bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide chip-to-chip connections between the logic 1172 and the logic 1174.

[0170] Although two logic units 1172, 1174 and a bridge 1182 are illustrated, the embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges since the bridge 1182 may be excluded when the logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Additionally, multiple logic units, dies, and bridges may be connected together in other possible configurations including three-dimensional configurations.

[0171] Figure 11CIllustrated is a packaged component 1190 that includes hardware logic die that are connected to a substrate 1180 (e.g., a base die). Graphics processing units, parallel processors, and / or compute accelerators as described herein may be composed of various silicon die that are manufactured separately. In this context, a die is an integrated circuit that is at least partially packaged and that includes different logic units that can be assembled with other die into a larger package. Die with various collections of different IP core logic may be assembled into a single device. Additionally, die may be integrated into a base die or base die using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP cores may be manufactured using different process technologies and configured during manufacturing, which avoids the complexity of converging multiple IPs into the same manufacturing process, especially for large SoCs with several styles of IP. Allowing the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. Additionally, decomposed IP is more easily modified to be independently power gated, and components not in use for a given workload can be turned off, thus reducing overall power consumption.

[0172] The hardware logic die may include dedicated hardware logic die 1172, logic or I / O die 1174, and / or memory die 1175. The hardware logic die 1172 and the logic or I / O die 1174 may be implemented at least partially in configurable logic or fixed function logic hardware and may include one or more portions of any of the (one or more) processor cores, (one or more) graphics processors, parallel processors, or other accelerator devices described herein. The memory die 1175 may be DRAM (e.g., GDDR, HBM) memory or cache (SRAM) memory.

[0173] Each die may be manufactured as a separate semiconductor die and coupled to the substrate 1180 via an interconnect fabric 1173. The interconnect fabric 1173 may be configured to route electrical signals between the various die and logic within the substrate 1180. The interconnect fabric 1173 may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect fabric 1173 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic, I / O, and memory die.

[0174] In some embodiments, the substrate 1180 is an epoxy-based laminated substrate. In other embodiments, the substrate 1180 may include other suitable types of substrates. The package assembly 1190 may be connected to other electrical devices via the package interconnect 1183. The package interconnect 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices, such as a motherboard, other chip sets, or a multi-chip module.

[0175] In some embodiments, the logic or I / O die 1174 and the memory die 1175 may be electrically coupled via a bridge 1187, which is configured to route electrical signals between the logic or I / O die 1174 and the memory die 1175. The bridge 1187 may be a dense interconnect fabric that provides routing for electrical signals. The bridge 1187 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide chip-to-chip connections between the logic or I / O die 1174 and the memory die 1175. The bridge 1187 may also be referred to as a silicon bridge or an interconnect bridge. For example, in some embodiments, the bridge 1187 is an Embedded Multi-die Interconnect Bridge (EMIB). In some embodiments, the bridge 1187 may simply be a direct connection from one die to another die.

[0176] The substrate 1180 may include hardware components for I / O 1191, cache memory 1192, and other hardware logic 1193. The structure 1185 may be embedded in the substrate 1180 to enable communication between various logic dies within the substrate 1180 and the logic 1191, 1193. In one embodiment, the I / O 1191, the structure 1185, the cache, the bridge, and other hardware logic 1193 may be integrated into a base die stacked on top of the substrate 1180. The structure 1185 may be a network-on-chip interconnect or another form of packet-switched fabric that exchanges data packets between components of the package assembly.

[0177] In various embodiments, the packaged component 1190 may include fewer or more components and dies interconnected by a structure 1185 or one or more bridges 1187. The dies within the packaged component 1190 can be arranged in a 3D arrangement or a 2.5D arrangement. Generally, the bridge fabric 1187 can be used to facilitate point-to-point interconnects, such as between logic or I / O dies and memory dies. The structure 1185 can be used to interconnect various logic and / or I / O dies (e.g., dies 1172, 1174, 1191, 1193) with other logic and / or I / O dies. In one embodiment, the cache memory 1192 within the substrate can act as a global cache for the packaged component 1190, act as part of a distributed global cache, or act as a dedicated cache for the structure 1185.

[0178] Figure 11D FIG. illustrates a packaged component 1194 including interchangeable dies 1195 according to an embodiment. The interchangeable dies 1195 can be assembled into standardized slots on one or more base dies 1196, 1198. The base dies 1196, 1198 can be coupled via a bridge interconnect 1197, which can be similar to other bridge interconnects described herein and can be, for example, an EMIB. Memory dies can also be connected to logic or I / O dies via a bridge interconnect. The I / O and logic dies can communicate via an interconnect structure. Each of the base dies can support one or more slots in a standardized format for one of logic or I / O or memory / cache.

[0179] In one embodiment, SRAM and power delivery circuitry can be fabricated into one or more of the base dies 1196, 1198, which can be fabricated using a different process technology relative to the interchangeable dies 1195, with the interchangeable dies 1195 stacked on top of the base dies. For example, the base dies 1196, 1198 can be fabricated using a larger process technology, while the interchangeable dies can be fabricated using a smaller process technology. One or more of the interchangeable dies 1195 can be memory (e.g., DRAM) dies. Different memory densities can be selected for the packaged component 1194 based on the power and / or performance of the product using the packaged component 1194. Additionally, logic dies with different numbers of types of functional units can be selected based on the power and / or performance of the product at the time of assembly. Further, dies containing different types of IP logic cores can be inserted into the interchangeable die slots, enabling a hybrid processor design that can mix and match IP blocks of different technologies. Example System-on-Chip Integrated Circuit

[0180] Figures 12 - 13BFIG. illustrates example integrated circuits and associated graphics processors that may be manufactured using one or more IP cores according to various embodiments described herein. In addition to what is illustrated, other logic and circuitry may be included, including additional graphics processor(s) / core(s), peripheral interface controllers, or general purpose processor cores.

[0181] Figure 12 is a block diagram of an example system-on-chip integrated circuit 1200 that may be manufactured using one or more IP cores according to an embodiment. The example integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210 and may additionally include an image processor 1215 and / or a video processor 1220, either of which may be a modular IP core from the same design facility or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235 and I 2 S / I 2 C controller 1240. Additionally, the integrated circuit may include a display device 1245 that is coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display interface 1255. Storage may be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface may be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 1270.

[0182] Figures 13A - 13B is a block diagram of an example graphics processor for use within a SoC according to embodiments described herein. Figure 13A illustrates an example graphics processor 1310 of a system-on-chip integrated circuit that may be manufactured using one or more IP cores according to an embodiment. Figure 13B illustrates an additional example graphics processor 1340 of a system-on-chip integrated circuit that may be manufactured using one or more IP cores according to an embodiment. Figure 13A The graphics processor 1310 is an example of a low-power graphics processor core. Figure 13B The graphics processor 1340 is an example of a higher-performance graphics processor core. Each of the graphics processors 1310, 1340 may be Figure 12 a variant of the graphics processor 1210.

[0183] As shown Figure 13A in FIG. 483, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A-1315N (e.g., 1315A, 1315B, 1315C, 1315D, up to 1315N-1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic such that the vertex processor 1305 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1315A-1315N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. The vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. The (one or more) fragment processors 1315A-1315N use the primitive data and vertex data generated by the vertex processor 1305 to produce a frame buffer that is displayed on a display device. In one embodiment, the (one or more) fragment processors 1315A-1315N are optimized to execute fragment shader programs as provided in the OpenGL API, and these fragment shader programs can be used to perform operations similar to those of pixel shader programs as provided in the Direct 3D API.

[0184] The graphics processor 1310 additionally includes one or more memory management units (MMUs) 1320A-1320B, (one or more) caches 1325A-1325B, and (one or more) circuit interconnects 1330A-1330B. The one or more MMUs 1320A-1320B provide virtual-to-physical address mapping for the graphics processor 1310 (including for the vertex processor 1305 and / or the (one or more) fragment processors 1315A-1315N), and this virtual-to-physical address mapping can reference vertex data or image / texture data stored in memory in addition to vertex data or image / texture data stored in the one or more caches 1325A-1325B. In one embodiment, the one or more MMUs 1320A-1320B can be synchronized with other MMUs within the system such that each processor 1205-1220 can participate in a shared or unified virtual memory system, and the other MMUs within the system include one or more MMUs associated with Figure 12 one or more application processors 1205, image processors 1215, and / or video processors 1220. According to an embodiment, the one or more circuit interconnects 131330A-1330B enable the graphics processor 1310 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0185] As Figure 13BAs shown, the graphics processor 1340 includes Figure 13A one or more MMUs 1320A - 1320B, caches 1325A - 1325B, and circuit interconnects 1330A - 1330B of the graphics processor 1310. The graphics processor 1340 includes one or more shader cores 1355A - 1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N - 1 and 1355N), which provide a unified shader core architecture where a single core or any type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary depending on the embodiment and implementation. Additionally, the graphics processor 1340 includes an inter - core task manager 1345, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1355A - 1355N and a tiling unit 1358 for accelerating tiling operations for tile - based rendering, where rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or to optimize the use of internal caches.

[0186] In some embodiments, as described herein, processing resources represent processing elements (e.g., GPGPU cores, ray - tracing cores, tensor cores, execution resources, execution units (EUs), stream processors, streaming multiprocessors (SMs), graphics multiprocessors) associated with a graphics processor or the graphics processor structure in a GPU (e.g., parallel processing units, graphics processing engines, multi - core groups, computing units, computing units of the next - generation graphics cores). For example, processing resources can be: a GPGPU core of a graphics multiprocessor, or one of the tensor / ray - tracing cores; a ray - tracing core, tensor core, or GPGPU core of a graphics multiprocessor; an execution resource of a graphics multiprocessor; one of the GFX cores, tensor cores, or ray - tracing cores of a multi - core group; one of the vector logic units or scalar logic units of a computing unit; an execution unit with an EU array or an EU array; an execution unit of execution logic; and / or an execution unit. Processing resources can also be, for example, a graphics processing engine, a processing cluster, a GPGPU, a GPGPU, a graphics processing engine, a graphics processing engine cluster, and / or an execution resource within a graphics processing engine. Processing resources can also be a graphics processor, a graphics processor, and / or processing resources within a graphics processor. Correctable Error Address Filtering in a Processing Architecture

[0187] Parallel computing is a type of computing in which the execution of many computations or processes is performed simultaneously. Parallel computing can take various forms, including but not limited to SIMD or SIMT. SIMD describes a computer with multiple processing elements that perform the same operation on multiple data points simultaneously. In one example, the Figures 5A - 5B refers to SIMD and its implementation in general-purpose processors in terms of EUs, FPUs, and ALUs. In a common SIMD machine, data is packed into registers, and each register contains an array of lanes. Instructions operate on the data found in lane n of one register and the data found in the same lane of another register. SIMD machines are advantageous in areas where a single instruction sequence can be applied simultaneously to a large amount of data. For example, in one embodiment, a graphics processing unit (e.g., GPGPU, GPU, etc.) can be used to perform SIMD vector operations using a compute shader program.

[0188] Embodiments can also be applied to use execution via Single Instruction Multiple Thread (SIMT) as an alternative to, or in addition to, the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to a combination of SIMD and SIMT. The following description is discussed in terms of SIMD machines. However, the embodiments herein are not limited to applications in the SIMD context and can also be applied to other parallel computing paradigms, such as, for example, SIMT. For ease of discussion and explanation, the following description generally focuses on SIMD implementations. However, the embodiments can be similarly applied to SIMT machines without modifying the described techniques and methods. Regarding SIMT machines, a similar pattern as discussed below can be followed to provide instructions to a systolic array and execute the instructions on a SIMT machine. Other types of parallel computer machines can also utilize the embodiments herein.

[0189] Embodiments can implement a GPGPU with a matrix acceleration circuitry system. Such a matrix acceleration circuitry system can be used to accelerate machine learning (ML) operations. A GPGPU with a matrix acceleration circuitry system is often deployed and hosted in a data center. Hardware resilience is a requirement for graphics architectures in the data center market segment. Architectures with reliability, availability, and serviceability (RAS) features are designed to meet the resilience goals. Reliability refers to how reliable the design's operation is. Availability refers to the uptime of the design's operation that can still be provided in the presence of errors. Serviceability refers to how easily the design can be serviced to resume reliable operation once an error occurs.

[0190] To do this, the design should be able to detect and record errors (to improve reliability), correct these errors if possible (to improve availability), and report to a higher-level system component (such as a driver) when the errors are uncorrectable. The system software can then take appropriate measures to service the error and bring the design back to reliable operation.

[0191] Server components (such as those operating in a data center and including CPUs and GPUs that support AI / ML operations) have high expectations for providing RAS capabilities. This includes detecting errors, recording errors, and reporting errors to a central component responsible for evaluating and resolving errors. One problem encountered when detecting, recording, and reporting errors in such systems is the over-reporting of correctable errors.

[0192] Correctable errors are those errors that can be corrected by hardware without intervention from the software provided for continued operation. For example, a storage device with an ECC SECDED (single error correction and double error detection) scheme can automatically correct single-bit errors. Errors on a storage device with parity protection and a replay / rewind mechanism are also considered correctable errors.

[0193] Uncorrectable errors are errors that cannot be corrected by hardware. These errors should be reported to the software for error recovery. Structural errors (such as double-bit errors on an ECC-protected storage device), single-bit errors on a parity-protected storage device, and functional protocol errors (such as an inability to complete a write operation) are examples of uncorrectable errors. Recovery from such errors depends on the impact of the error.

[0194] Regarding correctable errors, even though these types of errors can be corrected by hardware without intervention, such errors are still reported because information about correctable errors is relevant to the operation and performance of the system. In some cases, correctable errors may be over-reported. This is because, in some cases, correctable errors cannot be corrected until the value is overwritten. Repeated reads of an address with a correctable error may cause multiple error reports for the same individual error, resulting in a hard fault in the device when this is not the case.

[0195] The implementations herein solve the above technical problems by providing correctable error address filtering in a processing architecture. In one implementation, a method is provided for filtering out correctable errors at a correctable error source using address filtering. The implementation provides correctable error address filtering circuitry in each unit (such as a system of a graphics architecture) for detecting errors in the system. The correctable error address filtering circuitry of each unit is configured to control when a unit that has detected a correctable error should report the correctable error to a higher-level system component in an aggregation system for error reporting. The correctable error address filtering circuitry of a unit may keep track of the last "N" correctable errors detected by that unit, including the addresses associated with each detected correctable error.

[0196] In the implementations herein, when a new correctable error is detected, the address of the correctable error is compared to entries maintained by the correctable error address filtering circuitry that keeps track of the last "N" correctable errors. In one implementation, the correctable error address filtering circuitry may utilize a data structure such as a table to keep track of the last "N" correctable errors. If an address match is detected, the correctable error is not reported to a higher-level system component (e.g., the correctable error is masked). If an address match is not detected, the correctable error is reported. In the case where a correctable error is associated with a write to an address and a match is found by the correctable error address circuitry, the corresponding entry in the correctable error address filtering circuitry is cleared.

[0197] The correctable error address filtering circuitry of the implementations herein reduces the number of duplicate correctable errors reported in a system (such as a processing architecture including, but not limited to, a graphics architecture). This results in reduced hard fault reporting in the system, lower hardware costs and technical support supervision, and improved perceived performance of the system.

[0198] Figure 14 FIG. illustrates an embodiment in accordance with an implementation herein of a processor 1400 including correctable error address filtering. In one implementation, the processor 1400 is a graphics processor 1400. In one implementation, the graphics processor 1400 may include a system interface 1410, a graphics engine 1450, a security engine 1460, and a display engine 1470. However, note that the basic principles of the present disclosure are not limited to this implementation. For example, while the illustrated embodiment uses a system interface 1410 (such as a system graphics interface (SGI)), the techniques described herein are equally applicable to non-graphics interfaces.

[0199] The system interface 1410 compiles error data received from each source within the error counter 1430. The error sources can include units of the graphics processor 1400, such as the graphics engine 1450, the security engine 1460, the display engine 1470, and other non-graphics component(s) 1480. Each of the error sources can include error routing circuitry 1455, 1465, 1475, 1485 responsible for detecting errors and reporting the errors to the system interface 1410.

[0200] Errors can be classified into different types of errors. In the implementations herein, the types of errors can include correctable errors, uncorrectable fatal errors, or uncorrectable non-fatal errors. Correctable errors are those errors that can be corrected by hardware and without intervention provided by software for continued operation. For example, a storage device with an ECC SECDED (single error correction and double error detection) scheme can automatically correct single-bit errors. Errors on a storage device with parity protection and a replay / rewind mechanism are also considered correctable errors.

[0201] Uncorrectable errors are errors that cannot be corrected by hardware. These errors should be reported to software for error recovery. Structural errors (such as double-bit errors on an ECC-protected storage device), single-bit errors on a parity-protected storage device, and functional protocol errors (such as an inability to complete a write operation) are examples of uncorrectable errors. Recovery from such errors depends on the impact of the error.

[0202] In one implementation, the error routing circuitry 1455, 1465, 1475, 1485 of each distributed component of the graphics processor 1400 can use a standard error log format for a specific type of error (e.g., an uncorrectable error log format or a correctable error log format) to report errors. The error routing circuitry 1455, 1465, 1475, 1485 can then transmit the compiled error data to the system interface 1410 via an on-chip fabric (e.g., such as an IO system fabric (IOSF)).

[0203] The system interface 1410 can include an error aggregator 1420, which includes hardware circuitry for tracking and recording the reported errors from the distributed units of the processor 1400. The error aggregator can include an error counter 1430 and an error log 1440 for tracking and recording errors and their corresponding metadata for reporting to higher-level system components (e.g., debug and / or telemetry software) to determine the cause of the errors for, e.g., silicon health assessment in production.

[0204] Error aggregator 1420 can receive error report messages from one of the distributed units of the processor 1400. Error aggregator 1420 can identify the error type from the error report message and provide an interface to error counter 1430 to update the count for the received error type (e.g., count for correctable errors, count for uncorrectable errors, etc.).

[0205] Error aggregator 1420 can also route the error report message to the corresponding error log 1440. In some implementations, different error logs 1440 can be maintained for each type of error. For example, error log 1440 maintained by error counter 1430 can include an uncorrectable centralized error queue, an informational (info) centralized error queue, and a correctable centralized error queue. In some implementations, a single centralized error log 1440 is maintained for all errors.

[0206] Error log 1440 can be a fine-grained error log, which can be used to determine the cause of the error, for debugging, and for silicon health assessment in telemetry and production. In some embodiments, error log 1440 is configured to be sticky across thermal resets (i.e., retain their values across resets). Since some uncorrectable errors cause a thermal reset of the graphics processor 1400, the error information corresponding to such uncorrectable errors would be lost if it were not sticky.

[0207] Additionally, in some implementations, system interface 1410 can include error interrupt circuitry (not shown) for managing interrupts related to errors received at system interface 1410. In this example, a message signaled interrupt (MSI) can be generated to convey the error. In one implementation, once the count value in error counter 1430 exceeds a determined threshold, an interrupt (such as an MSI) is sent to a higher-level system component (such as a graphics driver). The MSI can invoke the graphics driver, which can access error counter 1430 and error log 1440 to identify the errors that should be parsed and read.

[0208] In one implementation, the processor 1400 implements correctable error address filtering to filter out correctable errors at the correctable error source using address filtering. The implementation provides correctable error address filtering circuitry 1457, 1467, 1477, 1487 in each of the units 1450, 1460, 1470, 1480 to detect errors in the processor 1400. The correctable address filtering circuitry 1457, 1467, 1477, 1487 of each unit 1450, 1460, 1470, 1480 is configured to control when the unit 1450, 1460, 1470, 1480 that detects a correctable error should report the correctable error to the system interface 1410 that aggregates error reports using the error counter 1430. Further details of the correctable error address filtering are described below with reference to Figure 15 Further details of the correctable error address filtering are described.

[0209] Figure 15 FIG. is a block diagram showing a detailed view of a processor 1500 that implements correctable error address filtering according to an implementation described herein. In one implementation, the processor 1500 is the same as the processor 1400 described with reference to Figure 14 In one implementation, the processor 1500 is a graphics processor 1500. The processor 1500 may include multiple error sources, including error source 0 1510-0 to error source N 1510-N (collectively referred to herein as error source 1510). The error source 1510 may be the same as Figure 14 the units 1450, 1460, 1470, 1480. In some implementations, the error source 1510 may include, but is not limited to, a general register file (GRF), caches such as a local shared cache (LSC) and an L2 cache, SRAM, a TLB, etc.

[0210] Each error source 1510 may include error routing circuitry 1520, which may be the same as Figure 14 the error routing circuitry 1455, 1465, 1475, 1485. The error routing circuitry 1520 may report the detected error, which includes both correctable and uncorrectable errors, to the error aggregator circuitry 1530. In one implementation, the error aggregator circuitry 1530 is the same as Figure 14 the error aggregator 1420. The error aggregator circuitry 1530 may include a correctable error counter 1532 and an error log 1534, which may be the same as Figure 14 the error counter 1430 and the error log 1440, respectively.

[0211] In one implementation, the error routing circuitry 1520 can implement correctable error address filtering. The error routing circuitry 1520 can include: a storage array with ECC 1522 for receiving and processing detected errors; and a correctable error address filter 1524 for filtering correctable errors based on the address corresponding to the error (e.g., the address being read or written).

[0212] In one implementation, the storage array with ECC 1522 can classify an error as a correctable error or an uncorrectable error. The uncorrectable errors are reported by the error routing circuitry to the error aggregator circuitry 1530 outside. The correctable errors are passed to the correctable error address filter 1524.

[0213] The correctable error address filter 1524 can track the last “N” correctable errors detected by the error source 1510, including the addresses associated with each detected correctable error. In the implementations herein, when a new correctable error is detected, the address of the correctable error is compared with the entries maintained by the correctable error address filtering circuitry that tracks the last “N” correctable errors. In one implementation, the correctable error address filtering circuitry can utilize a data structure such as a table to track the last “N” correctable errors. If an address match is detected, the correctable error is not reported to the error aggregator circuitry 1530. In one implementation, when there is an address match, the correctable error can be masked.

[0214] If an address match is not detected, the correctable error is reported to the error aggregator circuitry 1530. In one implementation, the oldest entry of the correctable error address filter 1524 is replaced with the new address of the correctable error.

[0215] In the case where a correctable error is associated with a write to an address and a match is found by the correctable error address circuitry, the corresponding entry in the correctable error address filtering circuitry is cleared from the correctable error address filter 1524 (e.g., from the table maintained by the correctable error address filter 1524). This is because a soft error (correctable error) that causes a bit flip will be corrected when writing to that bit.

[0216] In some implementations, the number of correctable errors to be tracked by the correctable error address filter is configurable. An example factor to consider when selecting the number of correctable errors to be tracked can include the physical interleaving of the codewords protected by ECC. Another example factor to consider when selecting the number of correctable errors to be tracked includes process information about the highest likelihood patterns for perturbations (e.g., 2×2, 3×3, 2×1, etc.).

[0217] In one implementation, the number of correctable errors to be tracked can be two errors. Configuring the number of correctable errors to be tracked as two errors can cover the case of soft errors that cause 2-bit perturbations in adjacent rows (i.e., single-bit perturbations in each row with its own ECC calculation), resulting in two addresses, each with a correctable error. In addition, the probability of multi-bit perturbations exceeding two bits is very low (e.g., <1%).

[0218] In some implementations, control bits can be provided to enable and disable the correctable error address filtering feature in the processor 1500.

[0219] Figure 16 is a flowchart illustrating an embodiment of a method 1600 for handling correctable error address filtering in a processing architecture. The method 1600 can be executed by processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof. For the sake of simplicity and clarity of presentation, the processes of the method 1600 are illustrated in a linear order; however, it is contemplated that any number of them can be executed in parallel, asynchronously, or in a different order. Further, for simplicity, clarity, and ease of understanding, many of the components and processes described with reference to Figures 1 - 15 the components and processes described may not be repeated or discussed below. In one implementation, a processing unit (such as, Figure 14 the processor 1400) can execute the method 1600.

[0220] The method 1600 begins at processing block 1610, where the processor can detect an error in a component of the distributed error reporting hierarchy of the processing architecture. Then, at block 1620, the processor can classify the error with an error type. In one implementation, the error type includes one of an uncorrectable error or a correctable error.

[0221] Subsequently, at block 1630, the processor can determine whether the address of the error matches any of the N entries maintained by the correctable error address filtering circuitry in response to an error classified with a correctable error type. Then, at decision block 1640, the processor can determine whether the address matches an entry maintained in the correctable error address filtering circuitry.

[0222] If so, method 1600 proceeds to block 1650, where the processor may mask an error in response to an error associated with a read of an address. In one implementation, masking the error causes the error not to be reported to the error aggregation hardware circuitry of the system interface. On the other hand, if at decision block 1640 the address corresponding to the error does not match any entry maintained in the correctable error address filter circuitry, then 1600 proceeds to block 1660, where the processor may report the error to the error aggregation hardware circuitry of the system interface.

[0223] Figure 17 FIG. is a flowchart of an embodiment of method 1700 for clearing entries in a correctable error address filter in a processing architecture. Method 1700 may be executed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof. For purposes of presentation simplicity and clarity, the processes of method 1700 are illustrated in a linear sequence; however, any number of them are contemplated to be executed in parallel, asynchronously, or in a different order. Further, for simplicity, clarity, and ease of understanding, many of the components and processes described with reference to Figures 1 - 16 may not be repeated or discussed hereinafter. In one implementation, a processing unit (such as, Figure 14 processor 1400) may execute method 1700.

[0224] Method 1700 begins at processing block 1710, where the processor may detect an error in a component of the distributed error reporting hierarchy of the processing architecture. Then, at block 1720, the processor may determine that the error is a correctable error.

[0225] Subsequently, at block 1730, the processor may determine that the address of the correctable error matches an entry of the N entries maintained by the correctable error address filter circuitry. Then, at block 1740, the processor may determine that a write to the address is associated with the correctable error. Finally, at block 1750, the processor may clear the entry from the correctable error address filter circuitry.

[0226] The following examples relate to further embodiments. Example 1 is an apparatus for facilitating correctable error address filtering in a processing architecture. The apparatus of Example 1 includes a processor that includes error routing hardware circuitry for: receiving data associated with an error detected in a source hardware component that hosts the error routing hardware circuitry; determining, based on the data, that the error is classified as a correctable error; comparing the address of the correctable error with entries maintained by a correctable error address filtering circuitry of the error routing hardware circuitry; masking the correctable error to prevent reporting of the correctable error to an error aggregation hardware circuitry of the processor in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry; and reporting the correctable error to the error aggregation hardware circuitry in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry.

[0227] In Example 2, the subject matter of Example 1 may optionally include: wherein the error routing hardware circuitry is further for: clearing an entry in the entries of the correctable error address filtering circuitry and adding data associated with the correctable error to the cleared entry in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry. In Example 3, the subject matter of any one of Examples 1-2 may optionally include: wherein, in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry and in response to the correctable error being associated with a write to the address, the error routing hardware circuitry is further for clearing the entry of the correctable error address filtering circuitry that matches the address.

[0228] In Example 4, the subject matter of any one of Examples 1-3 may optionally include: wherein the number of entries of the correctable error address filtering circuitry is configurable. In Example 5, the subject matter of any one of Examples 1-4 may optionally include: wherein the number of entries is two. In Example 6, the subject matter of any one of Examples 1-2 may optionally include: wherein uncorrectable errors detected in the source hardware component bypass the correctable error address filtering circuitry and are reported to the error aggregation hardware circuitry.

[0229] In Example 7, the subject matter of any one of Examples 1-6 may optionally include: wherein the processor includes a control bit for performing at least one of the following: enabling or disabling a correctable error address filtering circuitry. In Example 8, the subject matter of any one of Examples 1-7 may optionally include: wherein the processor includes a graphics processing unit (GPU). In Example 9, the subject matter of any one of Examples 1-8 may optionally include: wherein the device is at least one of a single instruction multiple data (SIMD) machine or a single instruction multiple thread (SIMT) machine.

[0230] Example 10 is a method for facilitating correctable error address filtering in a processing architecture. The method as in Example 10 may include: receiving, by an error routing circuitry of a source hardware component of a processing device, data associated with an error detected in the source hardware component, the source hardware component hosting the error routing circuitry; determining, by the error routing circuitry, based on the data, that the error is classified as a correctable error; comparing an address of the correctable error with entries maintained by a correctable error address filtering circuitry of the error routing circuitry; in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry, masking the correctable error to prevent reporting of the correctable error to an error aggregation hardware circuitry of the processing device; and in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, reporting the correctable error to the error aggregation hardware circuitry.

[0231] In Example 11, the subject matter as in Example 10 may optionally further include: in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, clearing an entry in the entries of the correctable error address filtering circuitry and adding data associated with the correctable error to the cleared entry. In Example 12, the subject matter as in Examples 10-11 may optionally further include: in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry and in response to the correctable error being associated with a write to the address, clearing the entry of the correctable error address filtering circuitry that matches the address.

[0232] In Example 13, the subject matter of Examples 10 - 12 may optionally include: wherein the number of entries in the correctable error address filtering circuitry is configurable. In Example 14, the subject matter of Examples 10 - 13 may optionally include: wherein uncorrectable errors detected in the source hardware component bypass the correctable error address filtering circuitry and are reported to the error aggregation hardware circuitry. In Example 15, the subject matter of Examples 10 - 14 may optionally include: wherein the processing device includes a control bit for performing at least one of the following: enabling or disabling the correctable error address filtering circuitry.

[0233] Example 16 is a non - transitory computer - readable storage medium for facilitating correctable error address filtering in a processing architecture. The non - transitory computer - readable storage medium of Example 16 has instructions stored thereon that, when executed by one or more processors, cause the processors to: receive, by an error routing circuitry of a source hardware component of the one or more processors, data associated with an error detected in the source hardware component, the source hardware component hosting the error routing circuitry; determine, by the error routing circuitry based on the data, that the error is classified as a correctable error; compare the address of the correctable error with entries maintained by a correctable error address filtering circuitry of the error routing circuitry; in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry, mask the correctable error to prevent reporting of the correctable error to an error aggregation hardware circuitry of the one or more processors; and in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, report the correctable error to the error aggregation hardware circuitry.

[0234] In Example 17, the subject matter of Example 16 may optionally include, wherein the one or more processors are further configured to: in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, clear an entry in the entries of the correctable error address filtering circuitry and add the data associated with the correctable error to the cleared entry. In Example 18, the subject matter of Examples 16 - 17 may optionally include, wherein the one or more processors are further configured to: in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry and in response to the correctable error being associated with a write to the address, clear the entry of the correctable error address filtering circuitry that matches the address.

[0235] In Example 19, the subject matter of Examples 16 - 18 may optionally include: wherein the number of entries in the correctable error address filtering circuitry is configurable. In Example 20, the subject matter of Examples 16 - 19 may optionally include: wherein an uncorrectable error detected in a source hardware component bypasses the correctable error address filtering circuitry and is reported to an error aggregation hardware circuitry.

[0236] Example 21 is a system for facilitating correctable error address filtering in a processing architecture. The system of Example 21 may optionally include: a memory; and a processor communicatively coupled to the memory, the processor including error routing hardware circuitry, wherein the error routing hardware circuitry is configured to: receive data associated with an error detected in a source hardware component that hosts the error routing hardware circuitry; determine, based on the data, that the error is classified as a correctable error; compare an address of the correctable error with entries maintained by a correctable error address filtering circuitry of the error routing hardware circuitry; in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry, mask the correctable error to prevent reporting of the correctable error to an error aggregation hardware circuitry of a processing device; and in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, report the correctable error to the error aggregation hardware circuitry.

[0237] In Example 22, the subject matter of Example 21 may optionally include, wherein the error routing hardware circuitry is further configured to: in response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, clear an entry in the entries of the correctable error address filtering circuitry and add data associated with the correctable error to the cleared entry. In Example 23, the subject matter of any one of Examples 21 - 22 may optionally include, wherein in response to the address of the correctable error matching an entry of the correctable error address filtering circuitry and in response to the correctable error being associated with a write to the address, the error routing hardware circuitry is further configured to clear the entry of the correctable error address filtering circuitry that matches the address.

[0238] In Example 24, the subject matter of any one of Examples 21 - 23 may optionally include: wherein the number of entries in the correctable error address filtering circuitry is configurable. In Example 25, the subject matter of any one of Examples 21 - 24 may optionally include: wherein the number of entries is two. In Example 26, the subject matter of any one of Examples 21 - 25 may optionally include: wherein an uncorrectable error detected in a source hardware component bypasses the correctable error address filtering circuitry and is reported to an error aggregation hardware circuitry.

[0239] In Example 27, the subject matter of any of Examples 21-26 may optionally include, wherein the processor includes a control bit for performing at least one of the following: enabling or disabling a correctable error address filtering circuitry. In Example 28, the subject matter of any of Examples 21-27 may optionally include, wherein the processor includes a graphics processing unit (GPU). In Example 29, the subject matter of any of Examples 21-28 may optionally include, wherein the device is at least one of a single instruction multiple data (SIMD) machine or a single instruction multiple thread (SIMT) machine.

[0240] Example 30 is a device for facilitating correctable error address filtering in a processing architecture, including: means for receiving, using an error routing circuitry of a source hardware component of a processing device, data associated with an error detected in the source hardware component, the source hardware component hosting the error routing circuitry; means for using the error routing circuitry to determine, based on the data, that the error is classified as a correctable error; means for comparing an address of the correctable error with entries maintained by a correctable error address filtering circuitry of the error routing circuitry; means for masking the correctable error to prevent reporting of the correctable error to an error aggregation hardware circuitry of the processor in response to a match of the address of the correctable error with an entry of the correctable error address filtering circuitry; and means for reporting the correctable error to the error aggregation hardware circuitry in response to a mismatch of the address of the correctable error with an entry of the correctable error address filtering circuitry. In Example 31, the subject matter of Example 30 may optionally include a device further configured to perform a method of any of Examples 11 to 15.

[0241] Example 32 is at least one machine-readable medium including a plurality of instructions that, when executed on a computing device, cause the computing device to perform a method of any of Examples 10-15. Example 33 is a device for facilitating correctable error address filtering in a processing architecture, the device being configured to perform a method of any of Examples 10 to 15. Example 34 is a device for facilitating conversion operations and special value use cases supporting an 8-bit floating-point format in a graphics architecture, including means for performing a method of any of Examples 10 to 15. Details in the examples may be used anywhere in one or more embodiments.

[0242] The foregoing specification and drawings are to be regarded in an illustrative rather than a restrictive sense. Those skilled in the art will understand that various modifications and changes may be made to the embodiments described herein without departing from the broader spirit and scope of the features as set forth in the appended claims.

Claims

1. A device comprising: A processor comprising an error routing hardware circuitry configured to: receiving data associated with an error detected in a source hardware component hosting the error routing hardware circuitry; determining, based on the data, that the error is classified as a correctable error; comparing the address of the correctable error to an entry maintained by correctable error address filtering circuitry of the error routing hardware circuitry; responsive to the address of the correctable error matching an entry of the correctable error address filtering circuitry, masking the correctable error to prevent reporting of the correctable error to error aggregation hardware circuitry of the processor; as well as In response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, reporting the correctable error to the error aggregation hardware circuitry.

2. The device according to claim 1, wherein: The error routing hardware circuit system is further used to: in response to the address of the correctable error not matching the entry of the correctable error address filtering circuit system, clear one of the entries of the correctable error address filtering circuit system, and add the data associated with the correctable error to the cleared entry of the entries.

3. The device according to claim 1, wherein: In response to the address of the correctable error matching an entry of the correctable error address filtering circuitry, and in response to the correctable error being associated with a write to the address, the error routing hardware circuitry is further operable to clear the entry of the correctable error address filtering circuitry that matches the address.

4. The device according to claim 1, wherein: The number of entries of the correctable error address filtering circuitry is configurable.

5. The device according to claim 4, wherein: The number of entries is two.

6. The device according to claim 1, wherein: Uncorrectable errors detected in the source hardware component bypass the correctable error address filtering circuitry and are reported to the error aggregation hardware circuitry.

7. The device of claim 1, wherein: The processor includes a control bit for at least one of enabling or disabling the correctable error address filtering circuitry.

8. The device of claim 1, wherein: The processor includes a graphics processing unit GPU.

9. The device of claim 8, wherein: The device is at least one of a Single Instruction Multiple Data SIMD machine or a Single Instruction Multiple Thread SIMT machine.

10. A method comprising: receiving, by error routing circuitry of a source hardware component of a processing device, data associated with an error detected in the source hardware component, the source hardware component hosting the error routing circuitry; determining, by the error routing circuitry based on the data, that the error is classified as a correctable error; comparing the address of the correctable error to an entry maintained by correctable error address filtering circuitry of the error routing circuitry; responsive to the address of the correctable error matching an entry of the correctable error address filtering circuitry, masking the correctable error to prevent reporting of the correctable error to error aggregation hardware circuitry of the processing device; as well as In response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, reporting the correctable error to the error aggregation hardware circuitry.

11. The method of claim 10, further comprising: In response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, one of the entries of the correctable error address filtering circuitry is cleared and the data associated with the correctable error is added to the cleared one of the entries.

12. The method of claim 10, further comprising: In response to the address of the correctable error matching an entry of the correctable error address filtering circuitry, and in response to the correctable error being associated with a write to the address, clearing the entry of the correctable error address filtering circuitry matching the address.

13. The method of claim 10, wherein: The number of entries of the correctable error address filtering circuitry is configurable.

14. The method of claim 10, wherein: Uncorrectable errors detected in the source hardware component bypass the correctable error address filtering circuitry and are reported to the error aggregation hardware circuitry.

15. The method of claim 10, wherein: The processing device includes a control bit for at least one of enabling or disabling the correctable error address filtering circuitry.

16. A non-transitory computer readable medium having instructions stored thereon, which instructions, when executed by one or more processors, cause the one or more processors to: receiving, by error routing circuitry of a source hardware component of the one or more processors, data associated with an error detected in the source hardware component, the source hardware component hosting the error routing circuitry; determining, by the error routing circuitry based on the data, that the error is classified as a correctable error; comparing the address of the correctable error to an entry maintained by correctable error address filtering circuitry of the error routing circuitry; responsive to the address of the correctable error matching an entry of the correctable error address filtering circuitry, masking the correctable error to prevent reporting of the correctable error to error aggregation hardware circuitry of the one or more processors; as well as In response to the address of the correctable error not matching an entry of the correctable error address filtering circuitry, reporting the correctable error to the error aggregation hardware circuitry.

17. The non-transitory computer readable medium of claim 16, wherein: The one or more processors are further configured to: in response to the address of the correctable error not matching an entry of the correctable error address filtering circuit system, clear an entry in the entries of the correctable error address filtering circuit system, and add the data associated with the correctable error to the cleared entry in the entries.

18. The non-transitory computer readable medium of claim 16, wherein: The one or more processors are further configured to: in response to the address of the correctable error matching an entry of the correctable error address filtering circuit system, and in response to the correctable error being associated with a write to the address, clear the entry of the correctable error address filtering circuit system that matches the address.

19. The non-transitory computer readable medium of claim 16, wherein: The number of entries of the correctable error address filtering circuitry is configurable.

20. The non-transitory computer readable medium of claim 16, wherein: Uncorrectable errors detected in the source hardware component bypass the correctable error address filtering circuitry and are reported to the error aggregation hardware circuitry.

21. At least one machine-readable medium comprising a plurality of instructions which, in response to being executed on a computing device, cause the computing device to perform the method of any one of claims 10-15.