Apparatus and method for managing data bias in a graphics processing architecture
Patent Information
- Application Number
- CN202311189470.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-04-07
- Filing Date
- 2018-04-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2038-04-08
Smart Images

Figure CN117057975B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on April 8, 2018, with a priority date of April 7, 2017, application number 201810306269.0, entitled "Apparatus and method for managing data bias in a graphics processing architecture". Background Technology Technical Field
[0003] This invention generally relates to the field of computer processors. More specifically, this invention relates to apparatus and methods for managing data bias in a graphics processing architecture.
[0004] Description of related technologies
[0005] Recent advancements have been made in graphics processing unit (GPU) virtualization. Virtualized graphics processing environments are used for applications such as media clouds, remote workstations / desks, interchangeable virtual instruments (IVIs), rich client virtualization, and more. Some architectures perform full GPU virtualization through capture and emulation to emulate a fully functional virtual GPU (vGPU) while delivering near-native performance by transferring performance-critical graphics memory resources.
[0006] As GPUs become increasingly important in supporting 3D, media, and GPGPU workloads in servers, GPU virtualization is becoming more and more prevalent. How to virtualize GPU memory access from virtual machines (VMs) is one of the key design factors. GPUs have their own graphics memory: dedicated video memory or shared system memory. When system memory is used for graphics, the guest physical address (GPA) needs to be translated into a host physical address (HPA) before being accessed by hardware.
[0007] There are several approaches to performing translations for GPUs. Some implementations perform translations via hardware support, but this can only deliver the GPU to a single VM. Another solution is a software approach that builds a shadow structure for the translation. For example, shadow page tables are implemented using certain architectures, such as those in the complete GPU virtualization solution mentioned above, which can support multiple VMs sharing a physical GPU.
[0008] In some implementations, guest / VM memory pages are backed by host memory pages. The virtual machine monitor (VMM) (sometimes called the "hypervisor") uses, for example, extended page tables (EPTs) to map guest physical addresses (PAs) to host PAs. Various memory-sharing techniques can be used, such as kernel page merging (KSM).
[0009] KSM merges pages from multiple VMs with identical content into a single write-protected page. That is, if a memory page in VM1 (mapped from guest PA1 to host PA1) has the same content as another memory page in VM2 (mapped from guest PA2 to host PA2), guest memory can be supported using only one host page (such as HPA_SH). In other words, both guest PA1 of VM1 and PA2 of VM2 are mapped to a write-protected HPA_SH. This saves memory used by the system and is particularly useful for guest read-only memory pages (such as code pages and zero pages). Using KSM, once a VM modifies the page content, copy-on-write (COW) technology can be used to remove shared memory.
[0010] The intermediary facilitates device performance and sharing within a virtualized system, where a single physical GPU is presented as multiple virtual GPUs to multiple clients with direct DMA, while privileged resources accessed by the clients remain captured and emulated. In some implementations, each client can run a native GPU driver, and device DMA directly accesses memory without hypervisor intervention. Attached Figure Description
[0011] The invention can be better understood from the following detailed description in conjunction with the accompanying drawings, wherein:
[0012] Figure 1 This is a block diagram of an embodiment of a computer system having a processor having one or more processor cores and a graphics processor;
[0013] Figure 2 This is a block diagram of one embodiment of a processor, the processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor;
[0014] Figure 3 This is a block diagram of one embodiment of a graphics processor, which may be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores;
[0015] Figure 4 This is a block diagram of an embodiment of a graphics processing engine for a graphics processor;
[0016] Figure 5 This is a block diagram of another embodiment of a graphics processor;
[0017] Figure 6 It is a block diagram of thread execution logic that includes an array of process elements;
[0018] Figure 7 The instruction format of the graphics processor execution unit according to an embodiment is shown;
[0019] Figure 8 This is a block diagram of another embodiment of a graphics processor, which includes a graphics pipeline, a media pipeline, a display engine, thread execution logic, and a rendering output pipeline.
[0020] Figure 9A This is a block diagram illustrating the graphics processor command format according to an embodiment;
[0021] Figure 9B This is a block diagram illustrating a sequence of graphics processor commands according to an embodiment;
[0022] Figure 10 An exemplary graphical software architecture of a data processing system according to an embodiment is shown;
[0023] Figure 11 An exemplary IP core development system according to an embodiment is shown;
[0024] Figure 12 An exemplary system-on-chip integrated circuit that can be fabricated using one or more IP cores according to embodiments is shown;
[0025] Figure 13 An exemplary graphics processor is shown that can be fabricated using one or more IP cores as a system-on-a-chip integrated circuit;
[0026] Figure 14 An additional exemplary graphics processor is shown that can be fabricated using one or more IP cores of a system-on-a-chip integrated circuit;
[0027] Figure 15 An exemplary graphics processing system is shown;
[0028] Figure 16 An exemplary architecture for full graphics virtualization is demonstrated;
[0029] Figure 17 An exemplary virtualized graphics processing architecture, including a virtual graphics processing unit (vGPU), is demonstrated.
[0030] Figure 18 An example of a virtualization architecture with IOMMU is shown;
[0031] Figure 19 An embodiment is shown in which graphics processing is performed on a server;
[0032] Figure 20 This illustrates one embodiment in which a Super Row Ownership Table (SLOT) is maintained to track GPU and host bias for the super row;
[0033] Figure 21-22 The implementation of super-row transactions for host-biased and GPU-biased scenarios is shown.
[0034] Figure 23 An embodiment is shown that includes a SLOT cache and the SLOT can be stored in GPU memory or host memory;
[0035] Figure 24 A method according to an embodiment of the present invention is shown;
[0036] Figure 25 This illustrates one embodiment of a symmetrical client-side implementation;
[0037] Figure 26 This illustrates one embodiment of an asymmetric client-side implementation;
[0038] Figure 27 This is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described in this application;
[0039] Figures 28A-28D A parallel processor component according to an embodiment is shown;
[0040] Figures 29A-29B This is a block diagram of a graphics multiprocessor according to an embodiment;
[0041] Figures 30A-30F An exemplary architecture is shown in which multiple GPUs are communicatively coupled to multiple multi-core processors; and
[0042] Figure 31 A graphics processor pipeline according to an embodiment is shown. Detailed Implementation
[0043] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the embodiments of the invention described below. However, it will be apparent to those skilled in the art that embodiments of the invention can be practiced without some of these specific details. In other instances, well-known structures and apparatuses are illustrated in block diagram form to avoid obscuring the fundamental principles of embodiments of the invention.
[0044] Exemplary graphics processor architecture and data types
[0045] System Overview
[0046] Figure 1This is a block diagram of a processing system 100 according to an embodiment. In various embodiments, the processing system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single-processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, the processing system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for mobile, handheld, or embedded devices.
[0047] Embodiments of the processing system 100 may include or incorporate a server-based game platform, game console, including a game and media console, mobile game console, handheld game console, or online game console. In some embodiments, the processing system 100 is a mobile phone, smartphone, tablet computing device, or mobile internet device. The processing system 100 may also include a wearable device (such as a smartwatch, smart glasses, augmented reality, or virtual reality device), coupled to or integrated into the wearable device. In some embodiments, the processing system 100 is a television or set-top box device having one or more processors 102 and a graphical interface generated by one or more graphics processors 108.
[0048] In some embodiments, each of the one or more processors 102 includes one or more processor cores 107 for processing instructions that, when executed, perform operations on the system and user software. In some embodiments, each of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). Multiple processor cores 107 may each process different instruction sets 109, which may include instructions for facilitating emulation of other instruction sets. Processor cores 107 may also include other processing means, such as digital signal processors (DSPs).
[0049] In some embodiments, processor 102 includes cache memory 104. Depending on the architecture, processor 102 may have a single internal cache or multiple levels of internal caches. In some embodiments, cache memory is shared among components of processor 102. In some embodiments, processor 102 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which can be shared among processor core 107 using known cache coherence techniques. Additionally, register file 106 is included in processor 102, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while others may be specific to the design of processor 102.
[0050] In some embodiments, processor 102 is coupled to processor bus 110, which is used to transmit communication signals, such as address, data, or control signals, between processor 102 and other components within processing system 100. In one embodiment, processing system 100 uses an exemplary 'central' system architecture including a memory controller central 116 and an input / output (I / O) controller central 130. Memory controller central 116 facilitates communication between memory devices and other components of processing system 100, while I / O controller central (ICH) 130 provides connectivity to I / O devices via a native I / O bus. In one embodiment, the logic of memory controller central 116 is integrated within the processor.
[0051] Memory device 120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or some other memory device with suitable performance for use as processing memory. In one embodiment, memory device 120 may operate as system memory of processing system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. Memory controller hub 116 is also coupled to an optional external graphics processor 112, which may communicate with one or more graphics processors 108 in processor 102 to perform graphics and media operations.
[0052] In some embodiments, ICH 130 connects peripheral components to memory device 120 and processor 102 via a high-speed I / O bus. I / O peripheral components include, but are not limited to, an audio controller 146, a firmware interface 128, a wireless transceiver 126 (e.g., Wi-Fi, Bluetooth), a data storage device 124 (e.g., a hard disk drive, flash memory, etc.), and a conventional I / O controller 140 for coupling conventional (e.g., Personal System 2 (PS / 2)) devices to the system. One or more Universal Serial Bus (USB) controllers 142 connect multiple input devices, such as a keyboard and mouse combination 144. A network controller 134 may also be coupled to ICH 130. In some embodiments, a high-performance network controller (not shown) is coupled to processor bus 110. It should be understood that the illustrated processing system 100 is exemplary and not limiting, as other types of data processing systems configured differently may also be used. For example, the I / O controller hub 130 may be integrated within one or more processors 102, or the memory controller hub 116 and the I / O controller hub 130 may be integrated within a discrete external graphics processor (such as external graphics processor 112).
[0053] Figure 2 This is a block diagram of an embodiment of processor 200, which has one or more processor cores 202A to 202N, an integrated memory controller 214, and an integrated graphics processor 208. Figure 2 Those elements having the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. Processor 200 may include, and include, additional cores 202N, indicated by dashed boxes. Each processor core 202A to 202N includes one or more internal cache units 204A to 204N. In some embodiments, each processor core may also access one or more shared cache units 206.
[0054] Internal cache units 204A to 204N and shared cache unit 206 represent the cache memory hierarchy within processor 200. The cache memory hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest-level cache is classified as LLC before external memory. In some embodiments, cache coherence logic maintains coherence between each cache unit 206 and 204A to 204N.
[0055] In some embodiments, the processor 200 may further include a group of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a group of peripheral buses, such as one or more peripheral component interconnect buses (e.g., PCI, PCI Express). The system agent core 210 provides management functions for each processor unit. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).
[0056] In some embodiments, one or more of the processor cores 202A to 202N include support for simultaneous multithreading. In this embodiment, the system agent core 210 includes components for coordinating and operating the cores 202A to 202N during multithreaded processing. Additionally, the system agent core 210 may also include a power control unit (PCU) including logic and components for regulating the power states of the processor cores 202A to 202N, and a graphics processor 208.
[0057] In some embodiments, processor 200 additionally includes a graphics processor 208 for performing graphics processing operations. In some embodiments, graphics processor 208 is coupled to a shared cache unit 206 and a system proxy core 210, the system proxy core including one or more integrated memory controllers 214. In some embodiments, display controller 211 is coupled to graphics processor 208 to drive graphics processor output to one or more coupled displays. In some embodiments, display controller 211 may be a separate module coupled to graphics processor via at least one interconnect, or it may be integrated within graphics processor 208 or system proxy core 210.
[0058] In some embodiments, ring-based interconnect units 212 are used to couple internal components of processor 200. However, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, including those well known in the art, may be used. In some embodiments, graphics processor 208 is coupled to ring interconnect 212 via I / O link 213.
[0059] Exemplary I / O link 213 represents at least one of several types of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 218 (such as eDRAM modules). In some embodiments, each of the processor cores 202A to 202N and the graphics processor 208 uses the embedded memory module 218 as a shared final-level cache.
[0060] In some embodiments, processor cores 202A to 202N are homogeneous cores executing the same instruction set architecture. In another embodiment, processor cores 202A to 202N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more of processor cores 202A to 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, processor cores 202A to 202N are homogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled to one or more power cores with lower power consumption. Additionally, processor 200 can be implemented on one or more chips or implemented as a SoC integrated circuit having, among other components, the components shown.
[0061] Figure 3 This is a block diagram of a graphics processing unit 300, which may be a discrete graphics processing unit or a graphics processing unit integrated with multiple processing cores. In some embodiments, the graphics processing unit communicates with memory via a mapped I / O interface to registers on the graphics processing unit and using commands placed in processor memory. In some embodiments, the graphics processing unit 300 includes a memory interface 314 for accessing memory. The memory interface 314 may be an interface to native memory, one or more internal caches, one or more shared external caches, and / or to system memory.
[0062] In some embodiments, the graphics processor 300 further includes a display controller 302 for driving display output data to a display device 320. The display controller 302 includes hardware for one or more overlapping planes of the display and a multi-layer video or user interface component. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding, decoding, or converting media codes to, from, or between one or more media encoding formats, including but not limited to: Moving Picture Experts Group (MPEG) (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC), and Society of Motion Picture and Television Engineers (SMPTE) 421M / VC-1, and Joint Picture Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG)).
[0063] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 to perform two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfer. However, in one embodiment, 2D graphics operations are performed using one or more components of a graphics processing engine (GPE) 310. In some embodiments, the GPE 310 is a computational engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0064] In some embodiments, GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering 3D images and scenes using processing functions acting on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed functional elements that perform various tasks within elements and / or generated execution threads of the 3D / media subsystem 315. While the 3D pipeline 312 can be used to perform media operations, embodiments of GPE 310 also include a media pipeline 316 specifically for performing media operations such as video post-processing and image enhancement.
[0065] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units to perform one or more specialized media operations, such as video decoding acceleration, video deinterfacing, and video encoding acceleration, in place of or on behalf of the video codec engine 306. In some embodiments, the media pipeline 316 further includes a thread generation unit to generate threads for execution on the 3D / media subsystem 315. The generated threads perform calculations on media operations on one or more graphics execution units included in the 3D / media subsystem 315.
[0066] In some embodiments, the 3D / media subsystem 315 includes logic for executing threads generated by the 3D pipeline 312 and the media pipeline 316. In one embodiment, the pipelines send thread execution requests to the 3D / media subsystem 315, the 3D / media subsystem including thread dispatch logic for arbitrating and dispatching requests to available thread execution resources. Execution resources include an array of graphics execution units for processing 3D and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory (including registers and addressable memory) for sharing data between threads and for storing output data.
[0067] Graphics processing engine
[0068] Figure 4This is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is... Figure 3 The image shows a version of GPE 310. Figure 4 Those elements having the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. For example, shown Figure 3 The 3D pipeline 312 and media pipeline 316 are included. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example, and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.
[0069] In some embodiments, GPE 410 is coupled to or includes command streamer 403, which provides a command stream to 3D pipeline 312 and / or media pipeline 316. In some embodiments, command streamer 403 is coupled to memory, which may be system memory, or one or more cache memories of internal cache memory and shared cache memory. In some embodiments, command streamer 403 receives commands from memory and sends these commands to 3D pipeline 312 and / or media pipeline 316. The commands are instructions obtained from a ring buffer storing instructions for 3D pipeline 312 and media pipeline 316. In one embodiment, the ring buffer may also include a batch command buffer storing multiple batches of commands. Commands for 3D pipeline 312 may also include references to data stored in memory, such as, but not limited to, vertex and geometry data for 3D pipeline 312 and / or image data and memory objects for media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process the commands by performing operations via logic within their respective pipelines or by dispatching one or more execution threads to the execution unit array 414.
[0070] In various embodiments, the 3D pipeline 312 can execute one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 414. The graphics core array 414 provides a unified block of execution resources. The multipurpose execution logic (e.g., execution units) within the graphics core array 414 includes support for various 3D API shader languages and can execute multiple concurrent threads associated with multiple shaders.
[0071] In some embodiments, the graphics core array 414 further includes execution logic for performing media functions such as video and / or image processing. In one embodiment, in addition to graphics processing operations, the execution unit also includes general-purpose logic programmable to perform parallel general-purpose computing operations. The general-purpose logic can be... Figure 1 (Multiple) processor cores 107 or Figure 2 The general logic within the cores 202A to 202N performs processing operations in parallel or in combination.
[0072] Output data generated by threads executing on the graphics core array 414 can be output to memory in a uniform return buffer (URB) 418. URB 418 can store data from multiple threads. In some embodiments, URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, URB 418 can also be used for synchronization between threads on the graphics core array and fixed-function logic within shared-function logic 420.
[0073] In some embodiments, the graphics core array 414 is scalable, such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance level of the GPE 410. In one embodiment, the execution resources are dynamically scalable, allowing them to be enabled or disabled as needed.
[0074] The graphics core array 414 is coupled to shared function logic 420, which includes multiple resources shared among the graphics cores in the graphics core array. Shared functions within the shared function logic 420 are hardware logic units that provide dedicated supplementary functions to the graphics core array 414. In various embodiments, the shared function logic 420 includes, but is not limited to, sampler 421, math 422, and inter-thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420. Shared functions are implemented when the requirement for a given dedicated function is insufficient to be included in the graphics core array 414. Instead, a single instance of the dedicated function is implemented as a separate entity within the shared function logic 420 and shared among execution resources within the graphics core array 414. The exact set of functions shared and included within the graphics core array 414 varies between embodiments.
[0075] Figure 5 This is a block diagram of another embodiment of the graphics processor 500. Figure 5 Those elements having the same reference numerals (or names) as those in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein.
[0076] In some embodiments, the graphics processor 500 includes a ring interconnect 502, a pipeline front-end 504, a media engine 537, and graphics cores 580A to 580N. In some embodiments, the ring interconnect 502 couples the graphics processor to other processing units, including other graphics processors or one or more general-purpose processor cores. In some embodiments, the graphics processor is one of a plurality of processors integrated within a multi-core processing system.
[0077] In some embodiments, the graphics processor 500 receives multiple batches of commands via a ring interconnect 502. Incoming commands are interpreted by a command streamer 503 in a pipeline front-end 504. In some embodiments, the graphics processor 500 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores(a plurality of) 580A to 580N. For 3D geometry processing commands, the command streamer 503 supplies commands to a geometry pipeline 536. For at least some media processing commands, the command streamer 503 supplies commands to a video front-end 534, which is coupled to a media engine 537. In some embodiments, the media engine 537 includes a video quality engine (VQE) 530 for video and image post-processing and a multi-format encoding / decoding (MFX) engine 533 for providing hardware-accelerated media data encoding and decoding. In some embodiments, the geometry pipeline 536 and the media engine 537 each generate execution threads for use with thread execution resources provided by at least one graphics core 580A.
[0078] In some embodiments, the graphics processor 500 includes Scalable Thread Execution Resource Characteristic (SRM) cores 580A to 580N (sometimes referred to as core shards), each SRM core having multiple sub-cores 550A to 550N and 560A to 560N (sometimes referred to as core sub-shards). In some embodiments, the graphics processor 500 may have any number of graphics cores 580A to 580N. In some embodiments, the graphics processor 500 includes a graphics core 580A, which has at least a first sub-core 550A and a second sub-core 560A. In other embodiments, the graphics processor is a low-power processor having a single sub-core (e.g., 550A). In some embodiments, the graphics processor 500 includes multiple graphics cores 580A to 580N, each graphics core including a set of first sub-cores 550A to 550N and a set of second sub-cores 560A to 560N. Each of the first set of sub-cores 550A to 550N includes at least a first set of execution units 552A to 552N and media / texture samplers 554A to 554N. Each of the second set of sub-cores 560A to 560N includes at least a second set of execution units 562A to 562N and samplers 564A to 564N. In some embodiments, each sub-core 550A to 550N and 560A to 560N shares a set of shared resources 570A to 570. In some embodiments, the shared resources include shared cache memory and pixel operation logic. Other shared resources may also be included in various embodiments of the graphics processor.
[0079] Execution unit
[0080] Figure 6 A thread execution logic 600 is shown, which includes an array of processing elements used in some embodiments of GPE. Figure 6 Those elements having the same reference numerals (or names) as those in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein.
[0081] In some embodiments, thread execution logic 600 includes a shader processor 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array including multiple execution units 608A to 608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 608A, 608B, 608C, 608D, up to 608N-1 and 608N) based on workload computational requirements. In one embodiment, the included components are interconnected via an interconnect structure linking each component within the components. In some embodiments, thread execution logic 600 includes one or more connections to memory (such as system memory or cache memory) via the instruction cache 606, the data port 614, the sampler 610, and one or more of the execution unit arrays 608A to 608N. In some embodiments, each execution unit (e.g., 608A) is an independent programmable general-purpose computing unit capable of executing multiple synchronous hardware threads while processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 608A to 608N is scalable to include any number of individual execution units.
[0082] In some embodiments, execution units 608A to 608N are primarily used to execute shader programs. Shader processor 602 can handle various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 604. In one embodiment, the thread dispatcher includes logic for arbitrating thread requests from the graphics and media pipeline and instantiating the requested thread on one or more execution units 608A to 608N. For example, a geometry pipeline (e.g., Figure 5 (536) can dispatch vertex processing, tessellation, or geometry processing threads to thread execution logic 600 ( Figure 6 The thread dispatcher 604 can also process runtime thread generation requests from the shader execution program. In some embodiments, the thread dispatcher 604 can also process runtime thread generation requests from the shader execution program.
[0083] In some embodiments, execution units 608A to 608N support instruction sets that include native support for many standard 3D graphics shader instructions, enabling minimal conversion to execute shader programs from graphics libraries (e.g., Direct3D and OpenGL). These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., computation and media shaders). Each of the execution units 608A to 608N is capable of executing multiple-issue single-instruction multiple-data (SIMD), and multithreaded operation enables an efficient execution environment in the face of high-latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. For pipelines with integer, single-precision floating-point and double-precision floating-point operations, SIMD branching capabilities, logical operations, transcendental operations, and other miscellaneous operations, execution is multiple-issue per clock cycle. While waiting for data from memory or a shared function, dependency logic within execution units 608A to 608N causes the waiting thread to sleep until the requested data has been returned. While the waiting thread is sleeping, hardware resources may be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution unit may perform operations on a pixel shader, a fragment shader, or another type of shader program that includes different vertex shaders.
[0084] Each execution unit in the 608A to 608N operates on an array of data elements. The number of data elements is the "execution size," or the number of instruction channels. An execution channel is a logical unit that performs data element access, masking, and flow control within instructions. The number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating-point units (FPUs) for a particular graphics processor. In some embodiments, the execution units 608A to 608N support both integer and floating-point data types.
[0085] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as compressed data types, and the execution unit will process these elements based on their data size. For example, when operating on a 256-bit wide vector, the 256-bit vector is stored in registers, and the execution unit operates on the vector as four individual 64-bit compressed data elements (four times the word length (QW) size), eight individual 32-bit compressed data elements (double the word length (DW) size), sixteen individual 16-bit compressed data elements (word length (W) size), or thirty-two individual 8-bit data elements (byte (B) size). However, different vector widths and register sizes are possible.
[0086] One or more internal instruction caches (e.g., 606) are included in the thread execution logic 600 to cache thread instructions of the execution unit. In some embodiments, one or more data caches (e.g., 612) are included for caching thread data during thread execution. In some embodiments, sampler 610 is included for providing texture sampling for 3D operations and media sampling for media operations. In some embodiments, sampler 610 includes dedicated texture or media sampling functions to process texture or media data during the sampling process before providing sampled data to the execution unit.
[0087] During execution, the graphics and media pipeline sends thread initiation requests to thread execution logic 600 via thread generation and dispatch logic. Once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 602 is invoked to further compute output information and write the results to output surfaces (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes values for vertex attributes interpolated across the rasterized object. In some embodiments, the pixel processor logic within shader processor 602 then executes a pixel or fragment shader program provided by an application programming interface (API). To execute the shader program, shader processor 602 dispatches threads to execution units (e.g., 608A) via thread dispatcher 604. In some embodiments, shader processor 602 uses texture sampling logic in sampler 610 to access texture data in a texture map stored in memory. Arithmetic operations are performed on the texture data and the input geometry data to calculate the pixel color data of each geometric fragment, or to discard one or more pixels without further processing.
[0088] In some embodiments, data port 614 provides a memory access mechanism for thread execution logic 600 to output processed data to memory for processing on the graphics processor output pipeline. In some embodiments, data port 614 includes or is coupled to one or more cache memories (e.g., data cache 612) to cache data via the data port for memory access.
[0089] Figure 7This is a block diagram illustrating a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set having multiple instruction formats. Solid lines represent components typically included in the execution unit instructions, while dashed lines represent optional components or components included only in subsets of the instructions. In some embodiments, the instruction format 700 described and illustrated is a macro instruction, as the macro instruction is an instruction supplied to the execution unit, as opposed to the micro-operations generated from instruction decoding (once the instruction is processed).
[0090] In some embodiments, the graphics processor execution unit natively supports instructions using the 128-bit instruction format 710. A 64-bit compressed instruction format 730 can be used for some instructions based on the selected instruction, instruction options, and number of operands. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted to the 64-bit instruction format 730. The native instructions available in the 64-bit instruction format 730 vary depending on the embodiment. In some embodiments, instructions are partially compressed using a set of index values in the index field 713. The execution unit hardware references a set of compression tables based on the index values and uses the output of the compression tables to reconstruct the native instructions using the 128-bit instruction format 710.
[0091] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel, which represents a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control over certain execution options, such as channel selection (e.g., prediction) and data channel ordering (e.g., blending). For instructions using the 128-bit instruction format 710, the execution size field 716 limits the number of data channels to be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compressed instruction format 730.
[0092] Some execution unit instructions have up to three operands, including two source operands (src0 720, src1 722) and a destination operand 718. In some embodiments, the execution unit supports dual-destination instructions, where one of these destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of an instruction may be an immediate (e.g., hard-coded) value passed with the instruction.
[0093] In some embodiments, the 128-bit instruction format 710 includes an access / address mode 726 field, which, for example, specifies whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits in the instruction.
[0094] In some embodiments, the 128-bit instruction format 710 includes an access / address mode 726 field, which specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, wherein the byte alignment of the access mode determines the access alignment of the instruction operands. For example, in a first mode, the instruction can use byte-aligned addressing for both source and destination operands, and in a second mode, the instruction can use 16-byte aligned addressing for both source and destination operands.
[0095] In one embodiment, the address mode portion of the access / address mode 726 field determines whether the instruction uses direct or indirect addressing. When using direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate number field in the instruction.
[0096] In some embodiments, instructions are grouped based on the 712-bit opcode field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The precise opcode grouping shown is merely exemplary. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares five most significant bits (MSB), where move (mov) instructions are in the form of 0000xxxxb, and logic instructions are in the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jump (jmp)) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mixture of instructions, including synchronous instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). Parallel math instruction group 748 includes component-based arithmetic instructions (e.g., add, multiply) in the form 0100xxxxb (e.g., 0x40). Parallel math group 748 performs arithmetic operations in parallel across data channels. Vector math group 750 includes arithmetic instructions (e.g., dp4) in the form 0101xxxxb (e.g., 0x50). Vector math group performs arithmetic operations, such as dot product, on vector operands.
[0097] Graphics Pipeline
[0098] Figure 8 This is a block diagram of another embodiment of the graphics processor 800. Figure 8 Those elements having the same reference numerals (or names) as those in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein.
[0099] In some embodiments, the graphics processor 800 includes a graphics pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a rendering output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system including one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or by commands issued to the graphics processor 800 via a ring interconnect 802. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted via a command streamer 803, which supplies instructions to individual components of the graphics pipeline 820 or the media pipeline 830.
[0100] In some embodiments, command streamer 803 directs the operation of vertex acquirer 805, which reads vertex data from memory and executes vertex processing commands provided by command streamer 803. In some embodiments, vertex acquirer 805 provides vertex data to vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, vertex acquirer 805 and vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A to 852B via thread dispatcher 831.
[0101] In some embodiments, execution units 852A to 852B are vector processor arrays having an instruction set for performing graphics and media operations. In some embodiments, execution units 852A to 852B have an attached L1 cache 851, which is dedicated to each array or shared between arrays. The cache may be configured as a data cache, an instruction cache, or a single cache partitioned to contain data and instructions in different partitions.
[0102] In some embodiments, the graphics pipeline 820 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some embodiments, a programmable shell shader 811 configures the tessellation operation. A programmable domain shader 817 provides back-end evaluation of the tessellation output. A tessellation unit 813 operates in the direction of the shell shader 811 and includes dedicated logic for generating a detailed set of geometric objects based on a coarse geometry model that is provided as input to the graphics pipeline 820. In some embodiments, if tessellation is not used, the tessellation component (e.g., shell shader 811, tessellation unit 813, domain shader 817) can be bypassed.
[0103] In some embodiments, the complete geometry object may be processed by the geometry shader 819 via one or more threads dispatched to the execution units 852A to 852B, or it may proceed directly to clipping / setting 829. In some embodiments, the geometry shader operates on the entire geometry object (rather than vertices or vertex patches such as those in previous stages of the graphics pipeline). If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 may be programmed by a geometry shader program to perform geometry tessellation when the tessellation unit is disabled.
[0104] Prior to rasterization, clipping / setting 829 processes vertex data. Clipping / setting 829 can be a fixed-function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer and depth testing unit 873 in the render output pipeline 870 dispatch pixel shaders to convert geometric objects into their per-pixel representations. In some embodiments, pixel shader logic is included in thread execution logic 850. In some embodiments, the application can bypass the rasterizer and depth testing unit 873 and access the unrasterized vertex data via the outflow unit 823.
[0105] The graphics processor 800 has an interconnect bus, interconnect structure, or some other interconnect mechanism that allows data and messages to be transferred among the main components of the graphics processor. In some embodiments, execution units 852A to 852B and(multiple) associated caches 851, texture and media samplers 854, and texture / sampler cache 858 are interconnected via data port 856 to perform memory accesses and communicate with the processor's rendering output pipeline components. In some embodiments, samplers 854, caches 851, 858, and execution units 852A to 852B each have a separate memory access path.
[0106] In some embodiments, the rendering output pipeline 870 includes a rasterizer and a depth testing unit 873 that converts vertex-based objects into associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / mask unit for performing fixed-function triangle and line rasterization. Associated rendering cache 878 and depth cache 879 are also available in some embodiments. Pixel manipulation unit 877 performs pixel-based operations on the data; however, in some instances, pixel operations associated with 2D operations (e.g., using mixed bit-block image transfer) are performed by the 2D engine 841, or alternatively by the display controller 843 using an overlay display plane at display time. In some embodiments, a shared L3 cache 875 is available for all graphics components, allowing data to be shared without using main system memory.
[0107] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front-end 834. In some embodiments, the video front-end 834 receives pipeline commands from a command streamer 803. In some embodiments, the media pipeline 830 includes a separate command streamer. In some embodiments, the video front-end 834 processes media commands before sending the commands to the media engine 837. In some embodiments, the media engine 837 includes a thread generation function for generating threads for dispatch to thread execution logic 850 via a thread dispatcher 831.
[0108] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and coupled to the graphics processor via a ring interconnect 802, or some other interconnect bus or mechanism. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 includes dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system-integrated display device (such as in a laptop computer) or an external display device attached via a display device connector.
[0109] In some embodiments, the graphics pipeline 820 and media pipeline 830 are configurable to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some embodiments, the graphics processor's driver software translates API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and computing APIs from the Khronos Group. In some embodiments, support may also be provided for Microsoft's Direct3D library. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the open-source computer vision library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a pipeline mapping from future APIs to the graphics processor's pipeline can be made.
[0110] Graphical Pipeline Programming
[0111] Figure 9A This is a block diagram illustrating a graphics processor command format 900 according to some embodiments. Figure 9B This is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 9A Solid lines in the diagram represent components that are generally included in the drawing command, while dashed lines represent components that are optional or only included in a subset of the drawing command. Figure 9A An exemplary graphics processor command format 900 includes data fields for identifying the target client 902 of the command, a command operation code (opcode) 904, and related data 906 for the command. Some commands also include a sub-opcode 905 and a command size 908.
[0112] In some embodiments, client 902 defines a client unit of a graphics device that processes command data. In some embodiments, a graphics processor command parser examines the client field of each command to adjust further processing of the command and route command data to the appropriate client unit. In some embodiments, the graphics processor client unit includes a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads opcode 904 and sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses information within data field 906 to execute the command. For some commands, an explicit command size 908 is expected to define the size of the command. In some embodiments, the command parser automatically determines the size of at least some commands in the command based on the command opcode. In some embodiments, commands are aligned via multiples of double word length.
[0113] Figure 9B The flowchart illustrates an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of a graphics processor uses a version of the illustrated command sequence to initiate, execute, and terminate a set of graphics operations. Sample command sequences are shown and described for illustrative purposes only, and embodiments are not limited to these specific commands or this command sequence. Moreover, the commands may be issued as a batch of commands in a command sequence, such that the graphics processor will process the command sequence in a manner that is at least partially simultaneous.
[0114] In some embodiments, the graphics processor command sequence 910 may begin with a pipeline dump clearing command 912 to cause any active graphics pipeline to complete its current pending commands. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate simultaneously. Pipeline dump clearing is performed to cause any pending commands to be completed by the active graphics pipeline. In response to pipeline dump clearing, the command parser for the graphics processor will stop command processing until the active rendering engine completes its pending operations and invalidates the associated read cache. Optionally, any data marked 'dirty' in the render cache may be dumped and cleared into memory. In some embodiments, pipeline dump clearing command 912 may be used for pipeline synchronization or before placing the graphics processor into a low-power state.
[0115] In some embodiments, a pipeline selection command 913 is used when a sequence of commands requires the graphics processor to explicitly switch between pipelines. In some embodiments, the pipeline selection command 913 is only required once in an execution context before a pipeline command is issued, unless the context requires issuing commands for two pipelines. In some embodiments, a pipeline dump clearing command 912 is required exactly before the pipeline switch via the pipeline selection command 913.
[0116] In some embodiments, pipeline control command 914 configures a graphics pipeline for operation and programs the 3D pipeline 922 and the media pipeline 924. In some embodiments, pipeline control command 914 configures the pipeline state of an active pipeline. In one embodiment, pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within an active pipeline before processing a batch of commands.
[0117] In some embodiments, the command for returning buffer state 916 is used to configure a set of return buffers for corresponding pipelined write data. Some pipelined operations require allocating, selecting, or configuring one or more return buffers, in which intermediate data is written during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, configuring return buffer state 916 includes selecting the size and number of return buffers for a set of pipelined operations.
[0118] The remaining commands in the command sequence vary based on the active pipeline used for the operation. Based on pipeline determination 920, the command sequence is tailored for either the 3D pipeline 922 starting at 3D pipeline state 930, or the media pipeline 924 starting at media pipeline state 940.
[0119] Commands for configuring 3D pipeline states 930 include 3D state setting commands for vertex buffer states, vertex element states, constant color states, depth buffer states, and other state variables to be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the specific 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass specific pipeline components (if those components will not be used).
[0120] In some embodiments, the 3D primitive 932 command is used to submit 3D primitives to be processed by the 3D pipeline. The command and associated parameters passed to the graphics processor via the 3D primitive 932 command are forwarded to the vertex acquisition function in the graphics pipeline. The vertex acquisition function uses the 3D primitive 932 command data to generate multiple vertex data structures. These vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 command is used to perform vertex operations on the 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution unit.
[0121] In some embodiments, the 3D pipeline 922 is triggered by executing command 934 or an event. In some embodiments, register writing triggers command execution. In some embodiments, execution is triggered via a 'go' or 'kick' command in a command sequence. In one embodiment, pipeline synchronization commands are used to trigger command execution so that the command sequence is cleared via a graphics pipeline dump. The 3D pipeline performs geometry processing on 3D primitives. Once the operation is complete, the resulting geometry is rasterized, and the pixel engine shades the resulting pixels. Additional commands for controlling pixel shading and pixel backend operations may also be included for these operations.
[0122] In some embodiments, when performing media operations, a sequence of graphics processor commands 910 follows the media pipeline 924 path. Generally, the specific purpose and manner of programming the media pipeline 924 depends on the media or computational operation to be performed. During media decoding, specific media decoding operations can be offloaded to the media pipeline. In some embodiments, the media pipeline can also be bypassed, and media decoding can be performed wholly or partially using resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, wherein the graphics processor is used to perform SIMD vector operations using computation shader programs that are not explicitly associated with rendering graphics primitives.
[0123] In some embodiments, the media pipeline 924 is configured in a manner similar to that of the 3D pipeline 922. A set of commands for configuring media pipeline state 940 is dispatched or placed in a command queue before the media object command 942. In some embodiments, the commands for media pipeline state 940 include data for configuring media pipeline elements that will be used to process media objects. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands for media pipeline state 940 also support directing one or more pointers to "indirect" state elements that contain a batch of state settings.
[0124] In some embodiments, media object command 942 supplies pointers to a media object for processing by the media pipeline. The media object includes a memory buffer containing video data to be processed. In some embodiments, all media pipeline states must be valid before issuing media object command 942. Once the pipeline states are configured and media object command 942 is queued, media pipeline 924 is triggered via execution command 944 or an equivalent execution event (e.g., register write). The output from media pipeline 924 can then be post-processed by operations provided by 3D pipeline 922 or media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.
[0125] Graphical software architecture
[0126] Figure 10 An exemplary graphics software architecture of a data processing system 1000 according to some embodiments is illustrated. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.
[0127] In some embodiments, the 3D graphics application 1010 includes one or more shader programs, which include shader instructions 1012. The shader language instructions may employ a high-level shader language, such as High-Level Shading Language (HLSL) or OpenGL Shading Language (GLSL). The application also includes executable instructions 1014, which employ a machine language suitable for execution by a general-purpose processor core 1034. The application also includes graphics objects 1016 defined by vertex data.
[0128] In some embodiments, the operating system 1020 is from Microsoft Corporation. The operating system 1020 may be a dedicated UNIX-like operating system or an open-source UNIX-like operating system using a variant of the Linux kernel. The operating system 1020 may support graphics APIs 1022, such as the Direct3D API, OpenGL API, or Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. This compilation may be just-in-time (JIT) compilation or pre-compilation of the application-executable shaders. In some embodiments, high-level shaders are compiled into low-level shaders during the compilation of the 3D graphics application 1010. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the standard Portable Intermediate Representation (SPIR) used by the Vulkan API.
[0129] In some embodiments, the user-mode graphics driver 1026 includes a back-end shader compiler 1027 for transforming shader instructions 1012 into a hardware-specific representation. When using the OpenGL API, shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses operating system kernel-mode functionality 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.
[0130] IP core implementation
[0131] One or more aspects of at least one embodiment can be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit (such as a processor). For example, the machine-readable medium may include instructions representing various logic within a processor. When read by a machine, these instructions can cause the machine to manufacture logic for performing the techniques described herein. Such representations (referred to as “IP cores”) are reusable units of logic for an integrated circuit that can be stored on a tangible, machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model can be supplied to consumers or manufacturing facilities that load the hardware model onto manufacturing machines that manufacture integrated circuits. Integrated circuits can be manufactured such that the circuits perform the operations described in association with any of the embodiments described herein.
[0132] Figure 11 This is a block diagram illustrating an IP core development system 1100 that can be used to manufacture integrated circuits to perform operations, according to an embodiment. The IP core development system 1100 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build entire integrated circuits (e.g., SOC integrated circuits). Design facility 1130 can generate software simulations 1110 of the IP core design using a high-level programming language (e.g., C / C++). Software simulation 1110 can be used to design, test, and verify the behavior of the IP core using simulation model 1112. Simulation model 1112 can include functional, behavioral, and / or timing simulations. Register transfer level (RTL) designs 1115 can then be created or synthesized from simulation model 1112. RTL design 1115 is an abstraction of the behavior of an integrated circuit (including associated logic performed using the modeled digital signals) that models the flow of digital signals between hardware registers. In addition to RTL design 1115, lower-level designs at logic or transistor levels can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.
[0133] The RTL design 1115 or an equivalent can be further synthesized into a hardware model 1120 by the design facility. This hardware model may employ a Hardware Description Language (HDL) or some other representation of the physical design data. The HDL can be further simulated or tested to validate the IP core design. The IP core design can be stored in non-volatile memory 1140 (e.g., hard disk, flash memory, or any non-volatile storage medium) for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted (e.g., via the Internet) through a wired connection 1150 or a wireless connection 1160. The manufacturing facility 1165 can then fabricate an integrated circuit at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.
[0134] Exemplary System-on-Chip Integrated Circuit
[0135] Figures 12 to 14 Exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein are shown. In addition to those shown, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0136] Figure 12This is a block diagram illustrating an exemplary system-on-chip integrated circuit 1200 that can be fabricated using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., a CPU), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, any of which may be modular IP cores from the same or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I2S / I2C controller 1240. Additionally, the integrated circuit may include a display device 1245 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1250 and a Mobile Industry Processor Interface (MIPI) display interface 1255. Storage may be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface may be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. In addition, some integrated circuits also include the embedded security engine 1270.
[0137] Figure 13 This is a block diagram illustrating an exemplary graphics processor 1310, a system-on-a-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. The graphics processor 1310 may be... Figure 12 A variant of the graphics processor 1210. The graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A to 1315N (e.g., 1315A, 1315B, 1315C, 1315D, up to 1315N-1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic, such that the vertex processor 1305 is optimized to perform vertex shader program operations, while the one or more fragment processors 1315A to 1315N perform fragment (e.g., pixel) shading operations for use in fragment or pixel shader programs. The vertex processor 1305 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. The fragment processors (multiple) 1315A to 1315N use the primitive and vertex data generated by the vertex processor 1305 to produce frame buffers displayed on a display device. In one embodiment, fragment processors (multiple) 1315A to 1315N are optimized to execute fragment shader programs provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs provided in the Direct 3D API.
[0138] Additionally, the graphics processor 1310 includes one or more memory management units (MMUs) 1320A to 1320B, one or more caches 1325A to 1325B, and (multiple) circuit interconnects 1330A to 1330B. The one or more MMUs 1320A to 1320B provide virtual-to-physical address mappings for the graphics processor 1310, including for vertex processors 1305 and / or (multiple) fragment processors 1315A to 1315N. These virtual-to-physical address mappings may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in the one or more caches 1325A to 1325B. In one embodiment, the one or more MMUs 1320A to 1320B may interact with other MMUs within the system, including those related to... Figure 12 One or more application processors 1205, image processor 1215, and / or video processor 1220 are associated with one or more MMUs for synchronization, enabling each processor 1205 to 1220 to participate in a shared or unified virtual memory system. According to an embodiment, one or more circuit interconnects 1330A to 1330B enable the graphics processor 1310 to interact with other IP cores within the SoC via the SoC's internal bus or via a direct connection.
[0139] Figure 14 This is a block diagram illustrating an additional exemplary graphics processor 1410 of a system-on-a-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. The graphics processor 1410 may be... Figure 12 A variant of the graphics processor 1210. The graphics processor 1410 includes... Figure 13 The integrated circuit 1300 includes one or more MMUs 1320A to 1320B, (multiple) caches 1325A to 1325B, and (multiple) circuit interconnects 1330A to 1330B.
[0140] The graphics processor 1410 includes one or more shader cores 1415A to 1415N (e.g., 1415A, 1415B, 1415C, 1415D, 1415E, 1415F, up to 1315N-1 and 1315N), which provide a unified shader core architecture, wherein a single core or type or core can execute all types of programmable shader code, including shader program code to implement vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present may vary in embodiments and implementations. Additionally, the graphics processor 1410 includes an inter-core task manager 1405, which acts as a thread dispatcher for distributing execution threads to one or more shader cores 1415A to 1415N and a chunking unit 1418 for accelerating chunking operations for tile-based rendering, wherein scene rendering operations are subdivided in the image space, for example to utilize local spatial consistency within the scene or optimize the use of internal caches.
[0141] Exemplary Graphics Virtualization Architecture
[0142] Some embodiments of the present invention are implemented on platforms utilizing full graphics processing unit (GPU) virtualization. Thus, an overview of GPU virtualization techniques employed in one embodiment of the present invention is provided below, followed by a detailed description of apparatus and methods for mode-driven page table shading.
[0143] One embodiment of the invention employs a full GPU virtualization environment running native graphics drivers in guest machines, and an intermediary that enables good performance, scalability, and secure isolation between guest machines. This embodiment provides each virtual machine (VM) with a virtual full-featured GPU that, in most cases, can directly access performance-critical resources without intervention from the hypervisor, while capturing and emulating privileged operations from guest machines at minimal cost. In one embodiment, a virtual GPU (vGPU) with full GPU features is presented to each VM. In most cases, the VM can directly access performance-critical resources without intervention from the hypervisor, while capturing and emulating privileged operations from guest machines to provide secure isolation between VMs. Each VM switches its vGPU context to share the physical GPU among multiple VMs.
[0144] Figure 15A high-level system architecture on which embodiments of the present invention can be implemented is illustrated. This high-level system architecture includes a graphics processing unit (GPU) 1500, a central processing unit (CPU) 1520, and system memory 1510 shared between the GPU 1500 and the CPU 1520. A rendering engine 1502 retrieves GPU commands from a command buffer 1512 in the system memory 1510 to accelerate graphics rendering using various different features. A rendering engine 1504 retrieves pixel data from a frame buffer 1514 and then sends the pixel data to an external monitor for display.
[0145] Some architectures use system memory 1510 as graphics memory, while other GPUs can use on-die memory. System memory 1510 can be mapped into multiple virtual address spaces via GPU page table 1506. A 2GB global virtual address space, called global graphics memory, can be accessed from GPU 1500 and CPU 1520 and is mapped via global page table. Native graphics memory space is supported as multiple 2GB native virtual address spaces, but access is limited to the rendering engine 1502 via native page table. Global graphics memory is mostly used as frame buffer 1514, but also as command buffer 1512. During hardware acceleration, a large amount of data access is performed to native graphics memory. GPUs with on-die memory employ a similar page table mechanism.
[0146] In one embodiment, the CPU 1520 programs the GPU 1500 using GPU-specific commands in a producer-consumer model, such as... Figure 15 As shown. Based on high-level programming APIs such as OpenGL and DirectX, the graphics driver programs GPU commands into command buffers 1512, including a main buffer and batch buffers. The GPU 1500 then fetches and executes the commands. The main buffer (circular buffer) can link other batch buffers together. The terms "main buffer" and "circular buffer" are used interchangeably below. Batch buffers are used to pass the majority (up to ~98%) of the commands for each programming model. Register tuples (head, tail) are used to control the circular buffers. In one embodiment, the CPU 1520 submits commands to the GPU 1500 by updating the tail, while the GPU 1500 fetches commands from the head and then notifies the CPU 1520 by updating the head after the commands have been executed.
[0147] As described above, one embodiment of the present invention is implemented in a full GPU virtualization platform with intermediary passing. Therefore, each VM is equipped with a full-featured GPU to run native graphics drivers within the VM. However, significant challenges arise in three areas: (1) the complexity of virtualizing an entire sophisticated modern GPU, (2) the performance issues resulting from multiple VMs sharing a GPU, and (3) complete security isolation between VMs.
[0148] Figure 16 A GPU virtualization architecture according to an embodiment of the present invention is illustrated, the GPU virtualization architecture including a hypervisor 1610 running on a GPU 1600, a privileged virtual machine (VM) 1620, and one or more user VMs 1631 to 1632. A virtualization stub module 1611 running in the hypervisor 1610 extends memory management to include an extended page table (EPT) 1614 for user VMs 1631 to 1632, and a privileged virtual memory management unit (PVMMU) 1612 for the privileged VM 1620, to implement capture and delivery strategies. In one embodiment, each VM 1620, 1631 to 1632 runs a native graphics driver 1628, which can directly access performance-critical resources such as frame buffers and command buffers using resource partitions as described below. To protect privileged resources, namely I / O registers and PTEs, corresponding accesses from user VMs 1631 to 1632 and the graphics driver 1628 in privileged VM 1620 are captured and forwarded to the virtualization mediator 1622 in privileged VM 1620 for emulation. In one embodiment, as shown in the figure, the virtualization mediator 1622 uses a hypercall to access the physical GPU 1600.
[0149] Additionally, in one embodiment, the virtualization mediator 1622 implements a GPU scheduler 1626 that runs concurrently with the CPU scheduler 1616 in the hypervisor 1610 to share the physical GPU 1600 among VMs 1631 to 1632. One embodiment uses the physical GPU 1600 to directly execute all commands submitted from the VMs, thus avoiding the complexity of emulating the rendering engine, the most complex part within the GPU. Simultaneously, resource passing of the frame buffer and command buffer minimizes the hypervisor 1610's interference with CPU access, while the GPU scheduler 1626 ensures that each VM quantum is used for direct GPU execution. Therefore, the illustrated embodiment achieves good performance when sharing a GPU among multiple VMs.
[0150] In one embodiment, the virtualization stub module 1611 selectively captures or forwards guest access to certain GPU resources. The virtualization stub module 1611 manipulates the EPT 1614 entry to selectively present or hide specific address ranges to user VMs 1631 through 1632, while using the reserved bit PTE in PVMMU 1612 for privileged VM 1620 to selectively capture or forward guest access to specific address ranges. In both cases, peripheral input / output (PIO) access is captured. All captured accesses are forwarded to the virtualization mediator 1622 for emulation, while the virtualization stub module 1611 uses a hypercall to access the physical GPU 1600.
[0151] As described above, in one embodiment, the virtualization mediator 1622 emulates the virtual GPU (vGPU) 1624 for privileged resource access and performs context switching between vGPUs 1624. Simultaneously, the privileged VM 1620 graphics driver 1628 is used to initialize the physical device and manage power. One embodiment employs a flexible distribution model, simplifying the binding between the virtualization mediator 1622 and the hypervisor 1610 by implementing the virtualization mediator 1622 as a kernel module within the privileged VM 1620.
[0152] The separate CPU / GPU scheduling mechanism is implemented via CPU scheduler 1616 and GPU scheduler 1626. This is because the cost of GPU context switching can be more than 1000 times that of CPU context switching (e.g., ~700µs vs. ~300ns). Furthermore, the number of CPU cores in a computer system may differ from the number of GPU cores. Therefore, in one embodiment, GPU scheduler 1626 is implemented separately from the existing CPU scheduler 1616. This separate scheduling mechanism results in the need for concurrent access to resources from both the CPU and GPU. For example, while the CPU is accessing the graphics memory of VM1 1631, the GPU may simultaneously access the graphics memory of VM2 1632.
[0153] As described above, in one embodiment, a native graphics driver 1628 is executed within each VM 1620, 1631 to 1632. This native graphics driver directly accesses a subset of performance-critical resources through privileged operations emulated by the virtualization mediator 1622. This decoupled scheduling mechanism leads to the resource partitioning design described below. To better support resource partitioning, one embodiment preserves a memory-mapped I / O (MMIO) register window to transmit resource partitioning information to the VM.
[0154] In one embodiment, the location and definition of virt_info have been incorporated into the hardware specification as a virtualization extension, so the graphics driver 1628 natively handles the extension, and future GPU builds follow the specification for backward compatibility.
[0155] Although Figure 16 While shown as a separate component, in one embodiment, the privileged VM 1620, which includes the virtualization mediator 1622 (and its vGPU instance 1624 and GPU scheduler 1626), is implemented as a module within the hypervisor 1610.
[0156] In one embodiment, virtualization mediator 1622 manages all VMs' vGPUs 1624 by capturing and emulating privileged operations. Virtualization mediator 1622 handles physical GPU interrupts and can generate virtual interrupts for designated VMs 1631 to 1632. For example, a physical completion interrupt of command execution might trigger a virtual completion interrupt and be passed to the render owner. The idea of emulating each semantic vGPU instance is simple; however, implementation requires significant engineering work and a deep understanding of GPU 1600. For example, some graphics drivers may access approximately 700 I / O registers.
[0157] In one embodiment, GPU scheduler 1626 implements a coarse-grained Quality of Service (QoS) policy. A specific amount of time can be selected as the time slice for each VM 1631 to 1632 to share the GPU 1600 resources. For example, in one embodiment, a time slice of 16ms is selected as the scheduling time slice because this value results in low perceptibility to human image changes. This relatively large quantum is also selected because the cost of GPU context switching is more than 1000 times that of CPU context switching, so the quantum cannot be as small as the time slice in CPU scheduler 1616. Commands from VMs 1631 to 1632 are submitted to GPU 1600 continuously until the guest / VM exhausts its time slice. In one embodiment, GPU scheduler 1626 waits for the guest ring buffer to become idle before switching, which could affect fairness since most GPUs today are non-preemptive. To minimize waiting overhead, a coarse-grained traffic control mechanism can be implemented by tracking command submissions to ensure that the backlog of commands is within certain limits at any given time. Therefore, the time drift between the allocated time slices and the execution time is relatively small, thus enabling a coarse-grained QoS strategy.
[0158] In one embodiment, during a rendering context switch, when switching rendering engines between vGPU 1624, the internal pipeline state and I / O register state are saved and restored, and a cache / TLB dump is performed. The internal pipeline state is not visible to the CPU but can be saved and restored via GPU commands. Saving / restoring the I / O register state is achieved by reading / writing the register list in the rendering context. The internal cache and Translation Lookaside Buffer (TLB) included in modern GPUs for accelerating data access and address translation must be dumped and cleared using commands at the rendering context switch to ensure isolation and correctness. In one embodiment, the steps for switching contexts are: 1) saving the current I / O state, 2) dumping and clearing the current context, 3) saving the current context using additional commands, 4) restoring the new context using additional commands, and 5) restoring the I / O state of the new context.
[0159] As described above, one embodiment uses a dedicated ring buffer to hold additional GPU commands. The (audited) guest ring buffer can be reused for performance improvements, but directly inserting commands into the guest ring buffer is unsafe because the CPU might continue to queue more commands, leading to overwritten content. To avoid contention, one embodiment switches from the guest ring buffer to its own dedicated ring buffer. At the end of the context switch, this embodiment switches from the dedicated ring buffer to the guest ring buffer of the new VM.
[0160] One embodiment reuses the privileged VM 1620 graphics driver to initialize the display engine, and then manages the display engine to display different VM frame buffers.
[0161] When two vGPU 1624s have the same resolution, only the frame buffer position is switched. For different resolutions, privileged VMs can use a hardware scaler, a common feature in modern GPUs that automatically scales the resolution. Both techniques take only a few milliseconds. In many cases, display management may not be necessary, such as when the VM is not displayed on a physical monitor (e.g., when the VM is located on a remote server).
[0162] like Figure 16 As shown, one embodiment passes access to the frame buffer and command buffer to accelerate performance-critical operations from VMs 1631 to 1632. For the 2GB global GPU memory space, GPU memory resource partitioning and address space expansion techniques can be employed. For the native GPU memory space, each space is also 2GB in size. Since native GPU memory can only be accessed by the GPU 1600, per VM native GPU memory can be implemented through rendering context switching.
[0163] As described above, one embodiment partitions the global graphics memory between VMs 1631 and 1632. As mentioned above, the separate CPU / GPU scheduling mechanism requires the CPU and GPU to access the global graphics memory of different VMs simultaneously; therefore, each VM must use its own resources for rendering at any given time, leading to the resource partitioning method for the global graphics memory.
[0164] Figure 17 Additional details are shown for one embodiment of a device 1700 with a graphics virtualization architecture, which includes multiple VMs (e.g., VMs 1730 and 1740) managed by a hypervisor 1710, including access to the full GPU feature array in a GPU 1720. In various embodiments, the hypervisor 1710 enables VMs 1730 or 1740 to use graphics memory and other GPU resources for GPU virtualization. Based on GPU virtualization technology, one or more virtual GPUs (vGPUs) (e.g., vGPUs 1760A and 1760B) can access the full functionality provided by the GPU 1720 hardware. In various embodiments, the hypervisor 1710 can track and manage the resources and lifecycle of vGPUs 1760A and 1760B as described herein.
[0165] In some embodiments, vGPU 1760A-B may include a virtual GPU device presented to VMs 1730 and 1740, and may be used to interact with native GPU drivers (e.g., as described above relative to...). Figure 16 (as described above). Then, VM 1730 or VM 1740 can access the entire GPU feature array and access the virtual graphics processor using the virtual GPU devices in vGPU 1760A-B. For example, once VM 1730 is captured in hypervisor 1710, hypervisor 1710 can manipulate the vGPU instance (e.g., vGPU 1760A) and determine whether VM 1730 can access the virtual GPU devices in vGPU 1760A. The vGPU context can be switched on a per-quantum or per-event basis. In some embodiments, context switching can occur per GPU rendering engine (such as 3D rendering engine 1722 or bitblock transmitter rendering engine 1724). Periodic switching allows multiple VMs to share the physical GPU in a manner transparent to the VM's workload.
[0166] GPU virtualization can take various forms. In some embodiments, device delivery can be used to enable VM 1730, where the entire GPU 1720 is presented to VM 1730 as if they were directly connected. Just as a single central processing unit (CPU) core can be designated for dedicated use by VM 1730, GPU 1720 can also be designated for dedicated use by VM 1730 (e.g., even for a limited time). Another virtualization model is a time-sharing model, where GPU 1720 or a portion thereof can be shared by multiple VMs (e.g., VM 1730 and VM 1740) in a multiplexed manner. In other embodiments, device 1700 may also use other GPU virtualization models. In various embodiments, the graphics memory associated with GPU 1720 can be partitioned and allocated to each vGPU 1760A-B in hypervisor 1710.
[0167] In various embodiments, the Graphics Translation Table (GTT) can be used by the VM or GPU 1720 to map graphics processor memory to system memory or to translate GPU virtual addresses to physical addresses. In some embodiments, the hypervisor 1710 can manage graphics memory mapping via a shadowed GTT, and the shadowed GTT can be maintained in a vGPU instance (e.g., vGPU 1760A). In various embodiments, each VM can have a corresponding shadowed GTT to maintain the mapping between graphics memory addresses and physical memory addresses (e.g., machine memory addresses in a virtualized environment). In some embodiments, the shadowed GTT can be shared and maintain mappings for multiple VMs.
[0168] In some embodiments, each VM 1730 or VM 1740 may include both per-process GTT and global GTT.
[0169] In some embodiments, device 1700 can use system memory as graphics memory. System memory can be mapped to multiple virtual address spaces via GPU page tables. Device 1700 can support a global graphics memory space and a per-process graphics memory address space. The global graphics memory space can be a virtual address space (e.g., 2GB) mapped via a global graphics translation table (GGTT). The lower portion of this address space is sometimes referred to as an opening accessible from GPU 1720 and CPU (not shown). The upper portion of this address space is referred to as a high-order graphics memory space or hidden graphics memory space that can only be used by GPU 1720. In various embodiments, a shaded global graphics translation table (SGGTT) can be used by VM 1730, VM 1740, hypervisor 1710, or GPU 1720 to translate graphics memory addresses to corresponding system memory addresses based on the global memory address space.
[0170] In full GPU virtualization, static global graphics memory (GRAM) partitioning schemes may face scalability issues. For example, with a 2GB GRAM space, the first 512 megabytes (MB) of virtual address space can be reserved for the opening, and the remaining portion (1536MB) can become the high-order (hidden) GRAM space. Using a static GRAM partitioning scheme, each VM enabling full GPU virtualization can be allocated 128MB of opening and 384MB of high-order GRAM space. Therefore, a 2GB GRAM space can only accommodate a maximum of four VMs.
[0171] Besides scalability issues, VMs with limited graphics memory space can also suffer from performance degradation. Sometimes, when media applications extensively utilize GPU media hardware acceleration, severe performance degradation can be observed during some of the media-heavy workloads of those applications. For example, decoding a single channel of 1080p H.264 / AVC bitstream may require at least 40MB of graphics memory. Therefore, decoding 10 channels of 1080p H.264 / AVC bitstream may require at least 400MB of graphics memory. Additionally, some graphics memory may need to be reserved for surface compositing / color conversion, switching display frame buffers during decoding, etc. In this case, 512MB of graphics memory per VM may be insufficient for the VM to run multiple video encoding or decoding operations.
[0172] In various embodiments, device 1700 may utilize on-demand SGGTT to handle GPU graphics memory overuse. In some embodiments, hypervisor 1710 may build SGGTT on demand, which may include all unused translations of graphics memory virtual addresses for the owner VMs of different GPU components.
[0173] In various embodiments, at least one VM managed by hypervisor 1710 may be allocated a global graphics memory address space and memory that is greater than a static partition. In some embodiments, at least one VM managed by hypervisor 1710 may be allocated or have access to the entire high-order graphics memory address space. In some embodiments, at least one VM managed by hypervisor 1710 may be allocated or have access to the entire graphics memory address space.
[0174] The hypervisor / VMM 1710 can use command parser 1718 to detect the potential memory working set of the GPU rendering engine for commands submitted by VM 1730 or VM 1740. In various embodiments, VM 1730 may have a corresponding command buffer (not shown) for holding commands from 3D workload 1732 or media workload 1734. Similarly, VM 1740 may have a corresponding command buffer (not shown) for holding commands from 3D workload 1742 or media workload 1744. In other embodiments, VM 1730 or VM 1740 may have other types of graphics workloads.
[0175] In various embodiments, command parser 1718 can scan commands from the VM and determine whether the commands contain memory operands. If so, the command parser can, for example, read the relevant graphics memory space mapping from the VM's GTT and then write it into a workload-specific portion of the SGGTT. After scanning the entire command buffer of the workload, the SGGTT that maintains the memory address space mapping associated with this workload can be generated or updated. Additionally, by scanning pending commands from VM 1730 or VM 1740, command parser 1718 can also improve the security of GPU operations (e.g., by mitigating malicious operations).
[0176] In some embodiments, an SGGTT can be generated to maintain the translation of all workloads across all VMs. In some embodiments, an SGGTT can be generated to maintain, for example, the translation of all workloads for only one VM. Workload-specific SGGTT portions can be constructed on demand by command parser 1718 to maintain the translation of a specific workload (e.g., 3D workload 1732 for VM 1730 or media workload 1744 for VM 1740). In some embodiments, command parser 1718 can insert SGGTTs into SGGTT queue 1714 and corresponding workloads into workload queue 1716.
[0177] In some embodiments, the GPU scheduler 1712 can construct such on-demand SGGTT at execution time. A particular hardware engine may use only a small portion of the graphics memory address space allocated to the VM 1730 at execution time, and GPU context switching occurs infrequently. To take advantage of this GPU characteristic, the hypervisor 1710 can use the VM 1730's SGGTT to maintain only the execution and pending translations for individual GPU components (rather than the entire portion of the global graphics memory address space allocated to the VM 1730).
[0178] The GPU scheduler 1712 of GPU 1720 can be decoupled from the CPU scheduler in device 1700. In some embodiments, to utilize hardware parallelism, the GPU scheduler 1712 can schedule workloads of different GPU engines (e.g., 3D rendering engine 1722, bitblock transmitter rendering engine 1724, video command stream converter (VCS) rendering engine 1726, and video enhanced command stream converter (VECS) rendering engine 1728) separately. For example, VM 1730 may be 3D enhanced, and 3D workload 1732 may need to be scheduled to 3D rendering engine 1722 at a time. Meanwhile, VM 1740 may be media enhanced, and media workload 1744 may need to be scheduled to VCS rendering engine 1726 and / or VECS rendering engine 1728. In this case, the GPU scheduler 1712 can schedule the 3D workload 1732 of VM 1730 and the media workload 1744 of VM 1740 separately.
[0179] In various embodiments, the GPU scheduler 1712 can track the SGGTTs in execution used by the corresponding rendering engine in the GPU 1720. In this case, the hypervisor 1710 can retain SGGTTs for each rendering engine to track all in-execution GPU memory working sets in the corresponding rendering engine. In some embodiments, the hypervisor 1710 can retain a single SGGTT to track all in-execution GPU memory working sets for all rendering engines. In some embodiments, this tracking can be based on a separate in-execution SGGTT queue (not shown). In some embodiments, this tracking can be based on tags on SGGTT queue 1714 (e.g., using a registry). In some embodiments, this tracking can be based on tags on workload queue 1716 (e.g., using a registry).
[0180] During scheduling, GPU scheduler 1712 may check the SGGTTs in SGGTT queue 1714 for the scheduled workloads in workload queue 1716. In some embodiments, in order to schedule the next VM for a particular rendering engine, GPU scheduler 1712 may check whether the graphics memory working set of the specific workload used by the VM of that rendering engine conflicts with a graphics memory working set being executed or to be executed by that rendering engine. In other embodiments, this conflict checking may be extended to checks performed by all other rendering engines using executing or to-be-executed graphics memory working sets. In various embodiments, this conflict checking may be based on the corresponding SGGTTs in SGGTT queue 1714 or on the SGGTTs held by hypervisor 1710 for tracking all executing graphics memory working sets in the corresponding rendering engines as discussed above.
[0181] If no conflict exists, the GPU scheduler 1712 can integrate the running and pending GPU memory worksets together. In some embodiments, the SGGTTs generated by the running and pending GPU memory worksets of a particular rendering engine can also be generated and stored, for example, in the SGGTT queue 1714 or other data storage device. In some embodiments, the SGGTTs generated by the running and pending GPU memory worksets of all rendering engines associated with a VM can also be generated and stored, if the GPU memory addresses of all these workloads do not conflict with each other.
[0182] Before submitting the selected VM workload to GPU 1720, hypervisor 1710 may write the corresponding SGGTT pages to GPU 1720 (e.g., to graphics translation table 1750). Thus, hypervisor 1710 can enable this workload to be executed using the correct mapping in the global graphics memory space. In various embodiments, all these translation entries may be written to graphics translation table 1750, to lower memory space 1754 or upper memory space 1752. In some embodiments, graphics translation table 1750 may contain a separate table per VM to hold these translation entries. In other embodiments, graphics translation table 1750 may also contain a separate table per rendering engine to accommodate these translation entries. In various embodiments, graphics translation table 1750 may contain at least the graphics memory address to be executed.
[0183] However, if a conflict is identified by the GPU scheduler 1712, then the GPU scheduler 1712 may delay the scheduling of this VM and instead attempt to schedule another workload of the same or different VM. In some embodiments, such a conflict can be detected if two or more VMs can attempt to use the same graphics memory address (e.g., for the same rendering engine or two different rendering engines). In some embodiments, the GPU scheduler 1712 may change its scheduler policy to avoid selecting one or more rendering engines that may conflict with each other. In some embodiments, the GPU scheduler 1712 may suspend the execution hardware engine to mitigate the conflict.
[0184] In some embodiments, memory overuse during GPU virtualization, as discussed herein, can coexist with a static global graphics memory space partitioning scheme. As an example, the opening in the lower memory space 1754 can still be used for static partitioning of all VMs. The high-order graphics memory space in the upper memory space 1752 can be used for the memory overuse scheme. Compared to the static global graphics memory space partitioning scheme, memory overuse during GPU virtualization allows each VM to utilize the entire high-order graphics memory space in the upper memory space 1752, which can allow some applications within each VM to use a larger graphics memory space for improved performance.
[0185] In a static global graphics memory space partitioning scheme, VMs that initially require a large portion of the memory to be protected may only be able to use a small fraction at runtime, while other VMs may be left without sufficient memory. In cases of memory overuse, the hypervisor can allocate memory to VMs on demand, and the saved memory can be used to support more VMs. In an SGGTT-based memory overuse scheme, only the graphics memory space needed for the workloads to be executed can be allocated at runtime, saving graphics memory space and supporting more VMs accessing the GPU 1720.
[0186] The current architecture supports hosting GPU workloads in cloud and data center environments. Full GPU virtualization is one of the fundamental supporting technologies used in GPU clouds. In full GPU virtualization, the virtual machine monitor (VMM), especially the virtual GPU (vGPU) driver, captures and emulates guest access to privileged GPU resources for security and multiplexing, while simultaneously allowing access to performance-critical resources such as the CPU, and vice versa. Once a GPU command is submitted, it is executed directly by the GPU without VMM intervention. Therefore, near-native performance is achieved.
[0187] The current system uses the GPU engine's system memory to access the Global Graphics Translation Table (GGTT) and / or the Per-Process Graphics Translation Table (PPGTT) to translate GPU graphics memory addresses to system memory addresses. Masking mechanisms can be used for the guest GPU page table's GGTT / PPGTT.
[0188] The VMM can use a shadowed PPGTT synchronized with the guest PPGTT. The guest PPGTT has write protection, allowing the shadowed PPGTT to be continuously synchronized with the guest PPGTT by capturing and emulating guest modifications of its PPGTT. Currently, the GGTT for each vGPU is shadowed and partitioned across each VM, and the PPGTT is shadowed and masked on each VM (e.g., per-process). Shadowing the GGTT page table is straightforward because the GGTT PDE table is kept within the PCI bar0 MMIO range. However, PPGTT shadowing relies on write protection of the guest PPGTT page table, and traditional shadowed page tables are complex (and therefore vulnerable) and inefficient. For example, CPU shadowed page tables incur a performance overhead of ~30% in current architectures. Therefore, in some systems, an enlightened shadowed page table is used, which modifies the guest graphics driver to collaboratively identify pages for page table pages and / or modifies the guest graphics driver when the page is released.
[0189] Embodiments of the present invention include a memory management unit (MMU), such as an I / O memory management unit (IOMMU), to remap client page numbers (GPNs) from client PPGTT mappings to host page numbers (HPNs) without relying on inefficient / complex shadowed PPGTTs. Simultaneously, one embodiment preserves the global shadowed GGTT page table for address inflation. These techniques are commonly referred to as Hybrid Layer Address Mapping (HLAM).
[0190] By default, IOMMU cannot be used in certain intermediary transport architectures because multiple VMs can only use a single Layer 2 translation. One embodiment of the present invention utilizes the following technique to solve this problem:
[0191] 1. Use IOMMU to perform two-level translation without shadowed PPGTT. Specifically, in one embodiment, the GPU translates from the graphics memory address (GM_ADDR) to the GPN, and the IOMMU translates from the GPN to the HPN, instead of the shadowed PPGTT from GM_ADDR to HPN, where write protection is applied to the guest PPGTT.
[0192] 2. In one embodiment, the IOMMU page table is managed for each VM and is switched (or may be partially switched) when a vGPU is switched. That is, when a VM / vGPU is scheduled, the corresponding VM's IOMMU page table is loaded.
[0193] 3. However, in one embodiment, the address of the GGTT mapping is shared, and since the vCPU can access the address of the GGTT mapping (e.g., an opening), the global shadow GGTT must remain valid even when the vGPU of the VM is not scheduled. Thus, one embodiment of the invention uses a hybrid layer address translation that preserves the global shadow GGTT but uses the guest PPGTT directly.
[0194] 4. In one embodiment, the GPN address space is partitioned to move GPN addresses mapped by the GGTT (which become inputs to the IOMMU, such as GPN) to a dedicated address range. This can be achieved by capturing and emulating the GGTT page table. In one embodiment, the GPN is modified from the GGTT with a large offset to avoid overlap with the PPGTT in the IOMMU mapping.
[0195] Figure 18 The architecture shown in one embodiment is illustrated, wherein the IOMMU 1830 enables device virtualization. The illustrated architecture includes two VMs 1801 and 1811 running on the hypervisor / VMM 1820 (however, the basic principles of the invention can be implemented with any number of VMs). Each VM 1801 and 1811 includes drivers 1802 and 1812 (e.g., native graphics drivers) that respectively manage guest PPGTTs and GGTTs 1803 and 1813. The illustrated IOMMU 1830 includes an HLAM module 1831 for implementing the hybrid layer address mapping technology described herein. It should be noted that in this embodiment, a shadowed PPGTT is not present.
[0196] In one embodiment, a GPN-HPN translation table 1833 for the entire guest VM (guest VM 1811 in the example) is prepared in the IOMMU mapping, and each vGPU switch triggers an IOMMU page table swap. That is, when each VM 1801, 1811 is scheduled, its corresponding GPN-HPN translation table 1833 is swapped. In one embodiment, the HLAM module distinguishes between GGTT GPN and PPGTT GPN and modifies the GGTT GPN so that the GGTT GPN does not overlap with the PPGTT GPN when a lookup is performed in the GPN-HPN translation table 1833. Specifically, in one embodiment, virtual GPN generation logic 1832 converts the GGTT GPN into a virtual GPN, which is then used to perform a lookup in the GPN-HPN translation table 1833 to identify the corresponding HPN.
[0197] In one embodiment, a virtual GPN is generated by shifting the GGTT by a specified (potentially large) offset to ensure that the mapped address does not overlap / conflict with the PPGTT GPN. Additionally, in one embodiment, since the CPU can access the GGTT mapped address at any time (e.g., via an opening), the global shadow GGTT will always be valid and remain in the IOMMU mapping of each VM.
[0198] In one embodiment, the HLAM module 1831 solution divides the IOMMU address range into two parts: a lower portion reserved for PPGTT GPN to HPN translation, and an upper portion reserved for GGTT virtual GPN to HPN translation. Since the GPN is provided by the VM / guest 1811, the GPN should be within the range of the guest memory size. In one embodiment, the guest PPGTT page table remains unchanged, and all GPNs from the PPGTT are sent directly to the graphics translation hardware / IOMMU via workload execution. However, in one embodiment, MMIO reads / writes from the guest VM are captured, and GGTT page table changes are captured and modified as described herein (e.g., by adding a large offset to the GPN to ensure it does not overlap with the PPGTT mapping in the IOMMU).
[0199] Remote virtualized graphics processing
[0200] In some embodiments of the invention, the server performs graphics virtualization, virtualizing the physical GPU on behalf of the client and running graphics applications. Figure 19 An embodiment is illustrated in which two clients 1901-1902 connect to server 1930 via network 1910 (such as the Internet and / or a private network). Server 1930 implements a virtualized graphics environment, wherein hypervisor 1960 allocates resources from one or more physical GPUs 1938, presenting the resources as virtual GPUs 1934-1935 to VMs / applications 1932-1933. Graphics processing resources can be allocated according to a resource allocation policy 1961, which allows hypervisor 1960 to allocate resources based on the requirements of applications 1932-1933 (e.g., higher-performance graphics applications require more resources), user accounts associated with applications 1932-1933 (e.g., some users pay extra for higher performance), and / or the current load on the system. The allocated GPU resources may include multiple sets of graphics processing engines, such as 3D engines, block image transfer engines, execution units, and media engines, etc.
[0201] In one embodiment, each user of clients 1901 to 1902 has an account on a service hosted on servers 1930. For example, the service may provide a subscription service to offer users remote access to online applications 1932 to 1933, such as video games, productivity applications, and multiplayer virtual reality applications. In one embodiment, in response to user input 1907 to 1908 from clients 1901 to 1902, the application is remotely executed on a virtual machine. Although not in Figure 19 As shown in the diagram, one or more CPUs can also be virtualized and used to execute applications (1932-1933), where graphics processing operations are offloaded to the vGPU (1934-1935).
[0202] In one embodiment, in response to the execution of graphics operations, vGPU 1934-1935 generates a series of image frames. For example, in a first-person shooter game, a user can specify input 1907 to move a character in a fantasy world. In one embodiment, the resulting image is compressed (e.g., by compression circuitry / logic, not shown) and streamed to clients 1901-1902 via network 1910. In one implementation, a video compression algorithm such as H.261 can be used; however, various different compression techniques can be used. Decoders 1905-1906 decode the input video stream and then render it on the corresponding displays 1903-1904 of clients 1901-1902.
[0203] use Figure 19 The system shown allows high-performance graphics processing resources, such as GPU 1938, to be allocated to different clients of a subscribed service. In an online gaming implementation, for example, server 1930 can host a new video game when it is released. The video game program code is then executed in a virtualized environment, and the resulting video frames are compressed and streamed to each client 1901 to 1902. Clients 1901 to 1902 in this architecture do not require a large amount of graphics processing resources. For example, even relatively low-power smartphones or tablets with decoders 1905 to 1906 will be able to decompress the video stream. Therefore, the latest graphics-intensive video games can be played on any type of client capable of compressing video. While video games are described as one possible implementation, the basic principles of the invention can be used for any form of application that requires graphics processing resources (e.g., graphic design applications, interactive and non-interactive ray-tracing applications, productivity software, video editing software, etc.).
[0204] Devices and methods for managing data bias in graphics processing architectures
[0205] Existing per-cacheline coherence between accelerator devices such as graphics processors and the processor is inefficient and expensive. Server multi-die solutions and multi-socket (e.g., PCIe graphics) make the coherence problem worse because the bandwidth of GPU-attached memory (e.g., HBM or GDDR) is significantly higher, while the listening bandwidth across dies or slots is much lower and the listening latency is high.
[0206] To address these issues, one embodiment of the present invention reduces the overhead of the coherence protocol shared by the CPU / accelerator, making the coherence performance approximately equal to the non-coherence performance. This embodiment maintains the hardware-implemented coherence mechanism for processor / GPU sharing and has no impact on CPU performance for non-shared workloads.
[0207] Specifically, one embodiment implements a bias protocol for naturally aligned sprawling regions (1KB-4KB) of memory, referred to as “super rows,” which are tracked using a “Super Row Ownership Table (SLOT)” or “bias table.” Super rows can be specified to any size, including, for example, a full memory page, 1 / n of memory, or a multiple of a cache line.
[0208] like Figure 20 As shown, in one implementation, GPU 2010 or other accelerator device includes SLOT management circuitry 2013, which specifies the range of superlines that GPU 2010 will process. A copy of this data is then transferred to GPU 2010's native memory 2015 (e.g., local memory), from which GPU 2010 can process the data. SLOT management circuitry 2013 may set bits for each superline to indicate that it has an accelerator / GPU bias. While SLOT management circuitry 2013 is shown as being within GPU 2010, in other embodiments, SLOT management circuitry 2013 may be integrated within CPU complex 2001 or system agent 2021. Alternatively, SLOT management circuitry 2013 may be distributed across GPU 2010, CPU complex 2001, and / or system agent 2021.
[0209] In one implementation, SLOT 2031 tracks who "owns" a given superrow at any given time. For example... Figure 20As shown, the complete SLOT table 2031 can be stored in memory 2030 (e.g., system memory), while cached versions 2011 and 2022 are stored on the die, within GPU 2010 and system agent 2021, respectively. In one embodiment, cached versions 2011 and 2022 may include a subset of the entire SLOT table 2031. Ownership transfer can be hardware-driven, software-initiated, or a combination of both.
[0210] In one embodiment, when the GPU 2010 needs to process a large block of data, it pulls the data into its native memory 2015 (e.g., HBM, GDDR memory, etc.) and updates the SLOT table 2031 to indicate that it is the current "owner" of the data block. For example, a bit can be set for each super row of the block to indicate that it is being processed by the GPU 2010. Typically, this would be a portion of the data that the CPU complex 2001 does not need to access frequently. When performing graphics operations, this portion of the data can be cached in the GPU's L3 cache 2011.
[0211] In one embodiment, GPU 2010 includes a native SLOT cache 2012 to store recently accessed portions of the main SLOT table 2031. Additionally, system agent 2021 may include a slot cache 2022 to provide efficient access to SLOT data for multiple cores of CPU complex 2001. A SLOT cache (not shown) may also be maintained on the CPU. Cache coherence circuitry ensures coherence between SLOT cache 2022 and SLOT cache 2012. Any form of cache coherence protocol may be used.
[0212] In one embodiment, the SLOT only covers HBM memory (e.g., 8GB) or GDDR memory directly attached to the accelerator or GPU. The entire SLOT table can be internal to the GPU for simplified implementation. Memory not connected to the accelerator (e.g., DDR) is not covered by the SLOT.
[0213] In one embodiment, CPU composite 2001 does not have a SLOT table and will not track any super-row based consistency. Only the accelerator or GPU 2010 will have a SLOT table 2031 and track ownership of native HBM or GDDR memory. In one implementation, all CPU composite 2001 requests for native memory 2015 will listen to the GPU SLOT table 2031 to check ownership. New consistency commands are used in this system to accelerate super-row ownership transfer. For example, a "super-row dump clear" command can be used to clear an entire super-row dump from CPU cache 2002 or GPU SLOT cache 2012.
[0214] One implementation leverages directly attached memory, such as stacked DRAM or HBM, to improve GPU performance and simplify application development for applications utilizing GPUs with directly attached memory. This implementation allows the GPU-attached memory to be mapped as a portion of system memory and accessed using shared virtual memory (SVM) techniques, such as those used in current IOMMU implementations, without suffering the typical performance drawbacks associated with full system cache coherency.
[0215] The ability to access GPU-attached memory as part of system memory without incurring heavy cache coherence overhead provides a beneficial operating environment for GPU offload. This ability to access memory as part of the system address map allows host software to build operands and access computation results without the overhead of traditional IO DMA data copying. Such traditional copying involves driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory access. Simultaneously, the ability to access GPU-attached memory without cache coherence overhead is critical to the execution time of offloaded computations. For example, in scenarios with heavy streaming write memory traffic, cache coherence overhead can halve the effective write bandwidth seen by the GPU. The efficiency of operand building, the efficiency of result access, and the efficiency of GPU computation all play a role in determining how effective GPU offload will be. If the cost of offload work (e.g., building operands, obtaining results) is too high, offload may not yield good results or may limit the GPU to only very large tasks. The efficiency of GPU computation execution will have the same effect.
[0216] An implementation applies different memory access and coherence techniques based on the entity initiating the memory access (e.g., GPU, core, etc.) and the accessed memory (e.g., host memory or GPU memory). These techniques are collectively referred to as a "coherence bias" mechanism, which provides two sets of cache coherence processes for GPU-attached memory. One set is optimized for efficient GPU access to its attached memory, and the second set is optimized for host access to GPU-attached memory and shared GPU / host access to GPU-attached memory. Furthermore, it includes two techniques for switching between these processes: one driven by application software and the other driven by autonomous hardware hints. In both sets of coherence processes, the hardware maintains full cache coherence.
[0217] like Figure 21As generally illustrated, an implementation applies to a computer system including a GPU 2010 and one or more computer processor chips having a processor (cores and I / O) 2103, wherein the GPU 2010 is coupled to the processor via a multi-protocol link 2110. In one implementation, the multi-protocol link 2110 is a dynamically multiplexed link that supports a variety of different protocols, including but not limited to the System on Chip (IOSF) protocol and / or the PCI High Speed (PCIe) protocol (i.e., for supporting producer / consumer traffic, device discovery and configuration, and interrupts), proxy coherence protocols such as Interface on Die (IDI), and memory access protocols such as System Memory Interface Protocol 3 (e.g., SMI3+). However, it should be noted that the basic principles of the invention are not limited to any particular set of protocols. Furthermore, it should be noted that, depending on the implementation, the GPU 2010 and the cores / I / O may be integrated on the same semiconductor chip or different semiconductor chips.
[0218] In the illustrated implementation, GPU memory bus 2112 couples GPU 2010 to GPU memory 2115, and a separate host memory bus 2111 couples core / I / O to host memory 2130. As mentioned, GPU memory 2115 may include high-bandwidth memory (HBM) or stacked DRAM (some examples of which are described in this application), and host memory 2130 may include DRAM such as dual data rate synchronous dynamic random access memory (e.g., DDR3 SDRAM, DDR4 SDRAM, etc.). However, the basic principles of the invention are not limited to any particular type of memory or memory protocol.
[0219] In one implementation, the GPU 2010 and the "host" software running on the processing core within the processor chip use two different sets of protocol flows (referred to as the "host bias" flow and the "GPU bias" flow) to access the GPU memory 2115. As described below, one implementation supports multiple options for modulating and / or selecting the protocol flow for a particular memory access.
[0220] The consistency biasing process is implemented in part at two protocol layers on the multi-protocol link 2110 between the GPU 2010 and one of multiple processor chips: the IDI protocol layer and the SMI3 protocol layer. In one implementation, the consistency biasing process is enabled by: (a) using existing opcodes in the IDI protocol in a new way, (b) adding new opcodes to the existing SMI3 standard, and (c) adding support for the SMI3 protocol to the multi-protocol link 2010 (existing links such as the R link that only include IDI and IOSF). Note that the multi-protocol link is not limited to supporting only IDI and SMI3; in one implementation, only support for at least those protocols is required.
[0221] As used in this application, Figure 21 The “host bias” process illustrated is a set of processes that centralizes all requests to GPU memory 2115, including those from the GPU itself, through a standard coherence controller 2109 in the processor chip to which GPU 2010 is attached. This causes GPU 2010 to take a roundabout path to access its own memory, but allows the use of the processor’s standard coherence controller 2109 to maintain consistency between accesses from both GPU 2010 and processor core / IO. In one implementation, this process uses standard IDI opcodes, via a multiprotocol link, to issue requests to the processor’s coherence controller 2109 in the same or similar manner as the processor cores issue requests to the coherence controller 2109. For example, the processor chip’s coherence controller 2109 may, on behalf of the GPU, issue UPI and IDI coherence messages (e.g., listeners) as a result of requests from GPU 2010 to all peer processor core chips (e.g., processor (cores and I / O) 2103) and system agent 2021, just as they would to requests from processor cores. In this way, consistency is maintained between data accessed by the GPU 2010 and the processor core / IO.
[0222] In one implementation, the coherence controller 2109 also conditionally publishes memory access messages to the GPU's memory controller 2106 via the multiprotocol link 2110. These messages are similar to those sent by the coherence controller 2109 to its processor die's native memory controller and include new opcodes that allow data to be returned directly to the agent inside the GPU 2010 (instead of forcing the coherence controller 2109 to return data to the processor via the multiprotocol link 2110), and are then returned to the GPU 2010 as an IDI response via the multiprotocol link 2110.
[0223] exist Figure 21 In one implementation of the “host-biased” mode shown, all requests from the processor core targeting GPU-attached memory (e.g., GPU memory 2115) are sent directly to the processor coherence controller 2109, just as if they were targeting ordinary host memory 2130. The coherence controller 2109 may apply its standard cache coherence algorithm and send its standard cache coherence messages, just as it would for accesses from the GPU 2010, and just as it would for accesses to ordinary host memory 2130. For such requests, the coherence controller 2109 may also conditionally send an SMI3 command via the multiprotocol link 2110; however, in this case, the SMI3 process returns data across the multiprotocol link 2110.
[0224] Figure 22 The “GPU bias” procedures shown are those that allow GPU 2010 to access its natively attached GPU memory 2115 without consulting the host memory’s cache coherency controller. More specifically, these procedures allow GPU 2010 to access its natively attached memory via memory controller 2106 without sending requests via multiprotocol link 2110.
[0225] In "GPU bias" mode, requests from the processor core / IO may be issued as described above for "host bias," but they are fulfilled as if they were issued as "non-cached" requests. This "non-cached" convention ensures that data conforming to the GPU bias process is never cached in the processor's cache hierarchy (e.g., 2011). This fact allows the GPU 2010 to access GPU bias data in its GPU memory 2115 without consulting the cache coherence controller 2109 on the processor.
[0226] In one implementation, support for the "non-cached" processor core access procedure is achieved using a globally observed one-use ("GO-UO") response on the processor bus (e.g., the IDI bus in some embodiments). This response returns a piece of data to the processor core and instructs the processor to use the value of that data only once. This prevents caching of the data and satisfies the requirement of the "non-cached" procedure. In systems with cores that do not support the GO-UO response, the "non-cached" procedure can be implemented using a sequence of multiple message responses on the memory protocol layer (e.g., the SMI3 layer) of the multi-protocol link 2110 and on the processor core bus.
[0227] Specifically, when a processor core is found to be targeting a "GPU bias" page at GPU 2010, the GPU establishes certain states to block future requests for the target cache line from the GPU and sends a special "GPU bias hit" response (e.g., an SMI3 message) on the multiprotocol link 2110. In response to this message, the processor's cache coherence controller 2109 returns data to the requesting processor core, followed by a listen-invalid message. When the processor core acknowledges the listen-invalidity as complete, the cache coherence controller 2109 sends another special "GPU bias block complete" message (e.g., another SMI3 message in one embodiment) back to GPU 2010 at the SMI3 layer of the multiprotocol link 2110. This complete message causes GPU 2010 to clear the aforementioned blocking state.
[0228] As mentioned above, the choice between GPU and host biasing processes can be driven by a biasing tracking data structure such as SLOT table 2307 in GPU memory 2115. This SLOT table 2307 can be a super-row granularity structure containing 1 or 2 bits of each memory page attached to the GPU (i.e., controlled at the granularity of memory pages, 1 / n memory pages, some cache line multiple, etc.). SLOT tables 2307 and 2331 can be implemented in the memory range of GPU-attached memory and / or in host memory (e.g., relative to...). Figure 21 (As described). Bias cache 2303 in the GPU (e.g., for frequently / recently used entries in cache SLOT table 2307). Alternatively, the entire SLOT table 2307 can be maintained within the GPU 2010.
[0229] In one implementation, prior to the actual access to GPU memory 2115, the SLOT table entry associated with each access to memory attached to the GPU is accessed, resulting in the following operation:
[0230] • In GPU bias, it was found that the native requests from GPU 2010 in its super row were directly forwarded to GPU memory 2115.
[0231] • Native requests from GPU 2010 that discover its super row in host bias are forwarded to the processor via multiprotocol link 2110 (e.g., as an IDI request in one embodiment).
[0232] • Requests from the processor (such as SMI3 requests) that are found in the GPU bias to be super-rows are completed using the “non-cached” process described above.
[0233] • The SMI3 request from the processor, which is found in the host bias, is completed like a normal memory read.
[0234] The bias state of the superrow can be changed through software-based mechanisms, hardware-assisted software-based mechanisms, or, for a limited set of cases, purely hardware-based mechanisms.
[0235] One mechanism for changing the bias state employs API calls (such as OpenCL), which invoke the GPU's device driver. This device driver then sends a message to the GPU 2010 (or enqueues a command descriptor) to instruct it to change the bias state and perform a cache dump clearing operation on the host for some transitions. The cache dump clearing operation is necessary for transitions from host bias to GPU bias, but not for the reverse transition.
[0236] In some cases, software may struggle to determine when to make a bias conversion API call and identify the superline requiring bias conversion. In these situations, the GPU can implement a bias conversion "hint" mechanism, whereby the GPU detects the need for bias conversion and sends a message to its driver indicating that need. This hint mechanism can be as simple as a lookup mechanism in response to SLOT table 2331, 2307, which is triggered when the GPU accesses the host bias superline or the host accesses the GPU bias superline and signals the event to the GPU driver via an interrupt.
[0237] Note that some implementations may require a second bias state bit to enable the bias transition state value. This allows the system to continue accessing those super lines while they are in the process of bias changing (i.e., when the cache is partially dumped and cleared, and Δ cache pollution caused by subsequent requests must be suppressed).
[0238] Figure 24 An exemplary process according to one embodiment is illustrated. This process can be implemented on the system and processor architecture described in this application, but is not limited to any particular system or processor architecture.
[0239] In 2401, a specific set of superrows is placed under GPU bias. As mentioned, this can be done by updating the entries for these superrows in the SLOT table to indicate that they are under GPU bias (e.g., by setting the bit associated with each superrow). In one implementation, once set to GPU bias, it is guaranteed that these superrows will not be cached in the host cache memory. In 2402, superrows are allocated from GPU memory (e.g., software allocates superrows by initiating a driver / API call).
[0240] At 2403, the operands are pushed to the allocated superline from the processor core. In one implementation, this can be achieved by software using API calls to flip the operand superline to the host bias (e.g., via OpenCL API calls). No data copying or cache dump clearing is required, and the operand data can terminate at some arbitration location in the host cache hierarchy at this stage.
[0241] In 2404, the GPU uses operands to produce results. For example, it can execute commands and process data directly from its native memory (e.g., native memory 2015 discussed above). In one implementation, the software uses the OpenCL API to flip operand superlines back to the GPU bias (e.g., update the bias table). As a result of the API call, a command / work descriptor is submitted to the GPU (e.g., via a shared queue on a dedicated command queue). The work descriptor / command can instruct the GPU to dump and clear operand superlines from the host cache, resulting in a cache dump clear (e.g., executed using CLFLUSH over the IDI protocol). In one implementation, the GPU executes without host-dependent consistency overhead and dumps data to the result superline.
[0242] At 2405, the result is pulled from the allocated superrow. For example, in one implementation, the software makes one or more API calls (e.g., via the OpenCL API) to flip the resulting superrow to the host bias. This action may cause some bias state to change, but does not cause any consistency or cache dump clearing actions. The host processor core can then access, cache, and share the resulting data as needed. Finally, at 2406, the allocated superrow is released (e.g., via software).
[0243] One or more of the above embodiments are particularly suitable for server implementations. Two different client-side implementations are also conceivable, including symmetric and asymmetric implementations.
[0244] In the symmetric implementation, both CPU and GPU accesses to the entire system memory space undergo SLOT checks. Figure 25 One specific implementation is shown in which a SLOT table 2551 is maintained in system memory 2550 and checked in response to access from cores 2501-2504 and GPU 2510. In this implementation, multiple layers of SLOT caching are maintained, including a slot cache 2521 within the core listener filter 2520, a SLOT cache 2512 within the GPU 2510, and a low-level SLOT cache 2540 that can store entries for both cores 2501-2504 and GPU 2510. In one embodiment, the low-level SLOT cache 2540 may be included as part of an LLC 2530 architecture. Access to the low-level SLOT cache 2540 may be provided to I / O device 2525 to participate in the biasing mechanism described herein. In one embodiment, a memory cache 2545 (e.g., HBM or high-speed DRAM memory located in front of main system memory 2550, which may utilize relatively lower-speed and lower-cost memory technologies) may also be used.
[0245] One embodiment of the asymmetric implementation may utilize partitioned memory. Figure 26 An exemplary partition of memory 2655 is shown, comprising a GPU-consistent partition 2651 and a system memory partition 2650. One objective of this embodiment is to simplify changes required for non-kernel applications. The system memory partition 2650 and the GPU-consistent partition 2651 can be exposed to the OS as Non-United Memory Access (NUMA) memory. In one embodiment, a driver allocates the entire GPU-consistent partition 2651 under driver load. Applications require driver-managed allocation of the GPU-consistent partition 2651.
[0246] In one implementation, SLOT is applied only to the GPU-consistent partition 2651. Ownership can be tracked using the SLOT cache 2512. The bias setting is indifferent to CPU and I / O access to that partition. Misses in all CPU final level cache (LLC) 2610 caches for the "GPU-consistent" space will need to be monitored via the system agent 2630. This implementation allows graphics data to be cached as "non-consistent" using LLC.
[0247] In one implementation, the graphics response to a kernel sniffing of a "GPU-consistent" partition is as follows. The GPU 2510 looks up the superrow in its bias table, first in the SLOT cache 2512, and then in the main SLOT table 2652 stored in memory 2655. If the GPU owns the superrow, it has free access to it without needing to cooperate with the CPU. If the superrow is not owned by the GPU, the sniffing response is sent normally to the CPU sniffing filter 2610.
[0248] If a superline is marked as “GPU biased” in SLOT 2652, a single-core listener for the superline can automatically trigger the “release” of the GPU bias for that superline. In one embodiment, the superline is first marked as “in transition to CPU bias”. New coherence requests for that superline from the GPU are then blocked. The GPU can send a listener response to the CPU, dump and clear all cache lines in that superline, and then remove the superline entry from the SLOT table. The blocking of new requests for that superline can then be lifted, and the transition back to GPU bias can then begin if necessary. Hooks for applications / drivers can be implemented to request the “acquisition” or “release” of the superline (e.g., software prefetching). These can be considered hints and may not be entirely dependent on this interface.
[0249] In some embodiments, a graphics processing unit (GPU) is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In other embodiments, the GPU may be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of how the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a job descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0250] In the following description, numerous specific details are set forth to provide a more thorough explanation. However, it will be apparent to those skilled in the art that the embodiments described herein can be practiced without one or more of these specific details. In other instances, well-known features have not been described to avoid obscuring the details of the embodiments of the invention.
[0251] System Overview
[0252] Figure 27 This is a block diagram illustrating a computer system 2700 configured to implement one or more aspects of the embodiments described herein. The computing system 2700 includes a processing subsystem 2701 having one or more processors 2702 and a system memory 2704, the one or more processors and the system memory communicating via an interconnect path, the interconnect path including a memory hub 2705. The memory hub 2705 may be a separate component within a chipset assembly or integrated within one or more processors 2702. The memory hub 2705 is coupled to an I / O subsystem 2711 via a communication link 2706. The I / O subsystem 2711 includes an I / O hub 2707 that enables the computing system 2700 to receive input from one or more input devices 2708. Additionally, the I / O hub 2707 enables a display controller (which may be included in one or more processors 2702) to provide output to one or more display devices 2710A. In one embodiment, one or more display devices 2710A coupled to the I / O hub 2707 may include native display devices, internal display devices, or embedded display devices.
[0253] In one embodiment, the processing subsystem 2701 includes one or more parallel processors 2712 coupled to a memory hub 2705 via a bus or other communication link 2713. The communication link 2713 can be one of any number of standards-based communication link technologies or protocols (such as, but not limited to, PCI Express), or a vendor-specific communication interface or communication architecture. In one embodiment, the one or more parallel processors 2712 form a computation-centric parallel or vector processing system including a large number of processing cores and / or processing clusters such as integrated many-core (MIC) processors. In one embodiment, the one or more parallel processors 2712 form a graphics processing subsystem that can output pixels to one of one or more display devices 2710A coupled via an I / O hub 2707. The one or more parallel processors 2712 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 2710B.
[0254] Within the I / O subsystem 2711, system storage unit 2714 can be connected to I / O hub 2707 to provide storage for computing system 2700. I / O switch 2716 can be used to provide an interface mechanism to enable connection between I / O hub 2707 and other components that can be integrated into the platform, such as network adapter 2718 and / or wireless network adapter 2719, as well as various other devices that can be added via one or more plug-in devices 2720. Network adapter 2718 can be an Ethernet adapter or another wired network adapter. Wireless network adapter 2719 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radio components.
[0255] The computing system 2700 may include other components not explicitly shown, such as USB or other port connectors, optical storage drivers, video capture devices, etc., and may also be connected to the I / O hub 2707. Figure 27 The communication paths for interconnecting various components can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or (multiple) other bus or point-to-point communication interfaces and / or protocols such as NV-Link high-speed interconnect or interconnect protocols known in the art.
[0256] In one embodiment, one or more parallel processors 2712 are incorporated with circuitry optimized for graphics and video processing, including, for example, video output circuitry, and said circuitry constitutes a graphics processing unit (GPU). In another embodiment, one or more parallel processors 2712 are incorporated with circuitry optimized for general-purpose processing while retaining the underlying computing architecture described in more detail herein. In yet another embodiment, components of the computing system 2700 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 2712, a memory hub 2705, processor(s)2702(s), and an I / O hub 2707 may be integrated into a system-on-a-chip (SoC) integrated circuit. Alternatively, components of the computing system 2700 may be integrated into a single package to form a system-in-package (SIP) configuration. In other embodiments, at least a portion of the components of the computing system 2700 may be integrated into a multi-chip module (MCM), which may interconnect with other MCMs to form a modular computing system.
[0257] It should be understood that the computing system 2700 shown herein is exemplary and variations and modifications are possible. The connection topology can be modified as needed, including the number and arrangement of bridges, the number of processors(multiple) 2702, and the number of parallel processors(multiple) 2712. For example, in some embodiments, system memory 2704 is connected directly to processors(multiple) 2702 instead of via bridges, while other devices communicate with system memory 2704 via memory hub 2705 and processors(multiple) 2702. In other alternative topologies, parallel processors(multiple) 2712 are connected to I / O hub 2707 or directly to one or more processors 2702, instead of to memory hub 2705. In other embodiments, I / O hub 2707 and memory hub 2705 may be integrated into a single chip. Some embodiments may include two or more groups of processors(multiple) 2702 attached via multiple sockets, which may be coupled to two or more instances of parallel processors(multiple) 2712.
[0258] Some specific components shown herein are optional and may not be included in all embodiments of the computing system 2700. For example, any number of plug-in cards or peripheral devices may be supported, or some components may be omitted. Furthermore, some architectures may be described using different terminology. Figure 27 Similar components are shown. For example, in some architectures, the memory hub 2705 may be referred to as the Northbridge, while the I / O hub 2707 may be referred to as the Southbridge.
[0259] Figure 28AA parallel processor 2800 according to an embodiment is illustrated. Various components of the parallel processor 2800 can be implemented using one or more integrated circuit devices such as a programmable processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). According to an embodiment, the illustrated parallel processor 2800 is... Figure 27 The above are variations of one or more parallel processors 2712.
[0260] In one embodiment, the parallel processor 2800 includes a parallel processing unit 2802. The parallel processing unit includes an I / O unit 2804 that enables communication with other devices, including other instances of the parallel processing unit 2802. The I / O unit 2804 may be directly connected to other devices. In one embodiment, the I / O unit 2804 is connected to other devices via a hub or switch interface, such as a memory hub 2705. The connection between the memory hub 2705 and the I / O unit 2804 forms a communication link 2713. Within the parallel processing unit 2802, the I / O unit 2804 is connected to a host interface 2806 and a memory crossbar switch 2816, wherein the host interface 2806 receives commands relating to performing processing operations, and the memory crossbar switch 2816 receives commands relating to performing memory operations.
[0261] When host interface 2806 receives a command buffer via I / O unit 2804, host interface 2806 can route work operations for executing these commands to front end 2808. In one embodiment, front end 2808 is coupled to scheduler 2810, which is configured to distribute commands or other work items to processing cluster array 2812. In one embodiment, scheduler 2810 ensures that processing cluster array 2812 is correctly configured and in an active state before distributing tasks to the processing clusters of processing cluster array 2812. In one embodiment, scheduler 2810 is implemented via firmware logic executed on a microcontroller. The microcontroller-implemented scheduler 2810 can be configured to perform complex scheduling and work assignment operations at a coarse-grained level, thereby enabling fast preemption and context switching of threads executing on processing array 2812. In one embodiment, host software can validate workloads for scheduling on processing array 2812 via one of a plurality of graphics processing doorbells. The workloads can then be automatically distributed across processing array 2812 by scheduler 2810 logic within the scheduler microcontroller.
[0262] The processing cluster array 2812 may include up to "N" processing clusters (e.g., cluster 2814A, cluster 2814B, up to cluster 2814N). Each cluster 2814A through 2814N of the processing cluster array 2812 can execute a large number of concurrent threads. The scheduler 2810 may use various scheduling and / or work distribution algorithms to allocate work to the clusters 2814A through 2814N of the processing cluster array 2812, and these algorithms may vary depending on the workload caused by each type of program or computation. Scheduling may be handled dynamically by the scheduler 2810, or it may be partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 2812. In one embodiment, the different clusters 2814A through 2814N of the processing cluster array 2812 may be assigned to process different types of programs or to perform different types of computations.
[0263] The processing cluster array 2812 can be configured to perform various types of parallel processing operations. In one embodiment, the processing cluster array 2812 is configured to perform general-purpose parallel computing operations. For example, the processing cluster array 2812 may include logic for performing processing tasks including filtering video and / or audio data, performing modeling operations including physical operations, and performing data transformations.
[0264] In one embodiment, the processing cluster array 2812 is configured to perform parallel graphics processing operations. In embodiments where the parallel processor 2800 is configured to perform graphics processing operations, the processing cluster array 2812 may include additional logic for supporting the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. Additionally, the processing cluster array 2812 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 2802 may transfer data from system memory via I / O unit 2804 for processing. During processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 2822) and then written back to system memory.
[0265] In one embodiment, when the parallel processing unit 2802 is used to perform graphics processing, the scheduler 2810 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 2814A to 2814N of the processing cluster array 2812. In some embodiments, portions of the processing cluster array 2812 can be configured to perform different types of processing. For example, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to produce a rendered image for display. Intermediate data generated by one or more of the clusters 2814A to 2814N can be stored in a buffer to allow intermediate data to be transferred between the clusters 2814A to 2814N for further processing.
[0266] During operation, the processing cluster array 2812 can receive processing tasks to be executed via scheduler 2810, which receives commands defining the processing tasks from front-end 2808. For graphics processing operations, processing tasks may include data to be processed, such as surface (patch) data, graph data, vertex data, and / or pixel data, as well as state parameters defining how the data is processed and indices of commands (e.g., which program to execute). Scheduler 2810 can be configured to retrieve indices corresponding to tasks or may receive indices from front-end 2808. Front-end 2808 can be configured to ensure that the processing cluster array 2812 is configured to be active before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.
[0267] Each of one or more instances of the parallel processing unit 2802 may be coupled to the parallel processor memory 2822. The parallel processor memory 2822 may be accessed via a memory crossbar switch 2816, which receives memory requests from the processing cluster array 2812 and the I / O unit 2804. The memory crossbar switch 2816 may access the parallel processor memory 2822 via a memory interface 2818. The memory interface 2818 may include a plurality of partition units (e.g., partition units 2820A, 2820B, up to partition units 2820N), each of which may be coupled to a portion (e.g., a memory cell) of the parallel processor memory 2822. In one embodiment, the number of partition units 2820A to 2820N is configured to be equal to the number of memory cells, such that a first partition unit 2820A has a corresponding first memory cell 2824A, a second partition unit 2820B has a corresponding memory cell 2824B, and an Nth partition unit 2820N has a corresponding Nth memory cell 2824N. In other embodiments, the number of partition units 2820A to 2820N may not be equal to the number of memory devices.
[0268] In various embodiments, memory cells 2824A to 2824N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In one embodiment, memory cells 2824A to 2824N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). Those skilled in the art will understand that the specific implementation of memory cells 2824A to 2824N can vary and can be selected from one of a variety of conventional designs. Render targets, such as frame buffers or texture maps, may be stored on memory cells 2824A to 2824N, thereby allowing partitioning cells 2820A to 2820N to write portions of each render target in parallel to efficiently utilize the available bandwidth of parallel processor memory 2822. In some embodiments, to support a unified memory design utilizing system memory along with native cache memory, native instances of parallel processor memory 2822 may be excluded.
[0269] In one embodiment, any of clusters 2814A to 2814N of the processing cluster array 2812 can process data to be written to any of the memory cells 2824A to 2824N within the parallel processor memory 2822. The memory crossbar switch 2816 can be configured to send the output of each cluster 2814A to 2814N to any partition cell 2820A to 2820N or another cluster 2814A to 2814N, which can perform additional processing operations on the output. Each cluster 2814A to 2814N can communicate with the memory interface 2818 via the memory crossbar switch 2816 to perform read or write operations for various external memory devices. In one embodiment, the memory crossbar switch 2816 may be connected to the memory interface 2818 to communicate with the I / O unit 2804, and may be connected to a native instance of the parallel processor memory 2822, thereby enabling processing units within different processing clusters 2814A to 2814N to communicate with system memory or other memory that is not native to the parallel processing unit 2802. In one embodiment, the memory crossbar switch 2816 may use virtual channels to separate traffic flows between clusters 2814A to 2814N and partition units 2820A to 2820N.
[0270] While a single instance of the parallel processing unit 2802 is shown within the parallel processor 2800, any number of instances of the parallel processing unit 2802 can also be included. For example, multiple instances of the parallel processing unit 2802 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. Even if different instances have different numbers of processing cores, different amounts of native parallel processor memory, and / or other configuration differences, different instances of the parallel processing unit 2802 can be configured to operate interactively. For example, and in one embodiment, some instances of the parallel processing unit 2802 may include higher precision floating-point units relative to other instances. Systems incorporating one or more instances of the parallel processing unit 2802 or the parallel processor 2800 can be implemented in various configurations and form factors, including but not limited to desktop computers, laptop or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0271] Figure 28B This is a block diagram of partitioning unit 2820 according to an embodiment. In one embodiment, partitioning unit 2820 is... Figure 28AAn example of one of partitioning units 2820A to 2820N. As shown, partitioning unit 2820 includes an L2 cache 2821, a frame buffer interface 2825, and a ROP 2826 (raster operation unit). The L2 cache 2821 is a read / write cache configured to perform load and store operations received from memory crossbar switches 2816 and ROP 2826. Read miss and urgent write-back requests are output from the L2 cache 2821 to the frame buffer interface 2825 for processing. Updates can also be sent to the frame buffer via the frame buffer interface 2825 for processing. In one embodiment, the frame buffer interface 2825 interfaces with one of the memory cells in the parallel processor memory, such as... Figure 28A Interacting with memory cells 2824A to 2824N (e.g., within parallel processor memory 2822).
[0272] In graphics applications, the ROP 2826 is a processing unit that performs raster operations such as stencil printing, z-testing, and blending. The ROP 2826 then outputs the processed graphics data stored in graphics memory. In some embodiments, the ROP 2826 includes compression logic for compressing depth or color data written to memory and decompressing depth or color data read from memory. The compression logic may be lossless compression logic using one or more of a variety of compression algorithms. The type of compression performed by the ROP 2826 can vary depending on the statistical characteristics of the data to be compressed. For example, in one embodiment, Δ color compression is performed on depth and color data on a per-tile basis.
[0273] In some embodiments, the ROP 2826 is included in each processing cluster (e.g., Figure 28A The data is stored within clusters 2814A to 2814N instead of partition units 2820. In this embodiment, read and write requests for pixel data are transmitted via a memory crossbar switch 2816 instead of pixel fragment data. The processed graphics data can be displayed on a display device such as... Figure 27 On one or more display devices 2710, routed by processor(s) 2702 for further processing, or by... Figure 28A One of the processing entities within the parallel processor 2800 is routed for further processing.
[0274] Figure 28C This is a block diagram of a processing cluster 2814 within a parallel processing unit according to an embodiment. In one embodiment, the processing cluster is... Figure 28AAn instance of one of the processing clusters 2814A to 2814N. Processing cluster 2814 can be configured to execute multiple threads in parallel, where the term "thread" refers to an instance of a specific program executing on a specific input dataset. In some embodiments, Single Instruction Multiple Data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, Single Instruction Multiple Threading (SIMT) technology is used to support the parallel execution of a large number of substantially synchronous threads using a common instruction unit configured to issue instructions to a set of processing engines within each of the processing cluster. Unlike SIMD execution mechanisms, where all processing engines typically execute the same instructions, SIMT execution allows different threads to more easily follow divergent execution paths through a given thread program. Those skilled in the art will understand that SIMD processing mechanisms represent a subset of the functionality of SIMT processing mechanisms.
[0275] The operation of the processing cluster 2814 can be controlled via the pipeline manager 2832, which distributes processing tasks to the SIMT parallel processors. The pipeline manager 2832... Figure 28A The scheduler 2810 receives instructions and manages the execution of those instructions via the graphics multiprocessor 2834 and / or texture unit 2836. The graphics multiprocessor 2834 shown is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors with different architectures can be included within the processing cluster 2814. One or more instances of the graphics multiprocessor 2834 can be included within the processing cluster 2814. The graphics multiprocessor 2834 can process data, and the data cross switch 2840 can be used to distribute the processed data to one of a plurality of possible destinations, including other shading units. The pipeline manager 2832 can facilitate the distribution of processed data by specifying destinations for the data to be distributed via the data cross switch 2840.
[0276] Each graphics multiprocessor 2834 within the processing cluster 2814 may include the same group of functional execution logic (e.g., arithmetic logic units, load-memory units, etc.). The functional execution logic can be configured in a pipelined manner, where new instructions can be issued before completing previous instructions. The functional execution logic supports various operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and calculations of various algebraic functions. In one embodiment, the same functional unit hardware can be used to perform different operations, and any combination of functional units can exist.
[0277] Instructions transmitted to the processing cluster 2814 constitute threads. A group of threads executing on a set of parallel processing engines is a thread group. Thread groups execute the same program on different input data. Each thread within a thread group can be assigned to a different processing engine within the graphics multiprocessor 2834. A thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 2834. When a thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycle of processing the thread group. A thread group may also include more threads than the number of processing engines within the graphics multiprocessor 2834. When a thread group includes more threads than the number of processing engines within the graphics multiprocessor 2834, processing can be performed on consecutive clock cycles. In one embodiment, multiple thread groups can be executed simultaneously on the graphics multiprocessor 2834.
[0278] In one embodiment, the graphics multiprocessor 2834 includes an internal cache memory for performing load and store operations. In one embodiment, the graphics multiprocessor 2834 may forgo the internal cache and instead use a cache memory (e.g., L1 cache 308) within the processing cluster 2814. Each graphics multiprocessor 2834 may also access partition units shared across all processing clusters 2814 (e.g., ...). Figure 28A The graphics multiprocessor 2834 has an L2 cache within partition units 2820A to 2820N, which can be used to transfer data between threads. The graphics multiprocessor 2834 can also access off-chip global memory, which may include one or more of native parallel processor memory and / or system memory. Any memory outside of the parallel processing unit 2802 can be used as global memory. Embodiments where the processing cluster 2814 includes multiple instances of the graphics multiprocessor 2834 can share common instructions and data that can be stored in the L1 cache 308.
[0279] Each processing cluster 2814 may include an MMU 2845 (Memory Management Unit) configured to map virtual addresses to physical addresses. In other embodiments, one or more instances of the MMU 2845 may reside in Figure 28A The memory interface 2818 is located within the MMU 2845. The MMU 2845 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses (more specifically, chunks) and optionally cache line indices. The MMU 2845 may include an address translation lookahead buffer (TLB) or cache that may reside within the graphics multiprocessor 2834 or the L1 cache or processing cluster 2814. Physical addresses are processed to distribute surface data access locality to achieve efficient request interleaving between partition units. Cache line indices can be used to determine whether a request for a cache line is a hit or a miss.
[0280] In graphics and computing applications, processing cluster 2814 can be configured such that each graphics multiprocessor 2834 is coupled to texture unit 2836 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. Texture data is read from an internal texture L1 cache (not shown) or, in some embodiments, from an L1 cache within the graphics multiprocessor 2834, and is retrieved as needed from an L2 cache, native parallel processor memory, or system memory. Each graphics multiprocessor 2834 outputs a processed task to data cross switch 2840 to provide the processed task to another processing cluster 2814 for further processing or to store the processed task in an L2 cache, native parallel processor memory, or system memory via memory cross switch 2816. preROP 2842 (pre-raster operation unit) is configured to receive data from graphics multiprocessors 2834 and direct the data to ROP units, which can be partitioned as described herein (e.g., Figure 28A The preROP 2842 unit is located in partitioning units 2820A to 2820N. The preROP 2842 unit can optimize color mixing, organize pixel color data, and perform address translation.
[0281] It should be understood that the core architecture described herein is exemplary and variations and modifications are possible. For example, any number of processing units such as a graphics multiprocessor 2834, a texture unit 2836, and a preROP 2842 can be included within the processing cluster 2814. Furthermore, although only one processing cluster 2814 is shown, the parallel processing units as described herein can include any number of instances of the processing cluster 2814. In one embodiment, each processing cluster 2814 can be configured to operate independently of other processing clusters 2814 using separate and different processing units, L1 caches, etc.
[0282] Figure 28D A graphics multiprocessor 2834 according to one embodiment is illustrated. In such an embodiment, the graphics multiprocessor 2834 is coupled to a pipeline manager 2832 of a processing cluster 2814. The graphics multiprocessor 2834 has an execution pipeline including, but not limited to, an instruction cache 2852, an instruction unit 2854, an address mapping unit 2856, a register file 2858, one or more general-purpose graphics processing unit (GPGPU) cores 2862, and one or more load / store units 2866. The GPGPU cores 2862 and the load / store units 2866 are coupled to a cache memory 2872 and a shared memory 2870 via a memory and cache interconnect 2868.
[0283] In one embodiment, instruction cache 2852 receives a stream of instructions to be executed from pipeline manager 2832. These instructions are cached in instruction cache 2852 and dispatched for execution by instruction unit 2854. Instruction unit 2854 can dispatch instructions as thread groups (e.g., threads), with each thread in the thread group assigned to a different execution unit within GPGPU core 2862. Instructions can access any of the native, shared, or global address spaces by specifying an address within a unified address space. Address mapping unit 2856 can be used to translate addresses in the unified address space into different memory addresses accessible by load / store unit 2866.
[0284] The register file 2858 provides a set of registers for the functional units of the graphics multiprocessor 2824.
[0285] Register file 2858 provides temporary storage for operands on data paths connected to functional units (e.g., GPGPU core 2862, load / store unit 2866) of graphics multiprocessor 2824. In one embodiment, register file 2858 is partitioned among each of the functional units such that each functional unit is allocated a dedicated portion of register file 2858. In one embodiment, register file 2858 is partitioned between different meridians being executed by graphics multiprocessor 2834.
[0286] Each GPGPU core 2862 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 2834. According to embodiments, the architecture of the GPGPU core 2862 may be similar or different. For example, in one embodiment, a first portion of the GPGPU core 2862 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In one embodiment, the FPU may implement the IEEE 754-2008 floating-point arithmetic standard or enable variable-precision floating-point arithmetic. Additionally, the graphics multiprocessor 2834 may also include one or more fixed-function or special-function units for performing specific functions such as copying rectangles or pixel blending operations. In one embodiment, one or more of the GPGPU cores may also contain fixed-function or special-function logic.
[0287] In one embodiment, the GPGPU core 2862 includes SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU core 2862 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. The SIMD instructions for the GPGPU core can be generated at compile time by a shader compiler, or automatically generated when executing a program written and compiled for a Single Program Multiple Data (SPMD) or SIMT architecture. Multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, and in one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logic unit.
[0288] The memory and cache interconnect 2868 is an interconnect network that connects each functional unit of the graphics multiprocessor 2834 to the register file 2858 and shared memory 2870. In one embodiment, the memory and cache interconnect 2868 is a cross-switch interconnect that allows the load / store unit 2866 to perform load and store operations between the shared memory 2870 and the register file 2858. The register file 2858 can operate at the same frequency as the GPGPU core 2862, thus data transfers between the GPGPU core 2862 and the register file 2858 have very short latency. The shared memory 2870 can be used to implement communication between threads executing on functional units within the graphics multiprocessor 2834. For example, the cache memory 2872 can be used as a data cache to cache texture data communicated between functional units and texture units 2836. The shared memory 2870 can also be used as a cached, managed program. In addition to the automatically cached data stored in cache memory 2872, threads executing on GPGPU core 2862 can also programmatically store data in shared memory.
[0289] Figures 29A to 29B An additional graphics multiprocessor according to an embodiment is shown. The graphics multiprocessors 2925 and 2950 shown are... Figure 28C Variants of the 2834 graphics multiprocessor. The 2925 and 2950 graphics multiprocessors shown can be configured as streaming multiprocessors (SM) capable of executing a large number of execution threads simultaneously.
[0290] Figure 29A A graphics multiprocessor 2925 according to an additional embodiment is shown. The graphics multiprocessor 2925 includes, relative to... Figure 28DThe graphics multiprocessor 2834 may include multiple additional instances of its execution resource units. For example, the graphics multiprocessor 2925 may include multiple instances of instruction units 2932A to 2932B, register files 2934A to 2934B, and multiple texture units 2944A to 2944B. The graphics multiprocessor 2925 may also include multiple sets of graphics or compute execution units (e.g., GPGPU cores 2936A to 2936B, GPGPU cores 2937A to 2937B, GPGPU cores 2938A to 2938B) and multiple sets of load / store units 2940A to 2940B. In one embodiment, the execution resource units have a common instruction cache 2930, a texture and / or data cache memory 2942, and a shared memory 2946.
[0291] Various components can communicate via interconnect structure 2927. In one embodiment, interconnect structure 2927 includes one or more cross switches for enabling communication between various components of the graphics multiprocessor 2925. In another embodiment, interconnect structure 2927 is a separate high-speed network structure layer on which each component of the graphics multiprocessor 2925 is stacked. Components of the graphics multiprocessor 2925 communicate with remote components via interconnect structure 2927. For example, GPGPU cores 2936A to 2936B, 2937A to 2937B, and 2978A to 2938B can all communicate with shared memory 2946 via interconnect structure 2927. Interconnect structure 2927 can arbitrate communication within the graphics multiprocessor 2925 to ensure fair bandwidth allocation among components.
[0292] Figure 29B A graphics multiprocessor 2950 according to an additional embodiment is shown. Figure 28D and Figure 29A As shown, the graphics processor includes multiple sets of execution resources 2956A to 2956D, each set of execution resources including multiple instruction units, register files, GPGPU cores, and load memory units. Execution resources 2956A to 2956D can work with (multiple) texture units 2960A to 2960D to perform texture operations, while sharing instruction cache 2954 and shared memory 2962. In one embodiment, execution resources 2956A to 2956D can share instruction cache 2954, shared memory 2962, and multiple instances of texture and / or data cache memories 2958A to 2958B. Various components can be connected via... Figure 29A The interconnect structure 2927 communicates with the interconnect structure 2952 similar to the interconnect structure 2927.
[0293] Those skilled in the art will understand that Figure 27 , Figures 28A to 28D and Figures 29A to 29B The architecture described herein is descriptive and does not limit the scope of embodiments of the invention. Therefore, the techniques described herein can be implemented on any suitably configured processing unit, including but not limited to: one or more mobile application processors; one or more desktop computer or server central processing units (CPUs), including multi-core CPUs; one or more parallel processing units such as… Figure 28A The parallel processing unit 2802; and one or more graphics processors or dedicated processing units, without departing from the scope of the embodiments described herein.
[0294] In some embodiments, a parallel processor or GPGPU, as described herein, is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., high-speed interconnects such as PCIe or NVLink). In other embodiments, the GPU may be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). Regardless of how the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a job descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0295] Technologies for GPU-to-host processor interconnects
[0296] Figure 30A An exemplary architecture is shown in which multiple GPUs 3010 to 3013 are communicatively coupled to multiple multi-core processors 3005 to 3006 via high-speed links 3040 to 3043 (e.g., bus, point-to-point interconnect, etc.). In one embodiment, high-speed links 3040 to 3043 support communication throughput of 4Gb / s, 30Gb / s, 80GB / s, or higher, depending on the implementation. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. However, the basic principles of the invention are not limited to any particular communication protocol or throughput.
[0297] Furthermore, in one embodiment, two or more of GPUs 3010 to 3013 are interconnected via high-speed links 3044 to 3045, which can be implemented using the same or different protocols / links as those used for high-speed links 3040 to 3043. Similarly, two or more of multi-core processors 3005 to 3006 can be connected via high-speed link 3033, which can be a symmetric multiprocessor (SMP) bus operating at speeds of 20Gb / s, 30Gb / s, 120GB / s, or higher. Alternatively, Figure 30A All communication between the various system components shown can be accomplished using the same protocol / link (e.g., via a common interconnect structure). However, as mentioned, the basic principles of the invention are not limited to any particular type of interconnect technology.
[0298] In one embodiment, each multi-core processor 3005 to 3006 is communicatively coupled to processor memories 3001 to 3002 via memory interconnects 3030 to 3031, and each GPU 3010 to 3013 is communicatively coupled to GPU memories 3020 to 3023 via GPU memory interconnects 3050 to 3053. Memory interconnects 3030 to 3031 and 3050 to 3053 may utilize the same or different memory access technologies. By way of example and not limitation, processor memories 3001 to 3002 and GPU memories 3020 to 3023 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint or Nano-RAM. In one embodiment, one portion of the memory may be volatile memory, while another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0299] As described below, although the various processors 3005 to 3006 and GPUs 3010 to 3013 can each be physically coupled to specific memories 3001 to 3002 and 3020 to 3023 respectively, a unified memory architecture can be implemented, in which the same virtual system address space (also known as the “effective address” space) is distributed across all the various physical memories. For example, processor memories 3001 to 3002 can each include 64 GB of system memory address space, and GPU memories 3020 to 3023 can each include 32 GB of system memory address space (resulting in a total of 256 GB of addressable memory space in the example described).
[0300] Figure 30B Additional details are shown regarding the interconnection between a multi-core processor 3007 and a graphics acceleration module 3046 according to one embodiment. The graphics acceleration module 3046 may include one or more GPU chips integrated on a line card coupled to the processor 3007 via a high-speed link 3040. Alternatively, the graphics acceleration module 3046 may be integrated on the same package or chip as the processor 3007.
[0301] The processor 3007 shown includes multiple cores 3060A to 3060D, each having a translational backstop buffer 3061A to 3061D and one or more caches 3062A to 3062D. These cores may include various other components (e.g., instruction fetch units, branch prediction units, decoders, execution units, reordering buffers, etc.) for executing instructions and processing data not shown to avoid obscuring the basic principles of the invention. Caches 3062A to 3062D may include Level 1 (L1) and Level 2 (L2) caches. Furthermore, one or more shared caches 3026 may be included in the cache hierarchy and shared by the respective groups of cores 3060A to 3060D. For example, one embodiment of the processor 3007 includes 24 cores, each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one of the L2 and L3 caches is shared by two adjacent cores. The processor 3007 and the graphics acceleration module 3046 are connected to the system memory 3041, which may include processor memories 3001 to 3002.
[0302] Consistency is maintained for data and instructions stored in various caches 3062A to 3062D, 3056 and system memory 3041 via inter-core communication through the consistency bus 3064. For example, each cache may have an associated cache consistency logic / circuit system to communicate via the consistency bus 3064 in response to a detected read or write to a particular cache line. In one embodiment, a cache snooping protocol is implemented via the consistency bus 3064 to snoop on cache accesses. Cache snooping / consistency techniques will be well understood by those skilled in the art, and to avoid obscuring the basic principles of the invention, they will not be described in detail here.
[0303] In one embodiment, proxy circuitry 3025 communicatively couples graphics acceleration module 3046 to coherence bus 3064, thereby allowing graphics acceleration module 3046 to participate in cache coherence protocols as a peer of the core. Specifically, interface 3035 provides connectivity to proxy circuitry 3025 via high-speed link 3040 (e.g., PCIe bus, NVLink, etc.), and interface 3037 connects graphics acceleration module 3046 to link 3040.
[0304] In one implementation, the accelerator integrated circuit 3036 provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines 3031, 3032, and N of the graphics acceleration module 3046. The graphics processing engines 3031, 3032, and N may each include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 3031, 3032, and N may include different types of graphics processing engines within the GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and block image transfer engines. In other words, the graphics acceleration module may be a GPU with multiple graphics processing engines 3031 to 3032, and N, or the graphics processing engines 3031 to 3032, and N may be separate GPUs integrated in a common package, line card, or chip.
[0305] In one embodiment, the accelerator integrated circuit 3036 includes a memory management unit (MMU) 3039 for performing various memory management functions such as virtual-to-physical memory translation (also known as effective-to-real memory translation) and memory access protocols for accessing system memory 3041. The MMU 3039 may also include a translation back buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translation. In one embodiment, cache 3038 stores commands and data for efficient access by graphics processing engines 3031 to 3032, N. In one embodiment, the data stored in cache 3038 and graphics memories 3033 to 3034, N is kept consistent with core caches 3062A to 3062D, 3056 and system memory 3011. As mentioned, this can be accomplished via proxy circuitry 3025, which participates in cache coherency mechanisms on behalf of cache 3038 and memories 3033 to 3034, N (e.g., sending updates to cache 3038 related to modifications / accesses to cache lines on processor caches 3062A to 3062D, 3056 and receiving updates from cache 3038).
[0306] A set of registers 3045 stores context data for threads executed by graphics processing engines 3031 to 3032, N, and context management circuitry 3048 manages the thread context. For example, context management circuitry 3048 can perform save and restore operations to save and restore the context of various threads during context switching (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during context switching, context management circuitry 3048 can store the current register value to a designated area in memory (e.g., identified by a context pointer). The context management circuitry can restore the register value upon returning to the context. In one embodiment, interrupt management circuitry 3047 receives and processes interrupts received from the system device.
[0307] In one implementation, the MMU 3039 translates the virtual / effective address from the graphics processing engine 3031 into a physical / actual address in system memory 3011. One embodiment of the accelerator integrated circuit 3036 supports multiple (e.g., 4, 8, 16) graphics acceleration modules 3046 and / or other accelerator devices. The graphics acceleration module 3046 may be dedicated to a single application executing on the processor 3007, or it may be shared among multiple applications. In one embodiment, a virtual graphics execution environment is presented, wherein the resources of graphics processing engines 3031 to 3032, N are shared with multiple applications or virtual machines (VMs). Resources may be subdivided into “shards” allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0308] Therefore, the accelerator integrated circuit acts as a bridge for the system of the graphics acceleration module 3046, and provides address translation and system memory caching services. Furthermore, the accelerator integrated circuit 3036 can provide virtualization facilities for the host processor to manage the virtualization of the graphics processing engine, interrupts, and memory management.
[0309] Because the hardware resources of graphics processing engines 3031 to 3032, N are explicitly mapped to the actual address space seen by the host processor 3007, any host processor can directly address these resources using valid address values. In one embodiment, one function of the accelerator integrated circuit 3036 is the physical separation of graphics processing engines 3031 to 3032, N, so that the graphics processing engines appear as independent units on the system.
[0310] As mentioned, in the illustrated embodiment, one or more graphics memories 3033 to 3034, M are coupled to each of the graphics processing engines 3031 to 3032, N, respectively. Graphics memories 3033 to 3034, M store instructions and data processed by each of the graphics processing engines 3031 to 3032, N. Graphics memories 3033 to 3034, M may be volatile memories, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories such as 3D XPoint or Nano-RAM.
[0311] In one embodiment, to reduce data traffic on link 3040, a biasing technique is used to ensure that the data stored in graphics memories 3033 to 3034, M is the data most frequently used by graphics processing engines 3031 to 3032, N, and preferably not used (or at least infrequently used) by cores 3060A to 3060D. Similarly, the biasing mechanism attempts to keep the data required by the cores (and preferably not graphics processing engines 3031 to 3032, N) within the caches 3062A to 3062D, 3056 of the cores and system memory 3011.
[0312] Figure 30C Another embodiment in which the accelerator integrated circuit 3036 is integrated within the processor 3007 is shown. In this embodiment, graphics processing engines 3031 to 3032, N communicate directly with the accelerator integrated circuit 3036 via high-speed link 3040 through interfaces 3037 and 3035 (this can also utilize any form of bus or interface protocol). The accelerator integrated circuit 3036 can perform operations related to... Figure 30B The operation described is the same, but given its close proximity to the coherence bus 3062 and caches 3062A to 3062D, 3026, it may operate at a higher throughput.
[0313] One embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The shared programming model may include a programming model controlled by the accelerator integrated circuit 3036 and a programming model controlled by the graphics acceleration module 3046.
[0314] In one embodiment of the dedicated process model, graphics processing engines 3031 to 3032, N are dedicated to a single application or process within a single operating system. A single application can centralize requests from other applications to graphics engines 3031 to 3032, N, thereby providing virtualization within a VM / partition.
[0315] In a dedicated process programming model, graphics processing engines 3031 to 3032, N can be shared by multiple VM / application partitions. The shared model requires a hypervisor to virtualize the graphics processing engines 3031 to 3032, N, allowing access by each operating system. For a single-partition system without a hypervisor, the graphics processing engines 3031 to 3032, N are owned by the operating system. In both cases, the operating system can virtualize the graphics processing engines 3031 to 3032, N to provide access to each process or application.
[0316] For a shared programming model, the graphics acceleration module 3046 or the individual graphics processing engines 3031 to 3032, N use a process handle to select process elements. In one embodiment, the process elements are stored in system memory 3011 and can be addressed using the effective address to physical address translation techniques described herein. The process handle may be a implementation-specific value provided to the host process when registering its context with the graphics processing engines 3031 to 3032, N (i.e., invoking system software to add process elements to the process element linked table). The lower 16 bits of the process handle may be an offset of the process element within the process element linked table.
[0317] Figure 30D An exemplary accelerator integration slice 3090 is shown. As used herein, a “slice” refers to a designated portion of the processing resources of the accelerator integrated circuit 3036. The application-effective address space 3082 within system memory 3011 stores process elements 3083. In one embodiment, process element 3083 is stored in response to a GPU call 3081 from an application 3080 executing on processor 3007. Process element 3083 contains the processing state of the corresponding application 3080. The job descriptor (WD) 3084 contained in process element 3083 may be a single job requested by the application, or it may contain a pointer to a job queue. In the latter case, WD 3084 is a pointer to a job request queue in the application address space 3082.
[0318] The graphics acceleration module 3046 and / or individual graphics processing engines 3031 to 3032, N can be shared by all or some processes in the system. Embodiments of the invention include infrastructure for establishing a processing state and sending a WD3084 to the graphics acceleration module 3046 to begin work in a virtual environment.
[0319] In one implementation, a dedicated process programming model is implementation-specific. In this model, a single process owns either the graphics acceleration module 3046 or a separate graphics processing engine 3031. Since the graphics acceleration module 3046 is owned by a single process, the hypervisor initializes the accelerator integrated circuit 3036 to obtain its assigned partition, and the operating system initializes the accelerator integrated circuit 3036 to obtain its assigned process when the graphics acceleration module 3046 is allocated.
[0320] In operation, the WD acquisition unit 3091 in the accelerator integration slice 3090 acquires the next WD 3084, which includes instructions for work to be performed by one of the graphics processing engines of the graphics acceleration module 3046. As shown, data from the WD 3084 can be stored in register 3045 and used by the MMU 3039, interrupt management circuitry 3047, and / or context management circuitry 3048. For example, one embodiment of the MMU 3039 includes a segment / page lookup circuitry system for accessing segment / page tables 3086 within the OS virtual address space 3085. The interrupt management circuitry 3047 can handle interrupt events 3092 received from the graphics acceleration module 3046. When performing graphics operations, the effective address 3093 generated by the graphics processing engines 3031 to 3032, N is translated into an actual address by the MMU 3039.
[0321] In one embodiment, the same set of registers 3045 is copied for each graphics processing engine 3031 to 3032, N and / or graphics acceleration module 3046, and this set of registers can be initialized by a hypervisor or operating system. Each of these copied registers can be included in the accelerator integration slice 3090. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.
[0322] Table 1 - Supervisor Initialization Registers
[0323]
[0324]
[0325] Table 2 shows exemplary registers that can be initialized by the operating system.
[0326] Table 2 - Operating System Initialization Registers
[0327] 2 Valid Address (EA) Context Save / Restore Pointer 3 Virtual address (VA) accelerators utilize record pointers 4 Virtual address (VA) memory segment table pointer 5 Authorization mask 6 Job descriptor
[0328] In one embodiment, each WD 3084 is specific to a particular graphics acceleration module 3046 and / or graphics processing engines 3031 to 3032, N. The WD contains all the information required for the graphics processing engines 3031 to 3032, N to complete their work, or the WD may be a pointer to a memory location where the application has established a queue of work commands to be completed.
[0329] Figure 30E Additional details of one embodiment of the shared model are shown. This embodiment includes a hypervisor physical address space 3098 in which a list of process elements 3099 is stored. The hypervisor physical address space 3098 is accessible via a hypervisor 3096 that virtualizes the graphics acceleration module engine of operating system 3095.
[0330] The shared programming model allows all or some processes from all or some partitions of the system to use the graphics acceleration module 3046. There are two programming models in which the graphics acceleration module 3046 is shared by multiple processes and partitions: time-sliced sharing and direct graphics sharing.
[0331] In this model, the hypervisor 3096 possesses the graphics acceleration module 3046 and makes its functionality available to all operating systems 3095. To enable the graphics acceleration module 3046 to support the virtualization of the hypervisor 3096, the graphics acceleration module 3046 may meet the following requirements:
[0332] 1) The application's job requests must be autonomous (i.e., no need to maintain state between jobs), or the graphics acceleration module 3046 must provide a context saving and restoring mechanism. 2) The graphics acceleration module 3046 guarantees completion of the application's job requests within a specified timeframe, including any conversion errors, or the graphics acceleration module 3046 provides the ability to preempt job processing. 3) When operating with a direct shared programming model, fairness among the graphics acceleration modules 3046 during the process must be guaranteed.
[0333] In one embodiment, for the shared model, application 3080 is required to make an operating system 3095 system call using the graphics acceleration module 3046 type, working descriptor (WD), authorization mask register (AMR) value, and context save / restore region pointer (CSRP). The graphics acceleration module 3046 type describes the target acceleration function of the system call. The graphics acceleration module 3046 type can be a system-specific value. The WD is specifically formatted for the graphics acceleration module 3046 and can be in the following forms: graphics acceleration module 3046 command; valid address pointer to a user-defined structure; valid address pointer to a command queue; or any other data structure describing the work to be performed by the graphics acceleration module 3046. In one embodiment, the AMR value is the AMR state for the current process. The value passed to the operating system is similar to that of the application setting the AMR. If the implementation of the accelerator integrated circuit 3036 and the graphics acceleration module 3046 does not support the User Authorization Mask Override Register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. Before placing the AMR in process element 3083, hypervisor 3096 may optionally apply the Current Authorization Mask Override Register (AMOR) value. In one embodiment, CSRP is one of registers 3045 containing the effective address of a region in application address space 3082 for the graphics acceleration module 3046 to save and restore context state. This pointer is optional if saving state between jobs is not required or when a job is preempted. The context save / restore region may be plugged-in system memory.
[0334] Upon receiving a system call, the operating system 3095 can verify that the application 3080 has been registered and authorized to use the graphics acceleration module 3046. The operating system 3095 then uses the information shown in Table 3 to invoke the hypervisor 3096.
[0335] Table 3 - Operating System Call Parameters for the Hypervisor
[0336]
[0337]
[0338] Upon receiving a call from the hypervisor, the hypervisor 3096 can verify that the operating system 3095 has been registered and authorized to use the graphics acceleration module 3046. The hypervisor 3096 then places the process element 3083 into a process element linked table corresponding to the graphics acceleration module 3046 type. The process element may contain the information shown in Table 4.
[0339] Table 4 - Process Element Information
[0340] 2 Authorization Mask Register (AMR) value (may be masked) 3 Valid Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual address (VA) accelerators utilize record pointers (AURP). 6 Virtual address of the Memory Segment Table Pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt vector table, exported from hypervisor call parameters 9 Status Register (SR) Value 10 Logical Partition ID (LPID) 11 The Real Address (RA) management accelerator utilizes record pointers 12 Storage Descriptor Register (SDR)
[0341] In one embodiment, the hypervisor initializes the multiple accelerator integration slice 3090 of register 3045.
[0342] like Figure 30F As shown, one embodiment of the invention employs a unified memory addressable via a common virtual memory address space for accessing physical processor memories 3001-3002 and GPU memories 3020-3023. In this embodiment, operations performed on GPUs 3010-3013 utilize the same virtual / effective memory address space to access processor memories 3001-3002 and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 3001, a second portion to second processor memory 3002, a third portion to GPU memory 3020, and so on. The entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memories 3001-3002 and GPU memories 3020-3023, thereby allowing any processor or GPU to access any physical memory having a virtual address mapped to said memory.
[0343] In one embodiment, the bias / coherence management circuitry 3094A to 3094E within one or more of the MMUs 3039A to 3039E ensures cache coherence between the host processor (e.g., 3005) and the caches of the GPUs 3010 to 3013, as well as biasing techniques that indicate the physical memory where certain types of data should be stored. Although in Figure 30F Several instances of bias / coherence management circuitry systems 3094A to 3094E are shown, but bias / coherence circuitry systems can also be implemented within the MMU of one or more host processors 3005 and / or within the accelerator integrated circuit 3036.
[0344] One embodiment allows GPU-attached memories 3020 to 3023 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology without suffering the typical performance drawbacks associated with system-wide cache coherence. The ability to access GPU-attached memories 3020 to 3023 as system memory avoids heavy cache coherence overhead, providing a favorable operating environment for GPU offloading. This arrangement allows host processor 3005 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. These traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, which are inefficient compared to simple memory accesses. Simultaneously, the ability to access GPU-attached memories 3020 to 3023 without cache coherence overhead can be critical for offloading computation execution time. For example, in scenarios with heavy streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 3010 to 3013. The efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation all play a crucial role in determining the effectiveness of GPU offloading.
[0345] In one implementation, the selection between GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which may be a page-granular structure comprising 1 or 2 bits per GPU-attached memory page (i.e., controlled at the memory page level). The bias table can be implemented within the stolen memory range of one or more GPU-attached memories 3020 to 3023, with or without a bias cache in GPUs 3010 to 3013 (e.g., caching frequently / recently used entries of the bias table). Alternatively, the entire bias table can be maintained within the GPU.
[0346] In one implementation, the bias table entries associated with each access to GPU-attached memories 3020-3023 are accessed before the actual access to GPU memory, such that: First, native requests from GPUs 3010-3013 that find pages in the GPU bias are directly forwarded to the corresponding GPU memories 3020-3023. Native requests from GPUs that find pages in the host bias are forwarded to processor 3005 (e.g., via a high-speed link as described above). In one embodiment, a request from processor 3005 that finds the requested page in the host processor bias completes like a normal memory read. Alternatively, requests for GPU-biased pages can be forwarded to GPUs 3010-3013. If the GPU is not currently using the page, the GPU can convert the page to host processor bias.
[0347] The page bias state can be changed through software-based mechanisms, hardware-assisted software mechanisms, or, for a finite set of cases, hardware-only mechanisms.
[0348] One mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn invokes the GPU's device driver. The driver then sends a message to the GPU (or enqueues a command descriptor), thereby instructing the GPU to change its bias state. For certain transitions, a cache dump clearing operation is performed on the host machine. This cache dump clearing operation is necessary for transitions from host processor 3005 bias to GPU bias, but not necessary for the reverse transition.
[0349] In one embodiment, cache coherence is maintained by temporarily presenting GPU bias pages that the host processor 3005 cannot cache. To access these pages, the processor 3005 may request access from the GPU 3010, which may grant access immediately or not, depending on the implementation. Therefore, to reduce communication between the processor 3005 and the GPU 3010, it is advantageous to ensure that the GPU bias pages are pages needed by the GPU but not by the host processor 3005, and vice versa.
[0350] Graphics processing pipeline
[0351] Figure 31 A graphics processing pipeline 3100 according to an embodiment is illustrated. In one embodiment, a graphics processor may implement the illustrated graphics processing pipeline 3100. The graphics processor may be included in a parallel processing subsystem such as those described herein. Figure 28A Within the parallel processor 2800, in one embodiment, the parallel processor is Figure 27 Variations of the (multiple) parallel processors 2712. As described herein, various parallel processing systems can be implemented via parallel processing units (e.g., Figure 28A One or more instances of parallel processing units 2802 are used to implement the graphics processing pipeline 3100. For example, shader units (e.g., Figure 29A The graphics multiprocessor 2834 can be configured to perform the functions of one or more of the vertex processing unit 3104, tessellation control processing unit 3108, tessellation evaluation processing unit 3112, geometry processing unit 3116, and fragment / pixel processing unit 3124. The functions of the data assembler 3102, primitive assemblers 3106, 3114, 3118, tessellation unit 3110, rasterizer 3122, and raster operation unit 3126 can also be handled by a processing cluster (e.g., Figure 3Other processing engines and corresponding partition units (e.g., within the processing cluster 214) Figure 2 The graphics processing pipeline 3100 is executed by partitioning units 220A to 220N. The graphics processing pipeline 3100 can also be implemented using one or more dedicated processing units. In one embodiment, one or more portions of the graphics processing pipeline 3100 can be executed by parallel processing logic within a general-purpose processor (e.g., a CPU). In one embodiment, one or more portions of the graphics processing pipeline 3100 can access on-chip memory (e.g., such as memory interface 3128) via a memory interface 3128. Figure 28A The parallel processor memory 2822 shown can have the memory interface as follows: Figure 28A An example of the memory interface 2818.
[0352] In one embodiment, the data assembler 3102 is a processing unit that collects vertex data of surfaces and primitives. The data assembler 3102 then outputs vertex data, including vertex attributes, to the vertex processing unit 3104. The vertex processing unit 3104 is a programmable execution unit that executes a vertex shader program to illuminate and transform vertex data as specified by the vertex shader program. The vertex processing unit 3104 reads data stored in cache, native, or system memory for processing vertex data and can be programmed to convert vertex data from an object-based coordinate representation to world space coordinate space or normalized device coordinate space.
[0353] The first instance of the primitive assembler 3106 receives vertex attributes from the vertex processing unit 50. The primitive assembler 3106 reads the stored vertex attributes as needed and constructs graphic primitives for processing by the tessellation control processing unit 3108. Graphic primitives include triangles, line segments, points, patches, etc., supported by various graphics processing application programming interfaces (APIs).
[0354] The tessellation control processing unit 3108 treats input vertices as control points for a geometric patch. These control points are transformed from an input representation from the patch (e.g., the patch's basis) into a representation suitable for surface evaluation by the tessellation evaluation processing unit 3112. The tessellation control processing unit 3108 can also calculate tessellation factors for the edges of the geometric patch. The tessellation factor applies to individual edges and quantifies the viewpoint-related level of detail associated with the edges. The tessellation unit 3110 is configured to receive the tessellation factors for the edges of the patch and subdivide the patch into multiple geometric primitives, such as line, triangle, or quadrilateral primitives, which are then transmitted to the tessellation evaluation processing unit 3112. The tessellation evaluation processing unit 3112 operates on the parameterized coordinates of the subdivided patch to generate a surface representation and vertex attributes associated with each vertex of the geometric primitives.
[0355] A second instance of the primitive assembler 3114 receives vertex attributes from the tessellation evaluation processing unit 3112, reads the stored vertex attributes as needed, and constructs graphic primitives for processing by the geometry processing unit 3116. The geometry processing unit 3116 is a programmable execution unit that executes a geometry shader program to transform the graphic primitives received from the primitive assembler 3114 as specified by the geometry shader program. In one embodiment, the geometry processing unit 3116 is programmed to subdivide the graphic primitives into one or more new graphic primitives and calculate parameters for rasterizing the new graphic primitives.
[0356] In some embodiments, the geometry processing unit 3116 can add or remove elements from the geometry stream. The geometry processing unit 3116 outputs parameters and vertices specifying new graphic primitives to the primitive assembler 3118. The primitive assembler 3118 receives the parameters and vertices from the geometry processing unit 3116 and constructs graphic primitives for processing by the viewport scaling, culling, and clipping unit 3120. The geometry processing unit 3116 reads data stored in the parallel processor memory or system memory for processing the geometry data. The viewport scaling, culling, and clipping unit 3120 performs clipping, culling, and viewport scaling, and outputs the processed graphic primitives to the rasterizer 3122.
[0357] Rasterizer 3122 can perform depth culling and other depth-based optimizations. Rasterizer 3122 also performs scan transformations on new graphic primitives to generate segments and outputs these segments and associated overlay data to segment / pixel processing unit 3124. Segment / pixel processing unit 3124 is a programmable execution unit configured to execute segment shader programs or pixel shader programs. Segment / pixel processing unit 3124 transforms segments or pixels received from rasterizer 3122 as specified by the segment or pixel shader program. For example, segment / pixel processing unit 3124 can be programmed to perform operations including but not limited to texture mapping, shading, blending, texture correction, and perspective correction to produce shaded segments or pixels output to raster operation unit 3126. Segment / pixel processing unit 3124 can read data stored in parallel processor memory or system memory for use when processing segment data. Segment or pixel shader programs can be configured to shade at sample, pixel, tile, or other granularities according to a sampling rate configured for the processing unit.
[0358] The raster operation unit 3126 is a processing unit that performs raster operations including but not limited to stenciling, z-testing, and blending, and outputs pixel data as processed graphic data for storage in a graphics memory (e.g., Figure 28A Parallel processor memory 2822 and / or such Figure 27The system memory 2704 is used for display on one or more display devices 2710 or for further processing by one or more processors 2702 or (a plurality of) parallel processors 2712. In some embodiments, the raster operation unit 3126 is configured to compress z or color data written to memory and decompress z or color data read from memory.
[0359] In embodiments, the terms "engine," "module," or "logic" may refer to, be part of, or include: an application-specific integrated circuit (ASIC), electronic circuitry, a processor (shared processor, dedicated processor, or group processor) and / or memory (shared memory, dedicated memory, or group memory), combinational logic circuitry, and / or other suitable components that provide the described functionality. In embodiments, the engine or module may be implemented as firmware, hardware, software, or any combination of firmware, hardware, and software.
[0360] Embodiments of the present invention may include the steps described above. These steps may be embodied as machine-executable instructions that can be used to cause a general-purpose or special-purpose processor to perform these steps. Alternatively, these steps may be performed by a specific hardware component containing hard-wired logic for performing these steps, or by any combination of programmable computer components and custom hardware components.
[0361] As described herein, instructions can refer to a specific configuration of hardware, such as an application-specific integrated circuit (ASIC) configured to perform certain operations or having predetermined functions or software instructions stored in memory implemented on a non-transitory computer-readable medium. Therefore, the techniques illustrated in the figures can be implemented using code and data stored and executed on one or more storage devices (e.g., end stations, network elements, etc.). Such electronic devices use computer machine-readable media to store and transmit (internal and / or using other electronic devices on a network) code and data, such as non-transitory computer machine-readable storage media (e.g., magnetic disks; optical disks; random access memory; read-only memory; flash memory; phase-change memory) and transient computer machine-readable communication media (e.g., electrical, optical, acoustic, or other forms of propagation signals—e.g., carrier waves, infrared signals, digital signals, etc.).
[0362] Furthermore, such electronic devices typically include a group of one or more processors coupled to one or more other components, such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., keyboards, touchscreens, and / or displays), and network connectivity. The coupling of this group of processors and other components is typically via one or more buses and bridges (also referred to as bus controllers). Storage devices and signals carrying network traffic represent one or more machine-readable storage media and machine-readable communication media, respectively. Thus, the storage devices of a given electronic device typically store code and / or data for execution on the group of one or more processors of that electronic device. Of course, different combinations of software, firmware, and / or hardware can be used to implement one or more portions of embodiments of the invention. Throughout this detailed description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without some of these specific details. In some instances, well-known structures and functions have not been described in detail to avoid obscuring the subject matter of the invention. Therefore, the scope and spirit of the invention should be determined according to the following claims.
Claims
1. A device for graphics processing, comprising: Multiple graphics processing unit cores; A memory controller for coupling the graphics processor core to multiple local memory devices; An interconnect for coupling at least one external processing device to the local memory device, the interconnect being used to ensure that data accessed from the local memory device remains consistent; as well as A memory management circuit is configured to map shared virtual memory (SVM) addresses to physical addresses of memory pages in one of a local memory device and another local memory device and a system memory device, wherein the SVM addresses are shared with the at least one external processing device to allow the external data processing device to access memory pages from the one local memory device and the other local memory device and the system memory device using the SVM addresses. The memory management circuitry is used to: identify multiple memory pages that are accessed more frequently by the graphics processor core than by the at least one external processing device; and responsively biasing the plurality of memory pages to support access by the graphics processor core, wherein the external processing device is configured to maintain accessibility to the biased plurality of memory pages supporting access by the graphics processor core. In order to bias the plurality of memory pages, the memory management circuitry is used to move the plurality of memory pages from one of the other local memory device and the system memory device to the one local memory device.
2. The device as described in claim 1, characterized in that, To bias the plurality of memory pages, the memory management circuitry is configured to provide access by the graphics processor core to the plurality of memory pages without first sending a request to the at least one external processing device.
3. The device as described in claim 1, characterized in that, The at least one external processing device includes a central processing unit (CPU).
4. The device as described in claim 1, characterized in that, Access to the plurality of memory pages is tracked by updating the data structure.
5. The device as described in claim 1, characterized in that, The local memory device includes high-bandwidth memory (HBM).
6. The device as described in claim 1, characterized in that, The at least one external processing device accesses the local memory device through the interconnect.
7. A method for graphics processing, comprising: At least one external processing device is coupled to multiple local memory devices, and multiple graphics processor cores are coupled to the local memory devices. Interconnects are used to ensure that data accessed from the local memory devices remains consistent. The shared virtual memory (SVM) address is mapped to the physical address of a memory page in one of a local memory device and another local memory device and the system memory device. The SVM address is shared with the at least one external processing device to allow the external processing device to use the SVM address to access memory pages from one of the local memory device and another local memory device and the system memory device. Identify multiple memory pages that are accessed more frequently by the graphics processor core than by the at least one external processing device; as well as The plurality of memory pages are responsively biased to support access by the graphics processor core, wherein the external processing device is configured to maintain accessibility to the biased plurality of memory pages supporting access by the graphics processor core. The biasing of the plurality of memory pages includes moving the plurality of memory pages from one of the other local memory device and the system memory device to the one local memory device.
8. The method as described in claim 7, characterized in that, Biasing the plurality of memory pages includes providing access by the graphics processor core to the plurality of memory pages without first sending a request to the at least one external processing device.
9. The method as described in claim 7, characterized in that, The at least one external processing device includes a central processing unit (CPU).
10. The method as described in claim 7, characterized in that, Access to the plurality of memory pages is tracked by updating the data structure.
11. The method as described in claim 7, characterized in that, The local memory device includes high-bandwidth memory (HBM).
12. The method as described in claim 7, characterized in that, The at least one external processing device accesses the local memory device through the interconnect.
13. A machine-readable medium having program code stored thereon, which, when executed by a machine, causes the machine to perform the following operations: At least one external processing device is coupled to multiple local memory devices, and multiple graphics processor cores are coupled to the local memory devices. Interconnects are used to ensure that data accessed from the local memory devices remains consistent. The shared virtual memory (SVM) address is mapped to the physical address of a memory page in one of a local memory device and another local memory device and the system memory device. The SVM address is shared with the at least one external processing device to allow the external processing device to use the SVM address to access memory pages from one of the local memory device and another local memory device and the system memory device. Identify multiple memory pages that are accessed more frequently by the graphics processor core than by the at least one external processing device; as well as The plurality of memory pages are responsively biased to support access by the graphics processor core, wherein the external processing device is configured to maintain accessibility to the biased plurality of memory pages supporting access by the graphics processor core. The program code, when executed by the machine, causes the machine to move the plurality of memory pages from one of the other local memory device and the system memory device to the one local memory device.
14. The machine-readable medium as claimed in claim 13, characterized in that, When executed by the machine, the program code causes the machine to provide access by the graphics processor core to the plurality of memory pages without first sending a request to the at least one external processing device.
15. The machine-readable medium as claimed in claim 13, characterized in that, The at least one external processing device includes a central processing unit (CPU).
16. The machine-readable medium as claimed in claim 13, characterized in that, Access to the plurality of memory pages is tracked by updating the data structure.
17. The machine-readable medium as claimed in claim 13, characterized in that, The local memory device includes high-bandwidth memory (HBM).
18. The machine-readable medium as claimed in claim 13, characterized in that, The at least one external processing device accesses the local memory device through the interconnect.
19. An apparatus for graphics processing, comprising: A means for coupling at least one external processing device to a plurality of local memory devices, wherein a plurality of graphics processor cores are coupled to the local memory devices, and interconnections are used to ensure that data accessed from the local memory devices remains consistent; A means for mapping a shared virtual memory (SVM) address to a physical address of a memory page in one of a local memory device and another local memory device and a system memory device, the SVM address being shared with the at least one external processing device to allow the external processing device to use the SVM address to access memory pages from one of the local memory device and the other local memory device and the system memory device. A means for identifying multiple memory pages accessed more frequently by the graphics processor core than by the at least one external processing device; as well as A means for responsively biasing the plurality of memory pages to support access by the graphics processor core, wherein the external processing device is configured to maintain accessibility to the biased plurality of memory pages supporting access by the graphics processor core. The means for biasing the plurality of memory pages includes means for moving the plurality of memory pages from one of the other local memory device and the system memory device to the one local memory device.
20. The device as claimed in claim 19, characterized in that, The means for biasing the plurality of memory pages includes means for providing access by the graphics processor core to the plurality of memory pages without first sending a request to the at least one external processing device.
21. The device as claimed in claim 19, characterized in that, The at least one external processing device includes a central processing unit (CPU).
Citation Information
Patent Citations
Multiprocessor system with independent direct access to bulk solid state memory resources
CN106462510A
Instruction and logic for support of code modification in translation lookaside buffers
US20160085685A1