SYNCHRONIZING MEMORY MANAGEMENT UNITS IN MULTI-DIELET PROCESSOR ARCHITECTURES
Patent Information
- Application Number
- DE102024137677
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-06
- Filing Date
- 2024-12-13
- Publication Date
- 2025-09-18
Smart Images

Figure 00000028_0000 
Figure 00000029_0000 
Figure 00000030_0000
Abstract
Description
AREA
[0001] This technology generally refers to multi-dielet processors. Specifically, the technology relates to distributed graphics and compute engines across multiple dies and the synchronization of memory operations in such processors. background
[0002] Demand for processors with extensive parallel processing capabilities, such as graphics processing units (GPUs), continues to grow. The processing requirements for such processors are also rapidly increasing in terms of workload complexity and size, as well as throughput.
[0003] The demands of a GPU, characterized by rapidly growing complexity, size, and throughput, mean that ever-increasing numbers of components must be packed onto the GPU semiconductor die—typically a thumbnail-sized square of flat semiconductor material such as silicon, cut from a wafer, on which the circuitry is fabricated. The more components housed on a die, the more functions an integrated circuit containing that die can provide. Therefore, chip designers strive to fit an ever-increasing number and variety of components onto each physical die.
[0004] There are physical limits to how many components can be packed onto a single die or chip. For example, packing more transistors generates more heat, which could damage the chip if cooling is not managed properly. More components, often smaller components, on a single die or chip can complicate interconnect implementation and lead to signaling problems and the like on the interconnects. Furthermore, despite Moore's Law, some components may have a minimum size beyond which they cannot be easily miniaturized.
[0005] Therefore, the effort to pack more components onto a single die or chip may encounter insurmountable limitations in terms of component count, component types, or physical size of processors, as workload demands continue to increase. To meet growing workload demands, in addition to increasing the number and types of components on a single die or chip, other options for expanding the processing capacities and capabilities of processors, such as GPUs, can be explored. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 shows a multi-dielet GPU according to some embodiments of the present disclosure. Fig. 2 shows another multi-dielet GPU according to some embodiments of the present disclosure. Fig. 3 schematically illustrates an example of a hierarchy arrangement of a translation lookaside buffer (TLB) in each dielet of a multi-dielet GPU according to some embodiments of the present disclosure. Fig. Figure 4 shows a flow diagram for a binding process in a multi-dielet processing system such as that shown in the Fig. 1-3, according to some embodiments of the present disclosure. Fig. Figure 5 shows a flowchart for TLB invalidation in a multi-dielet GPU such as the one shown in the Fig. 1-3, according to some embodiments of the present disclosure. Fig. Figure 6A illustrates a flowchart for a page table walk in a multi-dielet GPU as described in the Fig. 1-3, according to some embodiments. Fig. 6B shows an example flowchart of a method for address translation service (ATS) requests and responses in a multi-dielet GPU according to some embodiments of this disclosure. Fig. 6C shows a flowchart for an error reporting method that may be used in a multi-dielet GPU according to some embodiments of the present disclosure. Fig. 7 shows an exemplary GPU on a dielet of the multi-dielet GPU, with its frame buffer hub (FBHUB) and high speed hub (HSHUB) identified with respect to some of their connections, according to some embodiments of the present disclosure. Fig. 8 shows an example parallel processing unit of a GPU according to some embodiments. Fig. Figure 9A shows an exemplary general processing cluster (GPC) within the parallel processing unit of Fig. 8 according to some embodiments. Fig. 9B shows an exemplary memory partition unit of the parallel processing unit of Fig. 8. Fig. 10A illustrates an example streaming multiprocessor (SM) of Fig. 9A with an MMA state machine circuit according to some embodiments. Fig. Figure 10B conceptually illustrates four sub-areas that are used in an SM such as the one in Fig. 10A are implemented according to some embodiments. Fig. 11A is an example conceptual diagram of a processing system implemented using the parallel processing unit (PPU) of Fig. 8 is implemented. Fig. 11B is a block diagram of an example system in which the various architectures and / or functionalities of the various previous embodiments may be implemented. DETAILED DESCRIPTION OF NON-LIMITING EXEMPLARY EMBODIMENTS
[0006] To solve the problem of how to fit an ever-increasing number of components onto a single die, it has been proposed to interconnect multiple physical dies ("dielets") to form a larger and more complex processing system, such as a GPU. In this disclosure, each of the multiple dies in such a larger processing system is referred to as a "dielet." Such larger processing systems are referred to herein as "multi-dielet processing systems." Designing a processing system that extends beyond the confines of a single fabricated die, as in a multi-dielet processing system, offers a new path to scalability and eliminates some previously existing physical limitations.
[0007] For descriptive purposes, this disclosure refers to a multi-dielet GPU having two independently manufactured dielets, each dielet including one or more streaming multiprocessors (SMs), specific and general-purpose hardware engines, and associated routing components used for application performance on behalf of CPU- or GPU-initiated processes. In some embodiments, the dielets may be identical (i.e., identical or substantially identical hardware). In such cases, the multi-dielet GPU may have duplicate components between the two dielets. Depending on the application, such duplicate components may be leveraged or redundant. In some embodiments, the GPUs on the dielets may not be identical, and some GPUs may have different collections of hardware units.It is necessary that example embodiments consider various scenarios of combining multiple GPUs with different capabilities.
[0008] However, as such larger GPUs are formed, it is often necessary or efficient to shield the software from requiring knowledge of the physical layout of such a larger GPU. Such shielding may, for example, be required to ensure that the multi-dielet GPU is interoperable in many usage scenarios without requiring extensive adaptations of the software stack for each scenario. Such shielding of the software from knowledge of the detailed organization of the multi-dielet GPU allows, at least in some cases, the retrofitting of various multi-dielet GPUs into existing software (e.g., GPU driver software). Shielding can also "future-proof" multi-dielet designs by ensuring that the software can operate with such multi-dielet GPUs regardless of the specific design (e.g., number of dielets, types and number of hardware engines on each dielet, etc.).
[0009] The multi-dielet processing system of the embodiments is configured to be viewed entirely, or at least largely, as a single processor by software, such as driver software running on a CPU of a system including the multi-dielet processing system. One aspect of presenting the multi-dielet processing system to software as a single processor involves presenting a single logical memory management unit (MMU) to the software. Since each dielet requires memory translation and memory protection, and the dielets are to be presented to the software as a single processor, the multi-dielet processing system is configured such that the multiple MMUs, particularly the highest-level MMU on each of the dielets, are synchronized with each other to form a single logical MMU.
[0010] If the hardware mechanisms described in this disclosure did not provide synchronization between the MMUs of the multiple dielets, software would have to provide the necessary coordination between the respective MMUs. This would also require developers to specifically recognize that they are addressing multiple GPUs, similar to what is required for peer-to-peer combining of multiple independent GPUs. Such software-based coordination would not only have a negative impact on performance but would also incur high development costs.
[0011] Fig. 1 shows a multi-dielet GPU 100 according to some embodiments of the present disclosure. The multi-dielet GPU 100 includes two dielets—Dielet-1 102a and Dielet-2 102b. Each dielet includes a GPU, with Dielet-1 102a including a GPU-1 104a and Dielet-2 102b including a GPU-2 104b. In some embodiments, GPU-1 and GPU-2 may each correspond to a parallel processing unit (PPU) 800, as shown in Fig. 8 is shown.
[0012] The dielets may be interconnected by one or more high-speed interconnects 116 for exchanging commands, data, and / or control. In some embodiments, the high-speed interconnect is an NVLink interface. In an embodiment where the multi-dielet GPU 100 includes two dielets, a high-bandwidth chip-to-chip interface on each dielet may correspond to the one or more high-speed interconnects 116. Embodiments are not limited to the number of dielets included in the multi-dielet GPU or the number of GPUs on each dielet. The multi-dielet GPU may additionally include any number of dielets that do not include GPUs.
[0013] Each dielet includes a plurality of hardware engines. In the illustrated embodiment, dielet-1 includes hardware engine-1 106a and hardware engine-2 108a, and dielet-2 includes hardware engine-1 106b and hardware engine-2 108b. The hardware engines may include compute units, graphics units, encoding / decoding units, encryption / decryption units, etc. As mentioned above, in some embodiments, dielet-1 and dielet-2 may each include a processor such as the PPU in Fig. 8, which includes a variety of general processing clusters (GPCs) (e.g. Fig. 9A), and each GPC includes a plurality of streaming multiprocessors (SM) 940, which are located in the Fig. 9A and Fig. 10A. SMs can execute compute and graphics workloads and can be considered hardware engines. References to "hardware engines" herein include, but are not limited to, those used in graphics and compute applications, video encoding or decoding, units in GPC and SYS clusters, and units with DMA capabilities. Embodiments are not limited to a particular number of hardware engines and / or to one or more types of hardware engines on each GPU.
[0014] Each dielet also has at least one memory management unit (MMU). Similar to CPUs, GPUs use MMUs to translate virtual addresses used by programs or processes into physical addresses in memory. This is because programs typically use a larger virtual address space than the physical memory available on the GPU. The MMU acts as a bridge, mapping these virtual addresses to their corresponding physical memory locations, enabling efficient memory access for various GPU operations. MMUs in GPUs contribute to memory protection by enforcing access permissions. They can control which parts of memory different processes running on the GPU can access. This prevents unauthorized access, data corruption, and potential security vulnerabilities.In addition, MMUs can help manage memory allocation and usage for different tasks running concurrently on the GPU.
[0015] In some embodiments, each dielet may include one or more other MMUs, each configured to service component-level memory-related requests on the dielet. For example, some embodiments may include a GPC MMU for each GPC (see, e.g., Fig. 9A) in the GPU and / or an input / output MMU (IO MMU). The highest-level MMU on each dielet is referred to herein as the hub MMU or "HUBMMU." In some embodiments, the HUBMMU is a GPU MMU ("GMMU") with separate virtual address spaces and page tables for each process. In the illustrated embodiment, HUBMMU-1 110a is located on dielet-1 and HUBMMU-2 is located on dielet-2. Each HUBMMU enables hardware engines on its dielet to access memory by providing virtual address translation, memory protection, etc. All HUBMMUs of the multi-dielet GPU, HUBMMU-1 and HUBMMU-2 in this example, work together to synchronize memory accesses for entities on all dielets and present a single logical MMU to software.
[0016] The ability to provide a monolithic view of the multi-dielet GPU 100 is provided so that external entities, such as driver software 122, running on the CPU 120, which communicates with the multi-dielet GPU 100 via an interface 118 (e.g., a PCI interface), can treat the multi-dielet GPU 100 as a single monolithic GPU. That is, the monolithic GPU view enabled by the circuitry within the multi-dielet GPU 100, including the circuitry of the HUBMMUs, allows software, such as driver software 122, to be agnostic with respect to the specific structure of the multi-dielet GPU 100. US application No. 18 / 606,924, entitled “Method and Apparatus for Supporting Distributed Graphics and Compute Engines and Synchronization in Multi-Dielet-Parallel Processor Architectures,” filed on June 15, 2018, is hereby incorporated by reference.March 202 2024 and incorporated herein by reference in its entirety, describes a hardware implementation in the multi-dielet GPU to map between its dielet-local hardware engine identifiers and globally unique engine identifiers, allowing software to view the multi-dielet GPU as a single monolithic GPU.
[0017] Before describing the mechanisms provided in the embodiments of this disclosure for synchronizing MMUs across all dielets of a multi-dielet GPU, the importance of the MMU can be illustrated by describing its role in a memory access by a hardware engine. Considering an exemplary process executing on a hardware engine (e.g., SM) of Dielet-1 that requires access to an object in memory, the process first generates a virtual address that points to the object's location within GPU memory (local / remote) or CPU memory. This virtual address is then used in the instructions for accessing the object. The process or the corresponding hardware engine sends a request with the generated virtual address to the MMU.The MMU can look up the virtual address in its Translation Lookaside Buffer (TLB), which is a cache that stores recently translated addresses, thereby improving translation speed. If the virtual address is found in the TLB, the corresponding physical address in GPU (local / remote) or CPU memory is retrieved from the TLB entry. This physical address points to the actual location of the sought-after object in GPU (local / remote) or CPU memory. If the virtual address is not found in the TLB, a TLB miss occurs. The MMU then performs a more complex lookup operation using page tables. Page tables are data structures that map virtual addresses to physical addresses for larger memory regions. The MMU accesses the page tables stored in memory to determine the physical address corresponding to the virtual address.This process may involve traversing multiple levels of page tables, depending on the system's memory organization. After completing the page table traversal, the MMU obtains the physical address of the object in GPU memory (local / remote) or CPU memory. The requesting hardware engine obtains the physical address from the MMU and uses it to access the object in GPU memory (local / remote) or CPU memory.
[0018] During the exemplary method described above, the MMU may also determine whether the requesting process and / or hardware engine is authorized to access the requested memory. Unauthorized memory access attempts may cause the MMU to generate an error, which is reported to the software and / or CPU. In this way, the MMU is a key component in the functioning of a GPU.
[0019] Fig. 2 illustrates another multi-dielet GPU 200 according to some embodiments. Fig. 2 schematically illustrates, in one example, a 2-dielet processing system 200 with exemplary communication involving the HUBMMU.
[0020] The multi-dielet GPU 200 includes a dielet-1 202a and a dielet-2 202b. Dielet-1 includes a HUBMMU 210a, a high-speed hub (HSHUB) 214a, and a frame buffer hub (FBHUB) 212a. Dielet-2 includes HUBMMU 210b, HSHUB 214b, and FBHUB 212b.
[0021] The bind table is a data structure that stores information about the mapping between virtual memory regions and physical memory frames assigned to various processes running on the GPU. Each entry in the bind table can contain the range of virtual addresses used by a process, the starting address of the physical memory frame assigned to the corresponding virtual address range, and the access permissions (read, write, execute) assigned to the memory region for protection.
[0022] The Page Table Bind Cache (PDB cache) is a hardware cache located in the HUBMMU. It can store the roots of currently used page tables, from virtual to physical addresses, for various GPU processes. A TLB miss requires retrieving the page table root for that process and performing a page table traversal using the retrieved root.
[0023] When a program attempts to access data using a virtual address, the MMU first checks the page table bind cache. If the translation (mapping between virtual and physical addresses) for that particular address is found in the cache, the MMU can efficiently translate the address and grant access to the corresponding physical memory location. This avoids having to access the main page tables (which are typically slower than the cache), reducing memory access latency and improving overall performance.
[0024] The bind table and / or PDB cache in the HUBMMU 210a may be accessed via the Hub Translation Lookaside Buffer (HUBTLB) or a process that accesses the HUBTLB and results in a TLB miss. Other TLBs, such as the L2 TLB and Link / CE TLB, may access the HUBTLB when misses occur during their respective lookups. The TLB entries may be updated by entries provided by the HUBMMU as a result of the bind table lookup, the PDB cache lookup, or a page table walk. In some embodiments, the TLB entries may be updated by entries provided by the HUBMMU as a result of the bind table and PDB cache lookup and a page table walk.
[0025] Regarding memory requests, memory responses, bind requests and responses, TLB invalidation requests and responses, and various acknowledgments, the HUBMMU on a dielet communicates with other HUBMMUs located on other dielets in the multi-dielet GPU. In some embodiments, this communication may occur via the FBHUB 212a to the other dielets.
[0026] ATS requests, ATS responses, ATS shootdown / invalidations (ATSD), and associated acknowledgments may be communicated between the HUBMMU and the high-speed hub (HSHUB) 214a.
[0027] The communication between the HUBMMU 210b of the second dielet 202b and the various TLBs, HSHUB 214b and FBHUB 212b may be the same or similar as with HSHUB 214a of the first dielet 202a.
[0028] Furthermore, HUBMMU-1 and HUBMMU-2 may be configured to exchange messages. For example, binding requests and acknowledgements, TLB invalidation requests and acknowledgements, error reports, ATS requests, Virtual Address Bus / Breakpoint (VAB) dump requests, etc. In some embodiments, at least some of the messages sent from one HUBMMU to another are transmitted over a connection 224 that is different from the dielet-to-dielet connection (e.g., 116). In some embodiments, connection 224 is a dedicated HUBMMU-to-HUBMMU connection over a dielet-to-dielet connection such as 116.
[0029] In the illustrated 2-dielet processing system 200, Dielet-1 202a and Dielet-2 202b have different types of network interfaces—Dielet-1 has an NVLINK interface 226 and a PCIE interface 228, and Dielet-2 has a chip-to-chip interface 230. In an exemplary embodiment, PCIE interface 228 may be used for memory access, while NVLINK interface 226 and chip-to-chip (C2C) interface 230 may be used for connection to a CPU and other devices (e.g., other GPUs, multi-dielet processing systems, etc.). In some embodiments, Dielet-1 and Dielet-2 may be identical and have identical network interface hardware, but one or more of the network interfaces on each Dielet may be disabled or unused.For example, Dielet-1 and Dielet-2 may both have NVLINK, PCIE, and C2C ports, where the C2C port on Dielet-1 may be inactive and the NVLINK and PCIE interfaces on Dielet-2 may be inactive.
[0030] In the illustrated embodiments of the Fig. 1 and Fig. 2, each dielet contains a HUBMMU. However, in some embodiments, at least some dielets may contain component-level MMUs, such as one or more GPC-level MMUs ("GPCMMUs") and a high-speed hub-level MMU ("HSHUBMMUs"). In some embodiments, all three types of MMU modules (GPCMMUs, HUBMMUs, and HSHUBMMUs) are active in both dielets. GPCMMUs and HSHUBMMUs in one dielet do not communicate with GPCMMUs and HSHUBMMUs in another dielet. Rather, GPCMMUs / HSHUBMMUs in one dielet communicate with HUBMMUs of the same dielet, and then the HUBMMUs in one dielet communicate with the HUBMMUs in the other dielet.
[0031] In some embodiments, communication between HUBMMU and HUBMMU may occur over a private bus. The private bus may be used to keep MMUs in the two dielets synchronized for bind operations, TLB invalidations, error reporting, security vector updates, and ATS requests. Although the MMU private bus is used for synchronization commands as noted above, ATS responses, ACK communications, other memory commands, and data transfers (for GMMU page table walk-through memory access data, error packets, and VAB data dumps to memory) may be sent over the FBHUB of the corresponding dielet. This may provide for faster synchronization of the HUBMMU.
[0032] Thus, in some embodiments, the HUBMMU is active on both dielets of a two-dielet processing system, with minimal communication between them. The operations of the exemplary dielet MMU may include multiple message types and request / response types.
[0033] All engines and other clients on a dielet are served by the MMUs on that dielet. The HUBMMU in a dielet independently performs a page table walk, without requiring information from HUBMMUs of other dielets. The TLB in a dielet does not communicate with TLBs in other dielets. For example, in some embodiments, each dielet may have a HUBTLB and one or more component-level TLBs (e.g., uTLB, gpcL1TLB, gpcL2TLB, and LinkTLB), but none of the TLBs on a dielet communicate with any TLB on other dielets.
[0034] Engines and other clients send binding requests and TLB invalidation requests to the HUBMMU of the same dielet. HUBMMUs communicate with each other to keep MMUs synchronized, but both dielets independently perform the binding operation in parallel to generate a combined confirmation for the requester.
[0035] In some embodiments, the bind operation is always initiated by the primary dielet. The secondary dielet(s) forwards its local bind request to the primary dielet. The TLB invalidation operation can also always be initiated by the primary dielet. The secondary dielet(s) forwards its local invalidation requests to the primary dielet. Error reporting is always initiated by the primary dielet. Errors from the secondary dielet(s) are forwarded to the primary dielet. A VAB memory dump is always initiated by the primary dielet. VAB dump requests from the secondary dielet(s) are forwarded to the primary dielet. ATS / ATSD requests, responses, and acknowledgments (with HSHUB) are always performed by a secondary dielet.The primary dielet sends its ATS requests to the HUBMMU of the designated secondary dielet for forwarding to the HSHUB (ultimately to the System MMU (SMMU)).
[0036] The PDB cache in a dielet is not a mirror of the PDB caches of other dielets and operates independently when processing bind commands, page table walks, and invalidation. For bind commands, the PDB ID associated with the same PDB may be different on different dielets.
[0037] As described above, certain MMU operations are performed by only one dielet, and the MMU of other dielets operates synchronously with the executing dielet for such operations. Furthermore, in at least some embodiments, each dielet may be configured to operate as an independent GPU (e.g., for production and testing). Thus, embodiments of the present disclosure contemplate that each dielet in a multi-dielet processing system knows its dielet identifier and operational role.
[0038] For example, each HUBMMU can determine whether it is configured as a primary dielet, secondary dielet, or standalone dielet based on the configuration settings of one or more fuses and / or registers. Each HUBMMU can also determine its dielet identifier based on the settings of one or more fuses and / or registers. The dielet identifier is used to uniquely identify a dielet within the multi-dielet processing system.
[0039] In some embodiments, communication from HUBMMU to HUBMMU may occur over a connection through a dielet-to-dielet connection crossbar. The HUBMMU-to-HUBMMU connection may be configured to transmit packets for a specific set of operations: an MMU binding (e.g., initial binding packet from client(s)), a TLB invalidate, an ATS request (ATR), an ATS invalidate (ATSD), an ATS response (ATRsp) data packet from secondary to primary, an error packet (internal), a VAB dump request, a VAB (internal) mask data packet, and acknowledgments of the above operations. The connection between HUBMMU and HUBMMU may be reserved exclusively for the above-mentioned packets. In some embodiments, the connection may be shared with other components, such as the FBHUB, to allow these components to also transmit selected types of packets.
[0040] In some embodiments, the HUBMMU includes a buffer to queue packets to be transmitted and / or received packets. In some embodiments, the HUBMMU may implement a blocking thread and a non-blocking thread to transmit messages. For example, the blocking thread may transmit a bind, a TLB invalidation, a blocking fault, a VAB, an ATS, and associated acknowledgments, and the non-blocking thread may be used to transmit non-blocking faults and associated acknowledgments.
[0041] Fig. Figure 3 schematically illustrates an example of a hierarchical TLB arrangement in each dielet of a 2-dielet processing system according to some embodiments. In the 2-dielet processing system of Fig. 3, Dielet-1 302a is configured as the primary dielet, and Dielet-2 302b is configured as the secondary dielet.
[0042] The primary dielet 302a includes a HUBMMU (HUBMMU-1) 310a that communicates with an FBHUB (FBHUB-1) 312a and an HSHUB (HSHUB-1) 314a, which are also located in the dielet. The HUBMMU-1 310a includes a fill unit 332a configured to communicate with a dielet TLB hierarchy 334a. The fill unit 332a is also configured to communicate with the fill unit of the HUBMMU (HUBMMU-2) 310b in dielet-2 302b via an interface 324. The fill unit 332a may include a PDB cache and a bind table. It may also include an error state machine and / or an ATS wrapper function.
[0043] Dielet TLB hierarchy 334a may include a hub-level TLB and a Level 1 (L1) TLB for each GPC on the dielet to accelerate access to the respective L1 caches. Dielet TLB hierarchy 334a may also include one or more Level 2 (L2) TLBs for one or more of the GPCs on the dielet and / or L1 TLBs for respective links (e.g., network interfaces). TLB invalidations from the HUBMMU are communicated to the respective local (e.g., component-level) TLBs of the dielet TLB hierarchy, and the respective local TLBs return corresponding acknowledgments.
[0044] Dielet-2 302b may have an identical configuration or design of the TLB hierarchy, including dielet TLB hierarchy 334b. The design of the HUBMMU 310b, the fill unit 332b, the FBHUB 312b, and the HSHUB 314b may be identical or similar to the corresponding components in dielet-1 302a.
[0045] Fig. Figure 4 shows a flow chart 400 for a binding process in a multi-dielet processing system as described in the Fig. 1-3, according to some embodiments of the present disclosure. In the illustrated configuration, Dielet-1 is configured as the primary dielet and Dielet-2 is configured as the secondary dielet. The multi-dielet processing system is configured such that the primary dielet performs the bind operation.
[0046] The bind operation links an engine (e.g., graphics engine, Copy / DMA (CE) engine, security / integrity engine, etc.) to an instance block in the HUBMMU so that the HUBMMU can translate subsequent memory requests from that engine (or a client of that engine). The instance block contains the corresponding page table pointer.
[0047] In some embodiments, the HUBMMU does not distinguish which engines may be connected to which dielet, although in some cases only VEIDs (Graphics Engine), Base Address Register (BAR) engines, CEs, and PMA engines are expected to be accessed in both dielets.
[0048] Each dielet's HUBMMU can have a bind table and a PDB cache. The bind table and PDB cache in each HUBMMU can be configured (e.g., resized) to support all engines in all dielets in a multi-dielet GPU. However, both the bind table and PDB cache in one dielet are designed to operate independently of the bind tables and PDB caches in other dielets. Therefore, the same PDB ID can point to different page tables in different dielets. The bind table entries can be indexed by engine IDs, so that the same indexed entry contains the same bind information in all dielets.
[0049] The primary dielet is configured to initiate the binding process. The secondary dielets forward binding packets to the primary dielet to start the binding process.
[0050] When a client (e.g., a hardware engine) on Dielet-1 or Dielet-2 issues a binding request (e.g., 402a and 402b), it is sent to the HUBMMU on the same Dielet. On Dielet-1, the HUBMMU (HUBMMU-1) is configured to process binding requests because Dielet-1 is configured as the primary Dielet. However, when the HUBMMU (HUBMMU-2) on Dielet-2 receives the binding request 402b from a local client, it forwards the binding request to HUBMMU-1 (402c), because Dielet-2 is configured as the secondary Dielet.
[0051] The binding request packet can be generated either by the engine or as a result of the engine writing a register (e.g., a privilege register).
[0052] Since the HUBMMU-1 receives binding requests from all other dielets in the multi-dielet processing system, it serializes (404) the binding requests. For example, incoming binding requests can be stored in a queue.
[0053] The HUBMMU-1 may then mediate or arbitrate (406) among the serialized binding requests to select a binding request to be processed.
[0054] At 408, the selected bind request is processed. Processing the bind request may include allocating a physical memory region to the requesting client and determining the virtual address for the allocated region. A new entry may be added to the bind table, or an existing corresponding entry may be updated in the bind table. An entry may contain a mapping between the virtual address and the allocated memory, as well as associated permissions.
[0055] At 410, the HUBMMU-1 transmits the binding information received at 408, such as the contents of the entry in the bind table, to the other dielets. In some embodiments, the binding information may include the engine identifier or engine ID of the sender of the bind request, the PDB cache information, and the mapped instance block pointer. Otherwise, the bind table and the PDB cache may become out of sync.
[0056] At 412, the HUBMMU-2 receives the binding information from the HUBMMU-1 and updates its own bind table accordingly. For example, the HUBMMU-2 adds an entry to its bind table.
[0057] After updating its bind table, each dielet updates its local TLB at 414 (414a and 414b), and the respective TLBs confirm the updates at 416 (416a and 416b). The TLB update can be initiated by the HUBMMU, which propagates the binding information to all relevant TLBs on the same dielet.
[0058] After HUBMMU-1 receives acknowledgments from its local TLB updates and other dielets (Dielet-2 returns an acknowledgment 418), HUBMMU-1 notifies the sender of the bind request that the bind request is complete and returns the binding information to the requester. For example, if the selected bind request originated from Dielet-1, acknowledgment 420a, which includes the virtual address of the allocated memory and optionally other parameters, is sent to the requesting local client. If the selected bind request originated from Dielet-2, acknowledgment 420b, which includes the virtual address of the allocated memory and optionally other parameters, is sent to the HUBMMU on the requesting dielet, and the HUBMMU on the requesting dielet forwards the acknowledgment and the virtual address, etc., to the requesting local client.For example, HUBMMU-1 transmits the acknowledgment 420b to HUBMMU-2, which then sends it (422c) to the requesting local client on Dielet-2.
[0059] The bind operation is often the first prerequisite before work is assigned to a process. It establishes the link between the memory system, in this case the page table, and a context and the engine executing the work. After the engine initiates a bind and an acknowledgment is received, the engine can be assigned work by the work dispatcher.
[0060] This design ensures that with respect to a bind request, only the primary dielet accesses memory to serve a bind request, regardless of which dielet originates the request.
[0061] This procedure includes handshaking that occurs locally on each dielet to ensure that the local TLBs are synchronized at the component level, and handshaking that occurs between the dielets to ensure that they are synchronized.
[0062] Fig. 5 shows a flowchart 500 for a TLB invalidation in a multi-dielet GPU, as described in the Fig. 1-3, according to some embodiments of the present disclosure.
[0063] TLB invalidation is a process of removing stale entries from the translation lookaside buffers (TLBs) in a multi-dielet GPU. TLB invalidation can be triggered by changes in the mapping between virtual and physical addresses (e.g., memory writes, page table updates), context switches, address space changes, etc.
[0064] TLB invalidation can be triggered by an MMU, an engine or client, or by software.
[0065] In the illustrated configuration, Dielet-1 is set as the primary dielet and Dielet-2 as the secondary dielet. The multi-dielet processing system is configured so that the primary dielet implements the TLB invalidation operation.
[0066] TLB invalidation invalidates previously fetched translations (e.g., TLB rows) either from the GMMU page table walk or from the ATS (IOMMU / SMMU). However, TLB invalidation does not invalidate page table entries (PTEs) or binding information.
[0067] TLB invalidation can originate from the hardware engines or clients on the dielets, or be generated based on register writes (e.g., privileged register writes). Requesting a TLB invalidation in one dielet invalidates TLB lines in all dielets. However, the dielets are invalidated independently and in parallel.
[0068] A TLB invalidation packet may contain the following information, which is sent over the interface to other HUBMMUs: virtual address, PDB entry, whether the invalidation applies to an ATS entry, the originating hub ID, and the request ID.
[0069] In the illustrated embodiment, the TLB invalidation request is issued by an engine or another client on Dielet-1. The engine or client sends the request 502 to the local HUBMMU - HUBMMU-1. The request can originate from an engine, a client on an engine, a front-end context switch (FECS), a GPU system processor (GSP), a high-speed hub (HSHUB), or be generated by writing to a privilege register. For example, an engine or client can send a method-based invalidation request, the HSHUB can send an ATS invalidation request (ATSD), an FECS / GSP can send a uCode-based TLB invalidation request, or a HUBMMU TLB invalidation privilege register can be used to generate a PRI-based TLB invalidation.
[0070] At 504, the HUBMMU-1 transmits the invalidation request to all other dielets in the multi-dielet processing system.
[0071] At 506, HUBMMU-1 and HUBMMU-2 may process TLB invalidation in parallel (e.g., 506a and 506b). If the TLB invalidation information indicates that a system memory lock operation is required, the source MMU (in this example, HUBMMU-1) performs an I / O flush operation at the end of the invalidation operation. Each HUBMMU independently retrieves the PDB ID to be invalidated from the PDB cache based on the input PDB, so the PDB ID may be different for the same PDB.
[0072] At 508, the HUBMMU-2 returns an acknowledgment to the HUBMMU-1 indicating that the TLB invalidation on the Dielet-2 is complete.
[0073] In another example, if the TLB invalidation request originates from Dielet-2, it is sent by the requesting client to HUBMMU-2, and then forwarded by HUBMMU-2 to HUBMMU-1. This is because HUBMMU-1 is configured as the primary dielet to process TLB invalidations, and HUBMMU-2 is configured as the secondary dielet to forward TLB invalidations to the primary dielet for processing.
[0074] After the HUBMMU-1 has received acknowledgments from its local TLB invalidation and other dielets, it returns an acknowledgment to the requesting client at 510.
[0075] Fig. Figure 6A shows a flowchart 600 for a page table walk in a GPU with multiple dielets, as described in the Fig. 1-3, according to some embodiments. In the illustrated configuration, Dielet-1 is configured as the primary dielet and Dielet-2 is configured as the secondary dielet. The multi-dielet processing system is configured such that each dielet performs page table traversal for TLB misses independently of the local TLBs.
[0076] Each HUBMMU can have its own page directory entry (PDE) / page table entry (PTE) caches to accelerate the page table walk of the GPU MMU (GMMU). The GMMU page table walk is also referred to herein as the HUBMMU page table walk. Upon a local TLB miss, the HUBMMU, or more precisely, a fill unit within the HUBMMU, performs a GMMU page table walk based on the engine's bind point found in its own bind table and PDB cache.
[0077] At 602, a hardware engine issues a memory request to the local HUBMMU (e.g., HUBMMU-1 on Dielet-1, HUBMMU-2 on Dielet-2).
[0078] At 604, the local HUBMMU performs a TLB lookup and an error occurs.
[0079] At 606, the local HUBMMU performs a GMMU page table walk. An error occurs.
[0080] At 608, the local HUBMMU transmits a page table walk memory fetch via its local FBHUB. The page table walk memory fetch request is sent to the local FBHUB. The FBHUB sends the request via the memory crossbar to access the memory. The target memory can be in the local video memory connected to all dielets or in the system memory. The system memory can be connected to the multi-dielet processing system via a chip-to-chip (C2C) interface or a PCIe interface. The memory fetch response takes the reverse path.
[0081] At 610, the FBHUB retrieves the memory and receives the requested memory at 612.
[0082] At 614, the FBHUB returns the memory response to the local HUBMMU.
[0083] At 616, the HUBMMU updates tables based on the obtained memory. For example, the HUBMMU may update local TLBs by returning the page table entries (or information therefrom) to the local TLBs that contain the HUBTLB.
[0084] At 618, a memory response is returned to the requesting engine.
[0085] It should be noted that in some cases, the GMMU page table walk may require an Address Translation Service (ATS) call. In this case, the ATS request is processed like any other ATS request in the multi-dielet processing system.
[0086] In many systems, the GPU is provided with data from the host through one of the numerous memory management API calls provided by the CUDA framework, such as CudaMallocManaged and CudaMemCpy. Some systems are able to avoid using CUDA calls for memory management and access the same data on both the GPU and the CPU. This can be achieved through Address Translation Services (ATS) technology, which provides a unified virtual address space for data allocated with malloc and new. ATS allows the CPU and GPU to share a single per-process page table, allowing all CPU and GPU threads to access all system-allocated memory, which can be located on either CPU or GPU physical memory. The CPU heap, CPU thread stack, global variables, memory-mapped files, and interprocess memory are accessible to all CPU and GPU processes.
[0087] In some embodiments, the multi-dielet GPU is connected via the C2C interface (see, for example, Fig. 2) connected to a CPU, which in some implementations may have hardware-based memory coherence, allowing the transfer of only the required data rather than the migration of entire pages to and from the GPU. Simple synchronization primitives across GPU and CPU threads may also be enabled by enabling native atomic operations from both the CPU and GPU. ATS may, in some embodiments, utilize direct memory access (DMA) copy engines on the GPU to accelerate bulk transfers of pageable memory between a host and a GPU. Such embodiments may enable applications to overcommit the GPU's memory and directly utilize CPU (system) memory with high bandwidth. Access to such large amounts of memory also facilitates high-performance computing, graphics, virtual reality, and AI applications.
[0088] ATS can be used with a virtual address or from a physical system address. Two types of ATS can be used: it can either translate a virtual address to a physical address, or it can translate a virtual GPU address to a physical GPU address and then translate the physical GPU address to a physical system address.
[0089] In some embodiments, an ATS request may be required to be generated as part of the page table walk. The request may be generated by a HUBMMU during a page table walk. In some embodiments, each dielet may perform the page table walk independently and issue the ATS request at a specific level. The type of ATS request may be based on the page table walk level at which the request is issued.
[0090] In Fig. Figure 2 shows that both dielets support ATS requests (ATR) / ATS responses (ATRsp) between the HUBMMU (e.g., the fill unit of the HUBMMU) and the HSHUB. However, in embodiments, the Dielet-1 capability may remain unused. Instead, Dielet-1 forwards its ATR to Dielet-2.
[0091] For the Dielet-2, the HUBMMU filler unit can accept an ATR from the Dielet-1 or the Dielet-2 up to a preconfigured maximum (e.g., 256) and serialize the ATR. ATRs can be sent up to a predefined maximum (e.g., 256 x 2 = 512) on the HUBMMU-HSHUB interface (or, in some implementations, an interface between the HUBMMU filler unit and the HSHUB) on the Dielet-2, from where they are forwarded to the high-speed interface (e.g., C2C) configured to connect to the CPU. The corresponding acknowledgments are received by the Dielet-2 and sent back to the Dielet-1 for requests originating from the Dielet-1.
[0092] Note that ATR / ATRsp propagation is centralized (e.g., at Dielet-2), but generation occurs independently and in parallel at each Dielet. Note that another possibility is for each fill unit to send its ATR to its respective HSHUBs and then allow HSHUB-1 (instead of the fill unit of HUBMMU-1) to forward its ATR to Dielet-2 for transmission over the C2C interface. This design enables efficient processing from an ATR, since, in some embodiments, ATS communication occurs primarily over the C2C interface, and, in at least some embodiments, the only C2C interface of the multi-Dielet processing system resides on Dielet-2.
[0093] The Dielet-1 ATRs are forwarded to Dielet-2. Then, the HUBMMU (e.g., the HUBMMU's fill unit) on Dielet-2 sends the ATR (e.g., an ATR from Dielet-1 or Dielet-2) to the Dielet-2's HSHUB, to the C2C, to the System MMU (SMMU). The SMMU processes the ATR and returns a corresponding ATRsp to the Dielet-2's HSHUB.
[0094] The HSHUB of the Dielet-2 sends the ATRsp to the Dielet-2's HUBMMU (e.g., the HUBMMU's fill unit). A target ID in the ATR and the ATRsp can be used to identify the target. The ATRsp can also contain one or more Dielet IDs, ATS cache IDs, cache line IDs, and an entry ID. These parameters can be used by the source to subsequently update its local tables.
[0095] Fig. 6B shows an exemplary flowchart 610 of a method for ATS requests and responses in a multi-dielet GPU according to some embodiments of this disclosure. As mentioned above, ATS requests may be generated by the HUBMMUs on each dielet during a page table walk. The multi-dielet GPU may be configured such that, although the generation of ATRs is distributed across all dielets, the related communication with other devices (e.g., CPU, SMMU, IOMMU) and the reception of the corresponding response is centralized on a dielet configured to be in the secondary role. In the Fig. The scenario shown in Figure 6B is in a 2-dielet GPU like the one in one of the Fig. 1-3, according to one embodiment, the dielet-1 is configured as the primary dielet and the dielet-2 as the secondary dielet. As with respect to Fig. 2, the C2C interface is active on Dielet-2 and unavailable on Dielet-1 (i.e., absent or set to inactive). The described scenario applies to an ATR generated by the HUBMMU of Dielet-1 (HUBMMU-1).
[0096] At 612, HUBMMU-1 generates an ATR. This may be the result of a page table walk, as described above. Since the multi-dielet GPU is configured so that Dielet-2 serves as the serialization point for ATR / ATRsp, HUBMMU-1 forwards the ATR to Dielet-2.
[0097] At 614, the HUBMMU receives the ATR from the HUBMMU-1 from the Dielet-2 (HUBMMU-2) and serializes it with all other ATS requests (e.g., from the Dielet-1 or the Dielet-2). It then arbitrates between the serialized ATSs and sends the ATR to its local HSHUB.
[0098] At 616, the HSHUB on the Dielet-2 (HSHUB-2) sends the ATR over the C2C interface to any CPU, SMMU, etc.
[0099] At 618, the corresponding ATRsp is received via the C2C interface on the Dielet-2 and forwarded to the local HSHUB-2.
[0100] At 620, the HSHUB-2 forwards the ATRsp to the local HUBMMU-2.
[0101] At 622, the HUBMMU-2 forwards the ATRsp to the HUBMMU-1.
[0102] Based on the ATRsp, the HUBMMU-1 can update the local TLB(s) with the information contained in the ATRsp.
[0103] The ATSD invalidation is identical in terms of message flow to that of Fig. 5. ATS invalidation invalidates TLB entries marked as ATS entries. ATSD invalidation requests and responses are handled via the HSHUB-2, just like TLB invalidations. An ATSD invalidation request is sent from the SMMU to the HSHUB-2 and forwarded to the HUBMMU-2. The HUBMMU-2 treats it similarly to a locally generated TLB invalidation and forwards the invalidation request to the HUBMMU-1. The HUBMMU-1 initiates the corresponding TLB updates locally and notifies other dielets to perform a TLB update in parallel. Then, after the TLB updates have been acknowledged locally and by all other dielets, the HUBMMU-1 transmits the acknowledgment on the ATSD invalidation to the HUBMMU-2, which sends it to the SMMU via the HSHUB-2.
[0104] As noted above, in some embodiments, the ATS request is sent to the IOMMU / SMMU via a secondary dielet in the multi-dielet processing system. For example, in one embodiment, a 2-dielet processing system such as the one shown in Fig. 1-3, the ATS request is sent through the HSHUB of the secondary dielet because it is the dielet that has an active C2C interface over which ATS communication is performed in the particular embodiment. In some embodiments, PCIe ATS requests, if they are to be sent, may be sent through the primary dielet that has the active PCIe interface. Each HUBMMU may have an ATS wrapper function to forward ATS messages.
[0105] Errors are detected independently of each dielet's MMUs (e.g., TLBs and fill unit in the HUBMMU). However, the primary dielet's HUBMMU is designed to report errors regardless of where (e.g., on which dielet) the error originated. When an error occurs in the secondary dielet, the secondary dielet forwards the error information (e.g., error type / ID, error parameters, engine / client of the source of the faulty command / request, primary dielet ID, etc.) to the primary dielet, which then processes this information to create an error packet.
[0106] The error packet can then be written to memory if an error buffer exists, written to a register, or reported by other means. As mentioned above, the error information may include an identifier (ID) of the engine or other component that caused the error. This makes it easier for the CPU and / or other error handler to initiate operations to respond to the error.
[0107] Fig. Figure 6C shows a flowchart 630 for an error reporting method that may be used in a multi-dielet GPU, such as the 2-dielet GPU in one of the Fig. 1-3, according to some embodiments of the present disclosure.
[0108] At 632, an engine or a process on an engine in the dielet-2 sends a memory request to its local HUBMMU (HUBMMU-2).
[0109] At 634, the HUBMMU-2 detects an error (e.g., a memory access violation). Each HUBMMU can implement an error state machine to detect errors on the local dielet.
[0110] At 636, the HUBMMU-2 reports the error to the HUBMMU on board 1 (HUBMMU-1).
[0111] At 638, the HUBMMU-1 serializes the error reports and notifies the CPU, SMMU, or another error handler of the error. The error handler can then respond to the error using the information received in the error report.
[0112] A Vidmem Access Bit (VAB) is a technique used to track memory segments that have been accessed or modified since tracking was enabled or the last VAB dump and clear was performed at a specific time. It focuses on the physical addresses used by the program and provides valuable information for paging, debugging, or VM swapping. A VAB dump captures the system state at a specific time, usually triggered by an event such as a write to a privileged register or a channel method.
[0113] In some embodiments, the VAB is dumped by the HUBMMU that receives the VAB dump request (e.g., like TLB invalidation). Each dielet independently tracks the VAB in its VAB mask (e.g., 4K-bit VAB). In response to a VAB dump request, the following steps may be performed sequentially.
[0114] For example, a host on the primary dielet sends a VAB dump packet to the local HUBMMU. The host sends the VAB dump request as part of a TLB invalidation packet. However, no invalidation is performed. A field in the invalidation packet can be used to identify that the invalidation request is for a VAB dump.
[0115] The primary HUBMMU sends the VAB dump packet to the secondary HUBMMU (and vice versa if it is received at a secondary dielet).
[0116] Both HUBMMUs collect VAB independently from different tracking copies in the dielet.
[0117] The final merged VAB mask is stored in the HUBMMU. The secondary HUBMMU sends the final merged VAB mask to the primary HUBMMU in multiple blocks.
[0118] The primary HUBMMU can merge the VAB mask from the secondary HUBMMU, and the merged VAB mask is written to memory.
[0119] Fig. 7 shows an example of a GPU 700 on a dielet of the multi-dielet GPU, whose frame buffer hub (FBHUB) and high-speed hub (HSHUB) are identified with respect to some of their connections, according to some embodiments of the present disclosure. The GPU 700 includes two processors 705-1 and 705-2, each processor including a plurality of processing units 710 connected to a cache memory 715. Respective crossbars 712-1 and 712-2 connect processing units 710 and a memory 715 in processors 705-1 and 705-2. A high-speed hub (HSHUB) 718 connects all processing units 710 and the cache memory 715. A framebuffer hub (FBHUB) 720 enables the processing units 710, which are connected to the FBHUB 720 via the HSHUB 718, to access the system and / or host memory 722. Example of a GPU architecture
[0120] The following is an example architecture of a dielet in a multi-die GPU with reference to the Fig. 1-6. The following information is for illustrative purposes only and should not be construed as limiting in any way. Each of the following features may be included optionally, with or without excluding other described features.
[0121] Fig. Figure 8 illustrates, according to one embodiment, a parallel processing unit (PPU) 800 that may be present on a dielet of a multi-dielet GPU. In one embodiment, the PPU 800 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 800 is a latency-hiding architecture designed for parallel processing of many threads. A thread (e.g., a thread of execution) is an instantiation of a set of instructions configured for execution by the PPU 800. In one embodiment, the PPU 800 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a device such as a liquid crystal display (LCD).In other implementations, the PPU 800 may be used to perform general-purpose computations. In some other embodiments, the PPU 800 is configured to implement large neural networks in deep learning or other high-performance computing applications.
[0122] One or more PPUs 800 can be configured to accelerate thousands of high-performance computing (HPC), data center, and machine learning applications. The PPU 800 can be configured to accelerate numerous deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, personalized user recommendations, and the like.
[0123] As in Fig. 8, the PPU 800 includes an input / output (I / O) unit 805, a front-end unit 815, a scheduler unit 820, a work distribution unit 825, a hub 830, a crossbar (Xbar) 870, one or more general processing clusters (GPCs) 850, and one or more partition units 880. The PPU 800 may be connected to a host processor or other PPUs 800 via one or more high-speed NVLink 810 interconnects. The PPU 800 may be connected to a host processor or other peripheral devices via interconnect 802. The PPU 800 may also be connected to memory, including a number of memory devices 804. In one embodiment, the memory 804 may include a number of dynamic random access memory (DRAM) devices. The DRAM devices may be configured as a high-bandwidth memory (HBM) subsystem, with multiple DRAM chips stacked on top of each other in each device.
[0124] The NVLink 810 interconnect allows systems to scale and include one or more PPUs 800 in combination with one or more CPUs, supports cache coherence between the PPUs 800 and the CPUs, and supports CPU mastering. Data and / or commands can be transferred via the NVLink 810 through the hub 830 to / from other units of the PPU 800, such as one or more copy machines, a video encoder, a video decoder, a power management unit, etc. (not explicitly shown). The NVLink 810 is used in conjunction with Fig. 11A and Fig. 11B is described in more detail.
[0125] The I / O unit 805 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) over interconnect 802. The I / O unit 805 may communicate with the host processor directly over interconnect 802 or through one or more intermediary devices, such as a memory bridge. In one embodiment, the I / O unit 805 may communicate with one or more other processors, such as one or more PPUs 800, over interconnect 802. In one embodiment, the I / O unit 805 implements a Peripheral Component Interconnect Express (PCIe) interface for communicating over a PCIe bus, and the interconnect 802 is a PCIe bus. In other implementations, the I / O unit 805 may also implement other types of known interfaces for communicating with external devices.
[0126] The I / O unit 805 decodes packets received over connection 802. In one embodiment, the packets represent commands configured to cause the PPU 800 to perform various operations. The I / O unit 805 transmits the decoded commands to various other units of the PPU 800, as the commands specify. For example, some commands may be transmitted to the front-end unit 815. Other commands may be transmitted to the hub 830 or other units of the PPU 800, such as one or more copy machines, a video encoder, a video decoder, a power management unit, etc. (not explicitly shown). In other words, the I / O unit 805 is configured to route communication between and among the various logical units of the PPU 800.
[0127] In one embodiment, a program executed by the host processor encodes an instruction stream in a buffer that provides workloads to the PPU 800 for processing. A workload may include multiple instructions and data to be processed by those instructions. The buffer is an area in memory that both the host processor and the PPU 800 can access (e.g., read and write). For example, the I / O unit 805 may be configured to access the buffer in system memory connected to the connection 802 via memory requests transmitted over the connection 802. In one embodiment, the host processor writes the instruction stream to the buffer and then transmits a pointer to the beginning of the instruction stream to the PPU 800. The front-end unit 815 receives pointers to one or more instruction streams.The front-end unit 815 manages the one or more streams, reads instructions from the streams, and forwards instructions to the various units of the PPU 800.
[0128] The front-end unit 815 is coupled to a scheduler unit 820, which configures the various GPCs 850 to process tasks defined by the one or more streams. The scheduler unit 820 is configured to track status information related to the various tasks managed by the scheduler unit 820. The status may indicate which GPC 850 a task is assigned to, whether the task is active or inactive, what priority level is assigned to the task, etc. The scheduler unit 820 manages the execution of a plurality of tasks on the one or more GPCs 850.
[0129] The scheduler unit 820 is coupled to a work distribution unit 825 configured to distribute tasks for execution among the GPCs 850. The work distribution unit 825 may track a number of scheduled tasks received from the scheduler unit 820. In one embodiment, the work distribution unit 825 manages a pending task pool and an active task pool for each of the GPCs 850. The pending task pool may include a number of slots (e.g., 32 slots) containing tasks assigned to a particular GPC 850 for processing. The active task pool may include a number of slots (e.g., 4 slots) for tasks being actively processed by the GPCs 850. When a GPC 850 finishes executing a task, that task is removed from the GPC 850's active task pool and one of the other tasks is selected from the pending task pool and scheduled to run on the GPC 850.If an active task on the GPC 850 has been idle, for example, while waiting for a data dependency to be resolved, the active task can be removed from the GPC 850 and returned to the pending task pool while another task is selected from the pending task pool and scheduled to execute on the GPC 850.
[0130] The work distribution unit 825 communicates with one or more GPCs 850 via the XBar 870. The XBar 870 is an interconnection network that couples many of the units of the PPU 800 to other units of the PPU 800. For example, the XBar 870 may be configured to couple the work distribution unit 825 to a specific GPC 850. Although not explicitly shown, one or more other units of the PPU 800 may also be connected to the XBar 870 via the hub 830.
[0131] The tasks are managed by the scheduler unit 820 and forwarded by the work distribution unit 825 to a GPC 850. The GPC 850 is configured to process the task and generate results. The results may be consumed by other tasks within the GPC 850, forwarded to another GPC 850 via the XBar 870, or stored in memory 804. The results may be written to memory 804 via the partition units 880, which implement a memory interface for reading and writing data to / from memory 804. The results may be transferred to another PPU 804 or CPU via the NVLink 810. In one embodiment, the PPU 800 has a number U of partition units 880, which corresponds to the number of separate and distinct memory devices 804 coupled to the PPU 800. A partition unit 880 is described below in connection with Fig. 9B described in more detail.
[0132] In one embodiment, a host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 800. In one embodiment, multiple computing applications are executed concurrently by the PPU 800, and the PPU 800 provides isolation, quality of service (QoS), and independent address spaces for the multiple computing applications. An application can generate instructions (e.g., API calls) that cause the driver core to generate one or more tasks for execution by the PPU 800. The driver core issues tasks to one or more streams that are processed by the PPU 800. Each task can comprise one or more groups of related threads, referred to herein as a warp. In one embodiment, a warp comprises 32 related threads that can execute in parallel.Cooperating threads may refer to a plurality of threads that have instructions to execute the task and can exchange data via a shared memory (SMEM). Threads, cooperating threads, and a hierarchical grouping of threads such as cooperating thread arrays (CTAs) and cooperating group arrays (CGAs), according to some embodiments, are further described in U.S. application Ser. No. 17 / 691,621, the entire contents of which are hereby incorporated by reference. The SMEM, in some embodiments, is described in U.S. application Ser. No. 17 / 691,690, which is hereby incorporated by reference in its entirety.
[0133] Fig. 9A shows a GPC 850 of the PPU 800 from Fig. 8 according to one embodiment. As in Fig. 9A, each GPC 850 includes a number of hardware units for processing tasks. In one embodiment, each GPC 850 includes a pipeline manager 910, a pre-raster operations unit (PROP) 915, a raster engine 925, a work distribution crossbar (WDX) 980, a memory management unit (MMU) 990, and one or more data processing clusters (DPCs) 920. It is understood that the GPC 850 of the Fig. 9A instead of the Fig. 9A or may have other hardware units in addition to them.
[0134] In one embodiment, the operation of the GPC 850 is controlled by the pipeline manager 910. The pipeline manager 910 manages the configuration of the one or more DPCs 920 for processing tasks assigned to the GPC 850. In one embodiment, the pipeline manager 910 may configure at least one of the one or more DPCs 920 to implement at least a portion of a graphics rendering pipeline, a neural network, and / or a compute pipeline. For example, with respect to a graphics rendering pipeline, a DPC 920 may be configured to execute a vertex shader program on the programmable streaming multiprocessor (SM) 940. The pipeline manager 910 may also be configured to forward packets received from the work distribution unit 825 to the appropriate logical units within the GPC 850.For example, some packets may be forwarded to fixed function hardware units in the PROP 915 and / or the raster engine 925, while other packets may be forwarded to the DPCs 920 for processing by the primitive engine 935 or the SM 940.
[0135] The PROP unit 915 is configured to forward the data generated by the raster engine 925 and the DPCs 920 to a Raster Operations (ROP) unit, which in conjunction with Fig. 9B. The PROP unit 915 may also be configured to perform color mixing optimizations, organize pixel data, perform address translations, and the like.
[0136] Each DPC 920 included in the GPC 850 includes an M-Pipe Controller (MPC) 930, a Primitive Engine 935, and one or more SMs 940. The MPC 930 controls the operation of the DPC 920 by forwarding packets received from the Pipeline Manager 910 to the appropriate units in the DPC 920. For example, packets associated with a vertex may be forwarded to the Primitive Engine 935, which is configured to retrieve vertex attributes associated with the vertex from memory 804. In contrast, packets associated with a shader program may be transferred to the SM 940.
[0137] The SM 940 includes a programmable streaming processor configured to process tasks represented by a number of threads. Each SM 940 has multiple threads and is configured to execute a plurality of threads (e.g., 32 threads) from a specific group of threads concurrently. In one embodiment, the SM 940 implements a single-instruction, multiple-thread (SIMT) architecture, where each thread in a group of threads (e.g., a warp) is configured to process a different set of instructions based on the same set of instructions. All threads in the group of threads execute the same instructions.In another embodiment, the SM 940 implements a SIMT (Single-Instruction, Multiple Thread) architecture, where each thread in a group of threads is configured to process a different set of instructions based on the same set of instructions, but individual threads within the group of threads are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each warp, enabling concurrency between warps and serial execution within warps when threads diverge within the warp. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, enabling equal concurrency among all threads within and between warps.By maintaining execution state for each individual thread, threads executing the same instructions can be merged and executed in parallel for maximum efficiency. The SM 940 is discussed below in conjunction with . Fig. 10A is described in more detail. Fig. 10B conceptually illustrates, according to some embodiments, four sub-partitions 1091-1094 that may be stored in an SM such as the one shown in Fig. 10A shown SM, are implemented.
[0138] The MMU 990 provides an interface between the GPC 850 and the partition unit 880. The MMU 990 can handle virtual address to physical address translation, memory protection, and memory request arbitration. In one embodiment, the MMU 990 provides one or more translation lookaside buffers (TLBs) to perform virtual address to physical address translation in memory 804.
[0139] Fig. 9B shows a memory partition unit 880 of the PPU 800 of Fig. 8 according to one embodiment. As in Fig. 9B, the memory partition unit 880 includes a Raster Operations (ROP) unit 950, a Level Two (L2) cache 960, and a memory interface 970. The memory interface 970 is coupled to the memory 804. The memory interface 970 may implement 32-, 64-, 128-, 1024-bit data buses, or the like, for high-speed data transfer. In one embodiment, the PPU 800 includes U memory interfaces 970, one memory interface 970 per pair of partitioning units 880, with each pair of partitioning units 880 connected to a corresponding memory device 804. For example, the PPU 800 may be connected to up to Y memory devices 804, such as, for example, a plurality of memory interfaces 970. E.g. high-bandwidth memory stacks or double-rate version 5 graphics memory, synchronous dynamic random access memory, or other types of persistent memory.
[0140] In one embodiment, memory interface 970 implements an HBM2 memory interface, and Y is equal to half of U. In one embodiment, the HBM2 memory stacks are located in the same physical package as PPU 800, enabling significant power and area savings compared to conventional GDDR5 SDRAM systems. In one embodiment, each HBM2 stack has four memory chips, and Y is equal to 4, with the HBM2 stack having two 128-bit channels per chip for a total of 8 channels and a data bus width of 824 bits.
[0141] In one embodiment, memory 804 supports Single-Error Correcting Double-Error Detecting (SECDED) Error Correction Code (ECC) to protect data. ECC provides greater reliability for data processing applications sensitive to data corruption. Reliability is especially important in large cluster computing environments where PPUs 800 process very large data sets and / or run applications for extended periods of time.
[0142] In one embodiment, the PPU 800 implements a multi-level memory hierarchy. In one embodiment, the memory partitioning unit 880 supports unified memory to provide a single unified virtual address space for the memory of the CPU and PPU 800 and to enable data sharing between virtual memory systems. In one embodiment, the frequency of accesses by a PPU 800 to memory on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 800 that accesses the pages more frequently. In one embodiment, the NVLink 810 supports address translation services that allow the PPU 800 to directly access a CPU's page tables and allow the PPU 800 full access to CPU memory.
[0143] In one embodiment, copy engines transfer data between multiple PPUs 800 or between PPUs 800 and CPUs. The copy engines can generate page faults for addresses not mapped in the page tables. The memory partitioning unit 880 can then handle the page faults and map the addresses to the page table, after which the copy engine can perform the transfer. In a conventional system, memory is pinned between multiple processors for multiple copy engine operations (e.g., is not pageable), significantly reducing the available memory. With hardware page faulting, addresses can be passed to the copy engines without worrying about whether the memory pages are resident, and the copy process is transparent.
[0144] Data from memory 804 or other system memory may be retrieved by memory partition unit 880 and stored in L2 cache 960, which is located on-chip and shared by the various GPCs 850. As illustrated, each memory partition unit 880 has a portion of L2 cache 960 associated with a corresponding device 804. Lower-level caches may then be implemented in various units within the GPCs 850. For example, each of the SMs 940 may implement an L1 cache. The L1 cache is private memory dedicated to a particular SM 940. Data from L2 cache 960 may be retrieved and stored in any of the L1 caches for processing in the functional units of the SMs 940. The L2 cache 960 is coupled to memory interface 970 and XBar 870.
[0145] The ROP unit 950 performs graphics raster operations related to pixel color, such as color compression, pixel blending, and the like. The ROP unit 950 also implements depth checking in conjunction with the raster engine 925, obtaining a depth for a sample location associated with a pixel fragment from the culling engine of the raster engine 925. The depth is compared to a corresponding depth in a depth buffer for a sample location associated with the fragment. If the fragment passes the depth check for the sample location, the ROP unit 950 updates the depth buffer and transmits the result of the depth check to the raster engine 925. The number of partition units 880 may differ from the number of GPCs 850, so that each ROP unit 950 may be coupled to each of the GPCs 850.The ROP unit 950 tracks the packets received from the various GPCs 850 and determines to which GPC 850 a result generated by the ROP unit 950 is forwarded via the Xbar 870. Although the ROP unit 950 is in . Fig. 9B is included in the memory partition unit 880, in other implementations it may also be located outside the memory partition unit 880. For example, the ROP unit 950 may be housed in the GPC 850 or another unit.
[0146] Fig. 10A shows the streaming multiprocessor 940 from Fig. 9A according to one embodiment. As in Fig. 10A, the SM 940 includes an instruction cache 1005, one or more scheduler units 1010, a register file 1020, one or more processing cores 1050, one or more special function units (SFUs) 1052, one or more load / store units (LSUs) 1054, an interconnect network 1080, and a SMEM / L1 cache 1070.
[0147] As described above, the work distribution unit 825 distributes tasks for execution among the GPCs 850 of the PPU 800. The tasks are assigned to a particular DPC 920 within a GPC 850, and if the task is associated with a shader program, the task may be assigned to an SM 940. The scheduler unit 1010 receives the tasks from the work distribution unit 825 and manages the scheduling of instructions for one or more thread blocks assigned to the SM 940. The scheduler unit 1010 schedules thread blocks for execution as warps of parallel threads, where each thread block consists of at least one warp. In one embodiment, each warp includes 32 threads. The scheduling unit 1010 may manage a plurality of different thread blocks by assigning the different thread blocks to different warps and then dispatching instructions from the plurality of different cooperative groups to the different functional units (e.g.,Cores 1050, SFUs 1052 and LSUs 1054) during each clock cycle.
[0148] Cooperative group arrays (CGAs) provide a programming model for organizing groups of communicating threads, allowing developers to express the granularity at which threads communicate, thus enabling richer, more efficient parallel decompositions. Cooperative startup APIs support synchronization between thread blocks for the execution of parallel algorithms. Conventional programming models provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define groups of threads at a smaller granularity than thread blocks and synchronize within the defined groups to enable higher performance, design flexibility, and software reuse in the form of collective group-wide functional interfaces.
[0149] Cooperative group arrangements allow programmers to explicitly define groups of threads at sub-block (e.g., as small as a single thread) and multi-block granularity and perform collective operations on the threads, such as synchronization, within a cooperative group. The programming model supports clean composition across software boundaries, allowing libraries and utility functions to be safely synchronized within their local context without making assumptions about convergence. Cooperative group primitives enable new patterns of cooperative concurrency, including producer-consumer concurrency, opportunistic concurrency, and global synchronization across an entire grid of thread blocks.Hierarchical grouping of threads such as cooperating thread arrays (CTA) and cooperating group arrays (CGA) according to some embodiments are described in more detail in U.S. Application No. 17 / 691,621, the entire contents of which are hereby incorporated by reference in its entirety.
[0150] A dispatcher unit 1015 is configured to dispatch instructions to one or more of the functional units. In the embodiment, the scheduler unit 1010 includes two dispatcher units 1015, allowing two different instructions from the same warp to be dispatched during each clock cycle. In alternative embodiments, each scheduler unit 1010 may include a single dispatcher unit 1015 or additional dispatcher units 1015.
[0151] Each SM 940 includes a register file 1020 that provides a set of registers for the functional units of the SM 940. In one embodiment, the register file 1020 is partitioned among the individual functional units, so that each functional unit is assigned its own section of the register file 1020. In another embodiment, the register file 1020 is partitioned among the various warps executed by the SM 940. The register file 1020 provides temporary storage for operands associated with the data paths of the functional units.
[0152] Each SM 940 includes multiple processing cores 1050. In one embodiment, the SM 940 includes a large number (e.g., 128, etc.) of different processing cores 1050. Each core 1050 may include a fully pipelined, single-, double-, and / or mixed-precision processing unit including a floating-point arithmetic logic unit and an integer arithmetic logic unit. In one embodiment, the floating-point arithmetic logic units implement the IEEE 754-2008 standard for floating-point arithmetic.
[0153] Tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are present in cores 1050. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for training and inferencing neural networks.
[0154] In some embodiments, the transposition hardware is present in the processing cores 1050 or another functional unit (e.g., SFUs 1052 or LSUs 1054) and is configured to generate matrix data stored along diagonals and / or generate the original matrix and / or transposed matrix from the matrix data stored along diagonals. The transposition hardware may be provided within the SMEM 1070 to register a load path of the file 1020 of the SM 940.
[0155] Each SM 940 also includes a plurality of SFUs 1052 that perform specialized functions (e.g., attribute evaluation, reciprocal square root, and the like). In one embodiment, the SFUs 1052 may include a tree traversal unit (e.g., TTU 943) configured to traverse a hierarchical tree data structure. In one embodiment, the SFUs 1052 may include a texture unit (e.g., texture unit 942) configured to perform texture map filtering operations. In one embodiment, the texture units are configured to load texture maps (e.g., a 2D array of texels) from memory 1004 and sample the texture maps to generate sampled texture values for use in shader programs executed by the SM 940. In one embodiment, the texture maps are stored in the SMEM / L1 cache 970. The texture units implement texture operations such as:Filtering operations using mipmaps (e.g., texture maps with different levels of detail). In one embodiment, each SM 940 has two texture units.
[0156] Each SM 940 also includes a plurality of LSUs 1054 that perform load and store operations between the SMEM / L1 cache 1070 and the register file 1020. Each SM 940 has an interconnect network 1080 that connects each of the functional units to the register file 1020 and the LSU 1054 to the register file 1020 and the SMEM / L1 cache 1070. In one embodiment, the interconnect network 1080 is a crossbar that can be configured to connect each of the functional units to each of the registers in the register file 1020 and to connect the LSUs 1054 to the register file 1020 and to locations in the SMEM / L1 cache 1070.
[0157] The SMEM / L1 cache 1070 is an arrangement of on-chip memory that enables data storage and communication between the SM 940 and the primitive engine 935, as well as between threads within the SM 940. In one embodiment, the SMEM / L1 cache 1070 comprises 128 KB of memory capacity and is located on the path from the SM 940 to the partition unit 880. The SMEM / L1 cache 1070 can be used to cach read and write operations. One or more of the SMEM / L1 cache 1070, the L2 cache 960, and the memory 1004 are backup memory.
[0158] Combining data cache and SMEM functionality in a single memory block provides the best overall performance for both types of memory access. This capacity can be used as a cache by programs that don't utilize SMEM. For example, if SMEM is configured to use half of its capacity, texture and load / store operations can use the remaining capacity. Integration with the SMEM / L1 cache 1070 enables the SMEM / L1 cache 1070 to act as a high-throughput conduit for streaming data while providing high-bandwidth, low-latency access to frequently reused data.
[0159] In the context of this disclosure, an SM or "streaming multiprocessor" comprises a means as described in USP 7,447,873 to Nordquist, including improvements and evolutions thereof, and as implemented, for example, in many generations of NVIDIA GPUs. For example, an SM may comprise a plurality of processing engines or cores configured to concurrently execute a plurality of threads arranged in a plurality of SIMD (Single-Instruction, Multiple-Data) groups (e.g., warps), where each of the threads in the same SIMD group executes the same data processing program comprising a sequence of instructions for a different input object, and different threads in the same SIMD group execute using different processing engines or cores.An SM may also typically provide: (a) a local register file with multiple lanes, with each processing processor or core configured to access a different subset of the lanes; and instruction dispatch logic configured to select one of the SIMD groups and dispatch one of the instructions of the same data processing program to each of the multiple processing processors in parallel, with each processing processor executing the same instruction in parallel with every other processing processor using the subset of lanes of the local register file accessible to it. An SM also typically includes core interface logic configured to initiate the execution of one or more SIMD groups.As shown in the figures, such SMs were designed to provide a fast local SMEM that enables data sharing / reuse and synchronization between all threads of a CTA running on the SM.
[0160] When configured for general parallel computing, a simpler configuration can be used compared to graphics processing. In particular, the Fig. 9A, resulting in a much simpler programming model. In the general purpose parallel computing configuration, the work distribution unit 1025 allocates blocks of threads and dispatches them directly to the DPCs 920. The threads in a block execute the same program, using a unique thread ID in the computation to ensure that each thread produces unique results, using the SM 940 to execute the program and perform computations, the SMEM / L1 cache 1070 for inter-thread communication, and the LSU 1054 to read and write global memory via the SMEM / L1 cache 1070 and the memory partition unit 1080. When configured for general purpose parallel computing, the SM 940 can also write instructions that the scheduler unit 820 can use to start new work on the DPCs 920.
[0161] The PPU 800, or a multi-dielet GPU comprising multiple PPUs, as described in this disclosure, may be present in a desktop computer, a laptop computer, a tablet computer, servers, supercomputers, a smartphone (e.g., a wireless handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, and the like. In one embodiment, the PPU 800 is packaged on a single semiconductor substrate. In another embodiment, the PPU 800 is included in a system-on-a-chip (SoC) along with one or more other devices, such as additional PPUs 800, the memory 1004, a RISC (reduced instruction set computer) CPU, an MMU (memory management unit), a DAC (digital-to-analog converter), and the like. Example computer system
[0162] Systems with multiple GPUs and CPUs are being used in a wide variety of industries as developers discover and leverage greater parallelism in applications such as artificial intelligence. Powerful GPU-accelerated systems with tens to many thousands of compute nodes are being deployed in data centers, research facilities, and supercomputers to solve increasingly large problems. As the number of processing units within high-performance systems increases, the communication and data transfer mechanisms must scale to support the increased bandwidth.
[0163] Fig. 11A is a conceptual diagram of a processing system 1100 implemented with multi-dielet GPUs 1000 that may comprise two or more of the PPUs 800 of Fig. 8. The exemplary system 1100 may be configured to implement the methods disclosed in this application. The processing system 1100 includes a CPU 1130, a switch 1155, and a plurality of multi-dielet GPUs and corresponding memories 1004. The NVLink 1010 provides high-speed communication links between each of the multi-dielet GPUs. Although in Fig. 11A illustrates a specific number of NVLink 1010 and Interconnect 1002 connections, the number of connections to each multi-dielet GPU and CPU 1130 may vary. Switch 1155 interfaces interconnect 1002 with CPU 1130. Multi-dielet GPUs 1101, memories 1004, and NVLinks 1010 may be packaged on a single semiconductor platform to form a parallel processing module 1125. In one embodiment, switch 1155 supports two or more protocols to mediate between different connections and / or links.
[0164] In another embodiment (not shown), NVLink 1010 provides one or more high-speed communication links between each of the multi-dielet GPUs and CPU 1130, and switch 1155 interfaces interconnect 1002 with each of the multi-dielet GPUs. Multi-dielet GPUs 1101, memories 1004, and interconnect 1002 may be arranged on a single semiconductor platform to form parallel processing module 1125. In another embodiment (not shown), interconnect 1002 provides one or more communication links between each of the multi-dielet GPUs and CPU 1130, and switch 1155 interfaces each of the multi-dielet GPUs using NVLink 1010 to provide one or more high-speed communication links between the multi-dielet GPUs.In another embodiment (not shown), NVLink 1010 provides one or more high-speed communication links between the multi-dielet GPUs and CPU 1130 via switch 1155. In another embodiment (not shown), interconnect 1002 establishes one or more communication links directly between the individual multi-dielet GPUs. One or more of the NVLink 1010 high-speed communication links may be implemented as a physical NVLink link or as either an on-chip or on-die link using the same protocol as NVLink 1010.
[0165] In the context of the present description, a single semiconductor platform may refer to a single unified semiconductor-based integrated circuit fabricated on a die or chip. It should be noted that the term single semiconductor platform may also refer to multi-chip modules with increased connectivity that simulate on-chip operation and provide significant improvements over the use of a conventional bus implementation. Of course, the various circuits or devices may also be packaged separately or in various combinations of semiconductor platforms, depending on the user's preference. Alternatively, the parallel processing module 1125 may be implemented as a printed circuit board substrate, and each of the multi-dielet GPUs 1000 and / or memory 1004 may be packaged devices.In one embodiment, the CPU 1130, the switch 1155, and the parallel processing module 1125 are located on a single semiconductor platform.
[0166] In one embodiment, NVLink 1010 enables direct load / store / atomic access from CPU 1130 to memory 1104 of each multi-dielet GPU 1000. In one embodiment, NVLink 1010 supports coherent operations so that data read from memories 1104 can be stored in the cache hierarchy of CPU 1130, reducing cache access latency for CPU 1130. In one embodiment, NVLink 1010 includes support for Address Translation Services (ATS), allowing multi-dielet GPU 1000 to directly access page tables within CPU 1130. One or more of NVLinks 1010 may also be configured to operate in a low-power mode.
[0167] Fig.11B shows an example system 1165 in which the various architectures and / or functions of the various embodiments may be implemented. The example system 1165 may be configured to implement the methods disclosed in this application.
[0168] As shown, a system 1165 is provided that includes at least one central processing unit 1130 connected to a communications bus 1175. The communications bus 1175 may be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communications protocol. The system 1165 also includes a main memory 1140. Control logic (software) and data are stored in main memory 1140, which may take the form of random access memory (RAM).
[0169] System 1165 also includes input devices 1160, parallel processing system 1125, and display devices 1145, e.g., a conventional CRT (cathode ray tube), LCD (liquid crystal display), LED (light-emitting diode), plasma display, or the like. User inputs may be received by input devices 1160, e.g., keyboard, mouse, touchpad, microphone, and the like. Each of the aforementioned modules and / or devices may even be arranged on a single semiconductor platform to form system 1165. Alternatively, the various modules may be arranged separately or in various combinations of semiconductor platforms, depending on the user's preferences.
[0170] In addition, the system 1165 may be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a wired network, or the like) for communication purposes via a network interface 1135.
[0171] System 1165 may also include secondary storage (not shown). Secondary storage includes, for example, a hard disk drive and / or a removable storage drive, such as a floppy disk drive, a magnetic tape drive, a compact disk drive, a digital versatile disk (DVD) drive, a recording device, or a universal serial bus (USB) flash drive. The removable storage drive reads from and / or writes to removable storage in a known manner.
[0172] Computer programs or logical algorithms for computer control may be stored in main memory 1140 and / or secondary storage. Such computer programs, when executed, enable system 1165 to perform various functions. Memory 1140, storage, and / or any other storage are possible examples of computer-readable media.
[0173] The architecture and / or functionality of the various preceding figures may be implemented in the context of a general computer system, a circuit system, an entertainment game console system, an application-specific system, and / or any other desired system. For example, system 1165 may take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., a wireless handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a cellular phone device, a television, a workstation, a game console, an embedded system, and / or other type of logic.
[0174] An application program may be implemented via an application executed by a host processor, such as a CPU. In one embodiment, a device driver may implement an application programming interface (API) that defines various functions that can be used by the application program to generate graphical data for display. The device driver is a software program that includes a plurality of instructions that control the operation of the multi-dielet GPU 1000, which includes two or more PPUs. The API provides an abstraction for a programmer, allowing them to utilize specialized graphics hardware, such as the PPU 800, to generate the graphical data without requiring the programmer to utilize the specific instruction set for the PPU 800. The application may include an API call that is passed to the device driver for the multi-dielet GPU.The device driver interprets the API call and performs various operations to respond to the API call. In some cases, the device driver may perform operations by executing instructions on the CPU. In other cases, the device driver may perform operations, at least in part, by initiating operations on the multi-dielet GPU 1000 using an input / output interface between the CPU and the multi-dielet GPU. In one embodiment, the device driver is configured to implement the graphics processing pipeline using the hardware of the PPU 800.
[0175] In the multi-dielet GPU 1000, which includes two or more PPUs 800, various programs can be executed to implement the various processing stages for the application program. For example, the device driver can launch a kernel on the PPU 800 to perform a processing stage on an SM 940 (or multiple SMs 940). The device driver (or the initial kernel executed by the PPU 800) can also launch other kernels on the PPU 800 to perform other processing steps. If the application program processing includes a graphics processing pipeline, some of the stages of the graphics processing pipeline can be implemented on a fixed hardware unit, such as a rasterizer or a data assembler implemented in the PPU 800.It becomes clear that the results of a kernel can be processed by one or more intermediate fixed-function hardware units before being processed by a subsequent kernel on an SM 940.
[0176] The techniques disclosed herein may be incorporated into any processor that can be used to process a neural network, such as a central processing unit (CPU), a graphics processing unit (GPU), an intelligence processing unit (IPU), a neural processing unit (NPU), a tensor processing unit (TPU), a neural network processor (NNP), a data processing unit (DPU), a vision processing unit (VPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and the like. Such a processor may be incorporated into a personal computer (e.g., a laptop), a data center, an Internet of Things (IoT) device, a wearable device (e.g., a smartphone), a vehicle, a robot, or any other device that performs inference, training, or other processing of a neural network.Such a processor can be used in a virtualized system so that an operating system running in a virtual machine on the system can use the processor.
[0177] As an example, a processor incorporating the techniques disclosed herein may be employed to execute one or more neural networks in a machine to identify, classify, manipulate, handle, operate, modify, or navigate physical objects in the real world. Such a processor may, for example, be employed in an autonomous vehicle (e.g., a car, motorcycle, helicopter, drone, airplane, boat, submarine, delivery robot, etc.) to move the vehicle through the real world. Furthermore, such a processor may be employed in a robot in a factory to select and assemble components into an assembly.
[0178] For example, a processor utilizing the techniques disclosed herein may be employed to execute one or more neural networks to identify one or more features in an image, or to alter, generate, or compress an image. Such a processor may be employed, for example, to enhance an image rendered using raster, ray tracing (e.g., with NVIDIA RTX), and / or other rendering techniques. In another example, such a processor may be employed to reduce the amount of image data transferred from a rendering device to a display over a network (e.g., the Internet, a mobile telecommunications network, a Wi-Fi network, or any other wired or wireless network system). Such transfers may be used to transfer image data from a server or data center in the cloud to a user device (e.g.,a PC, video game console, smartphone, other mobile device, etc.) to enhance services that stream images, such as NVIDIA GeForce Now (GFN), Google Stadia, and the like.
[0179] As an example, a processor incorporating the techniques disclosed herein may be used to execute one or more neural networks for other types of applications that can utilize the benefits of a neural network. Such applications may include, for example, translating from one spoken language to another, identifying and suppressing noise in audio signals, detecting anomalies or defects in the production of goods and services, monitoring living and / or non-living objects, medical diagnosis, decision making, and the like.
[0180] As an example, a processor incorporating the techniques disclosed herein may be employed to implement neural networks such as large language models (LLMs) to generate content (e.g., images, videos, text, essays, audio, and the like), respond to user queries, solve problems in mathematical and other domains, and the like.
[0181] All patents, patent applications, and publications cited herein are incorporated by reference for all purposes as if expressly listed.
[0182] While the invention has been described in connection with what is presently considered to be the most practical and preferred embodiment, it is to be understood that the invention is not limited to the disclosed embodiment, but on the contrary is intended to cover various modifications and equivalent arrangements which are within the spirit and scope of the appended claims. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 18 / 606,924
[0016] US 17 / 691,621 [0132, 0149] US 17 / 691,690
[0132] US 7,447,873 f
[0159]
Claims
[1] Multi-dielet processing system comprising: a first memory management unit, MMU, on a first dielet; and a second MMU on a second dielet, wherein the first MMU on the first dielet is configured to process memory-related requests of a first type from the first dielet and the second dielet, and to process memory-related requests of a second type from the first dielet, and wherein the second MMU on the second dielet is configured to forward memory-related requests of the first type to the first MMU on the first dielet and to process memory-related requests of the second type from the second dielet. [2] The multi-dielet processing system of claim 1, wherein the second MMU is further configured to process memory-related requests of a third type received from the second dielet and the first dielet, and wherein the first MMU is further configured to forward memory-related requests of the third type to the second MMU. [3] The multi-dielet processing system of claim 2, wherein the first MMU is configured to serialize the memory-related requests of the first type received from the first dielet and the second dielet, and wherein the second MMU is configured to serialize the memory-related requests of the third type received from the first dielet and the second dielet. [4] A multi-dielet processing system according to claim 2 or 3, wherein the first MMU is configured to generate, in response to at least one memory-related message of the first type, a message to be transmitted on a network interface of a first type, wherein the second MMU is configured to generate, in response to at least one memory-related message of the third type, a message to be transmitted on a network interface of a second type, wherein the network interface of the second type is located on the second dielet. [5] A multi-dielet processing system according to any one of claims 2 to 4, wherein the first MMU and the second MMU are each configured to perform an operating role according to a respective predetermined configuration setting. [6] The multi-dielet processing system of claim 5, wherein a third MMU located on a third dielet of the multi-dielet processing system performs an independent operational role based on the configuration setting. [7] The multi-dielet processing system of claim 5 or 6, wherein the first MMU performs a primary operational role based on the corresponding configuration setting, and wherein the second MMU performs a secondary operational role based on the corresponding configuration setting. [8] A multi-dielet processing system according to any preceding claim, wherein the memory-related requests of the first type comprise a binding request. [9] The multi-dielet processing system of claim 8, wherein the first MMU is configured, in response to the binding request, to: obtain binding information corresponding to the binding request, to update a bind table and / or a page directory bind cache in a memory of the first dielet, and to transfer the binding information to the second dielet. [10] The multi-dielet processing system of claim 9, wherein the binding request originates from the second dielet and is transmitted from the second dielet to the first dielet, and wherein the second MMU is further configured to: update a bind table and / or a page directory bind cache in a memory of the first dielet in response to receiving the binding information from the first dielet, and to respond to a source of the binding request. [11] The multi-dielet processing system of claim 10, wherein in response to receiving the binding information, each of the first MMU and the second MMU is further configured to update at least one translation lookaside buffer, TLB, on the local dielet according to the binding information. [12] A multi-dielet processing system according to any preceding claim, wherein the memory-related requests of the first type further comprise an error report request, a TLB invalidation request, or a VAB dump request. [13] The multi-dielet processing system of claim 12, wherein the memory-related requests of the second type comprise an address translation request, and wherein both the first MMU and the second MMU are configured to perform a lookup function in a TLB and / or a page table in response to an address translation request. [14] The multi-dielet processing system of claim 13, wherein the memory-related requests of the second type further comprise an ECC interrupt, a GPU page table walk, a memory barrier request, a memory-to-memory input-output request, or a packet generation request for a work generation request. [15] A multi-dielet processing system according to any preceding claim, wherein the second MMU is configured to, in response to a memory barrier request from the second dielet, transmit requests preceding the memory barrier request from hardware engines in a first phase, and to transmit one or more input-output flush, I / O flush, requests in a second phase. [16] A multi-dielet processing system according to any one of the preceding claims, wherein the second MMU is configured to, in response to a memory barrier request from the second dielet, transmit requests preceding the memory barrier request from hardware engines in a first phase and to transmit an input-output flush, I / O flush, request to the first MMU in a second phase, and wherein the first MMU is configured to serialize the execution of the IO flush requests. [17] A multi-dielet processing system according to any one of the preceding claims, wherein the first MMU and the second MMU are each configured to enable programming of one or more registers on the first dielet and the second dielet, respectively, by software. [18] A multi-dielet processing system according to any preceding claim, wherein the first MMU and the second MMU are each configurable in one of three settings: primary, secondary, or standalone. [19] A multi-dielet processing system according to any preceding claim, further comprising a high-speed interface connecting the first dielet and the second dielet, wherein the second MMU is configured to forward the memory-related requests of the first type to the first MMU via the high-speed interface. [20] The multi-dielet processing system of claim 19, wherein the high-speed interface is a dedicated connection for dielet-to-dielet communication within the multi-dielet processing system. [21] A multi-dielet processing system according to any preceding claim, wherein the hardware engine remapping circuit is configured to determine a globally unique dielet identifier for each dielet based on a dielet-identifying fuse signal. [22] A multi-dielet processing system according to any preceding claim, wherein the hardware engine remapping circuit comprises a plurality of dielet-level remapping circuits, and each dielet-level remapping circuit is configured to operate independently of other dielet-level remapping circuits. [23] A multi-dielet processing system according to any preceding claim, wherein each dielet-level remapping circuit is configured to apply an offset to hardware engine identifiers in received messages, the offset being determined based on a dielet-identifying fuse signal. [24] A multi-dielet processing system according to any preceding claim, wherein each dielet comprises a plurality of processing engines. [25] A multi-dielet processing system according to any preceding claim, wherein a plurality of parallel processors are provided on respective different dielets, including the first dielet and the second dielet, in a single housing, the housing being adapted to be connected to a central processing unit, CPU, in a computer. [26] A multi-dielet processing system according to any preceding claim, wherein each of the dielets is a graphics processing unit, GPU. [27] A method performed by a multi-dielet processing system configured to process a plurality of types of memory-related requests, the multi-dielet processing system comprising a first memory management unit, MMU, on a first dielet and a second MMU on a second dielet, the method comprising: Processing, by the first MMU, memory-related requests of a first type from the first dielet and the second dielet and memory-related requests of a second type received from the first dielet; Forwarding memory-related requests of the first type by the second MMU to the first MMU; and Processing memory-related requests of the second type received from the second dielet by the second MMU.
Citation Information
Patent Citations
US-ANMELDUNGNR.17/691,621
USP7,447,873F
US-ANMELDUNGNR.17/691,690
US-ANMELDUNGNR.18/606.924