Accelerator structure

By extending the accelerator link (IAL) of the R-Link protocol, efficient memory access between the accelerator and the host device is achieved, solving the problem of low accelerator memory management efficiency in traditional systems, improving memory access efficiency and throughput, and reducing power consumption.

CN109582611BActive Publication Date: 2025-07-22INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201810988366.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-29
Filing Date
2018-08-28
Publication Date
2025-07-22
Estimated Expiration
2038-08-28

AI Technical Summary

Technical Problem

In traditional computing systems, the accelerator's memory management is inefficient, resulting in inefficient cache coherence mechanisms, limiting the accelerator's ability to access local memory at high bandwidth and deployment options.

Method used

The accelerator link (IAL) protocol is adopted, and the R-Link protocol is extended to support accelerators and input/output devices to realize the interconnection structure of coherent memory devices, provide dynamic multiplexing and bias checking, and optimize memory access between the accelerator and the host device.

Benefits of technology

Improves the accelerator's memory access efficiency and throughput, reduces power consumption, supports more efficient accelerator deployment and memory utilization, and reduces data replicas and complex software operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109582611B_ABST
    Figure CN109582611B_ABST
Patent Text Reader

Abstract

A structure controller for providing a coherent accelerator structure, comprising: a host interconnect communicatively coupled to a host device; a memory interconnect communicatively coupled to an accelerator memory; an accelerator interconnect communicatively coupled to an accelerator having a last-level cache (LLC); and an LLC controller configured to provide a bias check for memory access operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of interconnect devices, and more specifically but not exclusively, to systems and methods for coherent memory devices for Peripheral Component Interconnect Express (PCIe). Background Art

[0002] Computing systems include various components for managing the demand for processor resources. For example, a developer may include a hardware accelerator (or "accelerator") operatively coupled to a central processing unit (CPU). Generally, an accelerator is an autonomous element configured to perform functions delegated to it by the CPU. An accelerator can be configured for specific functions and / or can be programmable. For example, an accelerator can be configured to perform specific computations, graphics functions, etc. When the accelerator performs the designated function, the CPU is free to use its resources for other needs. In traditional systems, an operating system (OS) can manage the physical memory available within the computing system (e.g., "system memory"); however, the OS does not manage or allocate the memory local to the accelerator. As a result, memory protection mechanisms such as cache coherence introduce inefficiencies into accelerator-based configurations. For example, traditional cache coherence mechanisms limit the ability of the accelerator to access its connected local memory at very high bandwidths and / or limit the deployment options of the accelerator. Brief Description of the Drawings

[0003] The present invention can be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, according to standard practice in the industry, the various features are not necessarily drawn to scale and are for illustrative purposes only. Where scale is explicitly or implicitly shown, it only provides an illustrative example. In other embodiments, the dimensions of the various features may be increased or decreased arbitrarily for clarity.

[0004] Figure 1 An example operating environment is shown that can represent various embodiments in accordance with one or more examples of this specification.

[0005] Figure 2a An example of a fully coherent operating environment is shown in accordance with one or more examples of this specification.

[0006] Figure 2b An example of an incoherent operating environment is shown in accordance with one or more examples of this specification.

[0007] Figure 2c An example of a coherent engine without a biased operating environment is shown in accordance with one or more examples of this specification.

[0008] Figure 3Shows an example of an operating environment that can represent various embodiments according to one or more examples of the present specification.

[0009] Figure 4 Shows another example operating environment that can represent various embodiments according to one or more examples of the present specification.

[0010] Figure 5a and 5b Shows other example operating environments that can represent various embodiments according to one or more examples of the present specification.

[0011] Figure 6 Shows an embodiment of a logical flow according to one or more examples of the present specification.

[0012] Figure 7 Is a block diagram showing a structure according to one or more examples of the present specification.

[0013] Figure 8 Is a flowchart showing a method according to one or more examples of the present specification.

[0014] Figure 9 Is a Block diagram of Accelerator Link Memory (IAL.mem) read through PCIe operation according to one or more examples of the present specification.

[0015] Figure 10 Is a block diagram of IAL.mem write through PCIe operation according to one or more examples of the present specification.

[0016] Figure 11 Is a block diagram of IAL.mem data completion through PCIe operation according to one or more examples of the present specification.

[0017] Figure 12 Shows an embodiment of a structure composed of point-to-point links interconnecting a group of components according to one or more examples of the present specification.

[0018] Figure 13 Shows an embodiment of a hierarchical protocol stack according to one or more embodiments of the present specification.

[0019] Figure 14 Shows an embodiment of a PCIe transaction descriptor according to one or more examples of the present specification.

[0020] Figure 15 Shows an embodiment of a PCIe serial point-to-point structure according to one or more examples of the present specification. Detailed Description

[0021] Of the present specification The Accelerator Link (IAL) is an extension of the Rosetta Link (R-Link) multi-chip package (MCP) interconnect link. The IAL extends the R-Link protocol to support accelerators and input / output (IO) devices that may not be fully supported by the baseline R-Link or the Peripheral Component Interconnect Express (PCIe) protocol.

[0022] The following disclosure provides many different embodiments or examples for implementing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present disclosure. Of course, these are merely examples and not restrictive. Additionally, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for simplicity and clarity purposes and does not in itself govern the relationship between the various embodiments and / or configurations discussed. Different embodiments may have different advantages, and no particular advantage is necessarily required of any embodiment.

[0023] In the following description, numerous specific details are set forth, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurement / height, specific processor pipeline stages and operations, etc., in order to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention need not be practiced with these specific details.

[0024] In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / codes for the algorithms described, specific firmware codes, specific interconnect operations, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expressions of algorithms in code, specific power-down and gating techniques / logics, and other specific operating details of computer systems, are not described in detail to avoid unnecessarily obscuring the present invention.

[0025] Although the following embodiments may be described with reference to energy conservation and energy efficiency in a particular integrated circuit, such as in a computing platform or a microprocessor, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein may be applied to other types of circuits or semiconductor devices, which may also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks TM , and may also be used in other devices, such as handheld devices, tablet computers, other thin notebooks, system-on-chip (SOC) devices, and embedded applications.

[0026] Some examples of handheld devices include cellular telephones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld personal computers (PCs). Embedded applications typically include microcontrollers, digital signal processors (DSPs), systems on a chip (SoCs), network personal computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can perform the functions and operations taught below. Additionally, the apparatus, methods, and systems described herein are not limited to physical computing devices and can also relate to software optimizations for energy conservation and efficiency. As will become apparent in the following description, embodiments of the methods, apparatus, and systems described herein (whether referring to hardware, firmware, software, or combinations thereof) are crucial for a “green technology” future balanced with performance considerations.

[0027] Various embodiments can generally relate to techniques for providing cache coherence among multiple components within a processing system. In some embodiments, the multiple components can include a processor, such as a central processing unit (CPU), and a logic device communicatively coupled to the processor. In various embodiments, the logic device can include locally attached memory. In some embodiments, the multiple components can include a processor communicatively coupled to an accelerator having locally attached memory (e.g., logic device memory).

[0028] In some embodiments, the processing system can operate a coherence biasing process that is configured to provide multiple cache coherence processes. In some embodiments, the multiple cache coherence processes can include a device biasing process and a host biasing process (collectively referred to as the “biasing protocol flow”). In some embodiments, the host biasing process can route requests to the locally attached memory of the logic device through a coherence component of the processor, including requests from the logic device. In some embodiments, the device biasing process can route logic device requests for the logic device memory directly to the logic device memory, e.g., without consulting the coherence component of the processor. In various embodiments, the cache coherence processes can switch between the device biasing process and the host biasing process based on a biasing indicator determined using application software, hardware hints, combinations thereof, etc. Embodiments are not limited to this context.

[0029] The IAL described in this specification uses an optimized accelerator protocol (OAP), which is a further extension of the R-Link MCP interconnect protocol. In one example, the IAL can be used to provide an interconnect fabric to an accelerator device (in some examples, the accelerator device can be a heavy accelerator that performs, for example, graphics processing, intensive computing, SmartNIC services, or similar processing). The accelerator can have its own attached accelerator memory, and an interconnect fabric such as the IAL or, in some embodiments, a PCIe-based fabric can be used to attach the processor to the accelerator. The interconnect fabric can be a coherent accelerator fabric, in which case the accelerator memory can be mapped to the memory address space of the host device. The coherent accelerator fabric can maintain coherence within the accelerator as well as between the accelerator and the host device. This can be used to implement state-of-the-art memory and to provide coherence support for these types of accelerators.

[0030] Advantageously, the coherent accelerator fabric according to this specification can provide optimizations that improve efficiency and throughput. For example, the accelerator can have a certain number of banks, where the corresponding n last-level caches (LLCs) are each controlled by an LLC controller. This fabric can provide different kinds of interconnects to connect the accelerator and its caches to the memory and to connect the fabric to the host device.

[0031] By way of illustration, throughout this specification, a bus or interconnect that connects devices of the same nature is referred to as a "horizontal" interconnect, while an interconnect or bus that connects different devices upstream and downstream can be referred to as a "vertical" interconnect. The terms "horizontal" and "vertical" used here are for convenience only and do not imply any necessary physical arrangement of the interconnects or buses, or require that they must be physically orthogonal to each other on the die.

[0032] For example, the accelerator may include 8 banks, with corresponding 8 LLCs, which may be level 3 (L3) caches, each controlled by an LLC controller. The coherent accelerator fabric may be partitioned into multiple independent "slices". Each slice services a bank and its corresponding LLC, and operates substantially independently of other slices. In an example, each slice may utilize the bias operations provided by the IAL and provide parallel paths to the banks. Memory operations involving the host device may be routed through a fabric coherence engine (FCE) that provides coherence with the host device. However, the LLC of any single slice may also have a parallel bypass path that writes directly to the memory that connects the LLC directly to the bank, bypassing the FCE. For example, this may be achieved by providing bias logic (e.g., host bias or accelerator bias) within the LLC controller itself. The LLC controller may be physically separated from the FCE and may be upstream of the FCE in a vertical orientation, enabling accelerator-biased memory operations to bypass the FCE and write directly to the bank.

[0033] Embodiments of this specification may also achieve significant power savings by providing a power manager that selectively shuts down portions of the coherent fabric when not in use. For example, the accelerator may be a very large bandwidth accelerator that may perform many operations per second. When the accelerator is performing its acceleration function, it is heavily using the fabric and requires extremely high bandwidth so that the computed values can be flushed to memory in a timely manner after computation. However, once the computation is complete, the host device may not be ready to consume the data yet. In such a case, portions of the interconnect, such as the vertical bus from the FCE to the LLC controller and the horizontal buses between the LLC controller and the LLC itself, may be powered down. These may remain powered down until the accelerator receives new data to operate on.

[0034] The following table illustrates several classes of accelerators. Note that the baseline R-Link may only support the first two classes of accelerators, while the IAL may support all five classes of accelerators.

[0035]

[0036] Note that, in addition to producer-consumer, embodiments of these accelerators may require some degree of cache coherence to support the usage model. Thus, the IAL is a coherent accelerator link.

[0037] The IAL uses a combination of three protocols that are dynamically multiplexed onto a common link to enable the accelerator model disclosed above. These protocols include:

[0038] · System-on-Chip Fabric (IOSF) - A reformatted PCIe-based interconnect that provides a non-coherent in-order semantics protocol. The IOSF may include an on-chip implementation of all or part of the PCIe standard. The IOSF packages PCIe traffic so that it can be sent to a companion chip, such as a system-on-chip (SoC) or multi-chip module (MCM). The IOSF supports device discovery, device configuration, error reporting, interrupts, direct memory access (DMA) style data transfers, and various services provided as part of the PCIe standard.

[0039] • In-die Interconnect (IDI) - enables devices to issue coherent read and write requests to the processor.

[0040] Scalable Memory Interconnect (SMI) - enables the processor to access memory attached to the accelerator.

[0041] These three protocols can be used in different combinations (e.g., IOSF only, IOSF plus IDI, IOSF plus IDI plus SMI, IOSF plus SMI) to support the various models described in the table above.

[0042] As a baseline, IAL provides a single link or bus definition that can cover all five accelerator models through a combination of the aforementioned protocols. Note that producer-consumer accelerators are essentially PCIe accelerators. They only require the IOSF protocol, which is already a reformatted version of PCIe. IOSF supports some Accelerator Interface Architecture (AiA) operations, such as support for the enqueue (ENQ) instruction, which industry standard PCIe devices may not support. Therefore, IOSF provides added value over PCIe for this type of accelerator. Producer-consumer plus accelerators are accelerators that can use only the IDI layer and IOSF layer of IAL.

[0043] In some embodiments, software assisted device memory and autonomous device memory may require the SMI protocol over the IAL, including the inclusion of special operation codes (opcodes) over the SMI and special controller support for flows associated with those opcodes in the processor. These additions support the consistency deviation model of the IAL. Usages may use all of the IOSF, IDI, and SMI.

[0044] The giant cache usage also uses IOSF, IDI, and SMI, but new qualifiers may also be added to the IDI and SMI protocols that are specifically designed for giant cache accelerators (i.e., not used in the device memory model discussed above). The giant cache may add new special controller support in the processor that is not required for any other usage.

[0045] IAL refers to these three protocols as IAL.IO, IAL.cache, and IAL.mem. The combination of these three protocols provides the required performance benefits for five accelerator models.

[0046] To achieve these benefits, IAL can use the R-Link (for MCP) or Flexbus (for discrete) physical layer to allow for dynamic multiplexing of the IO, cache, and mem protocols.

[0047] However, some form factors do not support the R-Link or Flexbus physical layer natively. In particular, Class 3 and Class 4 device memory accelerators may not support R-Link or Flexbus. Existing examples of these can use standard PCIe, which limits the device to a dedicated memory model rather than providing a coherent memory with a write-back memory address space that can be mapped to the host device. This model is restrictive because the memory attached to the device cannot be directly addressed by software. This can lead to suboptimal data marshalling between host and device memory over a bandwidth-constrained PCIe link.

[0048] Accordingly, embodiments of the present specification provide consistency semantics that follow the same deviation model-based definitions defined by IAL, which retain the benefits of consistency without incurring traditional overheads. All of these can be provided over an existing PCIe physical link.

[0049] Thus, some of the advantages of IAL can be realized on a physical layer that does not provide dynamic multiplexing between the IO, cache, and mem protocols provided by R-Link and Flexbus. Advantageously, enabling the IAL protocol over PCIe for certain classes of devices reduces the ecosystem entry burden for devices using the physical PCIe link. This enables the utilization of existing PCIe infrastructure, including the use of off-the-shelf components such as switches, root ports, and endpoints. This also allows for easier cross-platform use of devices with attached memory, using either traditional dedicated memory models or coherent system-addressable memory models suitable for the installation facility.

[0050] To support Class 3 and Class 4 devices (software-assisted memory and autonomous device memory) as described above, the components of IAL can be mapped as follows:

[0051] IOSF or IAL.io can use standard PCIe. This can be used for device discovery, enumeration, configuration, error reporting, interrupts, and DMA-style data transfer.

[0052] SMI or IAL.mem can use an SMI tunnel over PCIe. Details of the SMI tunnel over PCIe are described below, including the following Figure 9, the tunnels described in 10 and 11.

[0053] In some embodiments of this specification, IDI or IAL.cache is not supported. IDI enables a device to issue coherent read or write requests to the host memory. Although IAL.cache may not be supported, the methods disclosed herein can be used to enable bias-based coherence for the memory attached to the device.

[0054] To achieve this result, an accelerator device can use one of its standard PCIe memory base address register (BAR) regions for the size of its attached memory. To this end, the device can implement a specified vendor-specific extended capability (DVSEC), similar to the standard IAL, to point to the BAR region that should be mapped to the coherent address space. In addition, the DVSEC can declare additional information such as memory type, latency, and other attributes, which help the basic input / output system (BIOS) map this memory to the system address decoder in the coherent region. Then, the BIOS can program the memory base address and limit the host physical address in the device.

[0055] This allows the host to read the memory attached to the device using the standard PCIe memory read (MRd) opcode.

[0056] However, for writes, non-posted semantics may be required because access to metadata may be needed upon completion. To obtain NP MWr on PCIe, the following reserved encoding can be used:

[0057] · Fmt[2:0] - 011b

[0058] · Type[4:0] - 11011b

[0059] Using a novel non-posted memory write (NP MWr) on PCIe has the additional benefit of enabling the AiA ENQ instruction to efficiently submit work to the device.

[0060] To achieve the best quality of service, embodiments of this specification can implement three different virtual channels (VC0, VC1, and VC2) to separate different traffic types, as follows:

[0061] · VC0 → All memory-mapped input / output (MMIO) and configuration (CFG) traffic, both upstream and downstream

[0062] · VC1 → IAL.mem writes (from host to device)

[0063] · VC2 → IAL.mem reads (from host to device)

[0064] Note that, since IAL.cache or IDI is not supported, embodiments of this specification may not allow an accelerator device to issue coherent reads or writes to the host memory.

[0065] Embodiments of this specification may also have the ability to flush cache lines from the host (required for host-to-device bias flips). This can be done at the cache line granularity using non-allocated zero-length writes from the device on PCIe. Non-allocated semantics are described using transactions and processing hints on transaction layer packets (TLPs).

[0066] ·TH = 1, PH = 01

[0067] This allows the host to invalidate a given line. The device can issue a read after the page bias flip to ensure all lines are flushed. The device can also implement a content addressable memory (CAM) to ensure that while the flip is in progress, no new requests for the line are received from the host.

[0068] The system and method of a coherent memory device over PCIe will now be described more specifically with reference to the accompanying drawings. It should be noted that throughout the drawings, certain reference numerals may be repeated to indicate that a particular device or block is identical or substantially identical throughout the drawings. However, this is not meant to imply any particular relationship between the various disclosed embodiments. In some examples, a class of elements may be referred to by a particular reference numeral ("widget 10"), while individual species or examples of that class may be referred to by hyphenated numbers ("first particular widget 10-1" and "second particular widget 10-2").

[0069] Figure 1 An example operating environment 100 is shown that may represent various embodiments in accordance with one or more examples of this specification. Figure 1 The operating environment 100 depicted may include an apparatus 105 having a processor 110 (e.g., a central processing unit (CPU)). The processor 110 may include any type of computing element, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a virtual processor, such as a virtual central processing unit (VCPU), or any other type of processor or processing circuitry. In some embodiments, the processor 110 may be one or more processors from a company available from processor family. Although in Figure 1Only one processor 110 is depicted, but the device may include multiple processors 110. The processor 110 may include processing elements 112, such as processing cores. In some embodiments, the processor 110 may include a multi-core processor having multiple processing cores. In various embodiments, the processor 110 may include a processor memory 114, which may include, for example, a processor cache or local cache memory to efficiently access data processed by the processor 110. In some embodiments, the processor memory 114 may include random access memory (memory); however, the processor memory 114 may be implemented using other memory types, such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, and the like.

[0070] As Figure 1 shown, the processor 110 may be communicatively coupled to the logic device 120 via a link 115. In various embodiments, the logic device 120 may include a hardware device. In various embodiments, the logic device 120 may include an accelerator. In some embodiments, the logic device 120 may include a hardware accelerator. In various embodiments, the logic device 120 may include an accelerator implemented in hardware, software, or any combination thereof.

[0071] Although an accelerator is used as an example of the logic device 120 in this specific embodiment, the embodiments are not limited thereto, as the logic device 120 may include any type of device, processor (e.g., a graphics processing unit (GPU)), logic unit, circuit, integrated circuit, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), memory unit, computing unit, and / or the like that can operate according to some embodiments. In embodiments where the logic device 120 includes an accelerator, the logic device 120 may be configured to perform one or more functions of the processor 110. For example, the logic device 120 may include an accelerator operable to perform graphics functions (e.g., a GPU or graphics accelerator), floating-point operations, fast Fourier transform (FFT) operations, and the like. In some embodiments, the logic device 120 may include an accelerator configured to operate using various hardware components, standards, protocols, and the like. Non-limiting examples of the types of accelerators and / or accelerator technologies that can be used by the logic device may include OpenCAPI TM , CCIX, GenZ, NVLink TM , accelerator interface architecture (AiA), cache coherent agent (CCA), global mapping and coherent device memory (GCM), graphics media accelerator (GMA), for directed input / output (IO) Virtualization technologies (e.g., VT-d, VT-x, etc.), shared virtual memory (SVM), etc. Embodiments are not limited to this context.

[0072] The logical device 120 may include a processing element 122, such as a processing core. In some embodiments, the logical device 120 may include multiple processing elements 122. The logical device 120 may include a logical device memory 124, e.g., a locally attached memory configured for the logical device 120. In some embodiments, the logical device memory 124 may include local memory, cache memory, etc. In various embodiments, the logical device memory 124 may include random access memory (RAM); however, the logical device memory 124 may be implemented using other memory types, such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, etc. In some embodiments, at least a portion of the logical device memory 124 may be visible or accessible to the processor 110. In some embodiments, at least a portion of the logical device memory 124 may be visible or accessible to the processor 110 as system memory (e.g., as an accessible portion of the system memory 130).

[0073] In various embodiments, the processor 110 may execute the driver 118. In some embodiments, the driver 118 may be used to control various functional aspects of the logical device 120 and / or manage communication with one or more applications using the logical device 120 and / or computational results generated by the logical device 120. In various embodiments, the logical device 120 may include and / or may access bias information 126. In some embodiments, the bias information 126 may include information associated with a coherence biasing process. For example, the bias information 126 may include information indicating which cache coherence process may be valid for the logical device 120 and / or a particular process, application, thread, memory operation, etc. In some embodiments, the bias information 126 may be read, written, or otherwise managed by the driver 118.

[0074] In some embodiments, link 115 may include a bus component, such as a system bus. In various embodiments, link 115 may include a communication link (e.g., a multi-protocol link) operable to support multiple communication protocols. The supported communication protocols may include standard load / store IO protocols for component communication, including serial link protocols, device cache protocols, memory protocols, memory semantic protocols, directory bit support protocols, networking protocols, coherence protocols, accelerator protocols, data storage protocols, point-to-point protocols, fabric-based protocols, on-package (or on-chip) protocols, fabric-based on-package protocols, and / or similar protocols. Non-limiting examples of supported communication protocols may include Peripheral Component Interconnect (PCI) protocol, Peripheral Component Interconnect Express (PCIe or PCI-E) protocol, Universal Serial Bus (USB) protocol, Serial Peripheral Interface (SPI) protocol, Serial ATA Attachment (SATA) protocol, QuickPath Interconnect (QPI) protocol, UltraPath Interconnect (UPI) protocol, Optimized Accelerator Protocol (OAP), Accelerator Link (IAL), In-Device Interconnect (IDI) protocol (or IAL.cache), In-Out-Scale Fabric (IOSF) protocol (or IAL.io), Scalable Memory Interconnect (SMI) protocol (or IAL.mem), SMI Generation 3 (SMI3), and / or similar protocols. In some embodiments, link 115 may support in-device protocols (e.g., IDI) and memory interconnect protocols (e.g., SMI3). In various embodiments, link 115 may support in-device protocols (e.g., IDI), memory interconnect protocols (e.g., SMI3), and fabric-based protocols (e.g., IOSF).

[0075] In some embodiments, device 105 may include system memory 130. In various embodiments, system memory 130 may include the main system memory for device 105. System memory 130 may store data and sequences of instructions executed by processor 110 or any other device or component of device 105. In some embodiments, system memory 130 may be RAM; however, system memory 130 may be implemented using other memory types, such as dynamic DRAM, SDRAM, combinations thereof, etc. In various embodiments, system memory 130 may store software applications 140 (e.g., “host software”) executable by processor 110. In some embodiments, software applications 140 may be associated with logical device 120, either by using or otherwise. For example, software applications 140 may be configured to use the computational results generated by logical device 120.

[0076] The apparatus may include coherence logic 150 to provide cache coherence processes. In various embodiments, coherence logic 150 may be implemented in hardware, software, or a combination thereof. In some embodiments, at least a portion of coherence logic 150 may be disposed within processor 110, partially disposed within processor 110, or otherwise associated with processor 110. For example, in some embodiments, coherence logic 150 for cache coherence element or process 152 may be disposed within processor 110. In some embodiments, processor 110 may include a coherence controller 116 to perform various cache coherence processes, such as cache coherence process 152. In some embodiments, cache coherence process 152 may include one or more standard cache coherence techniques, functions, methods, processes, elements (including hardware or software elements), protocols, etc., performed by processor 110. Generally, cache coherence process 152 may include standard protocols for managing the caches of the system so as not to lose data or overwrite data before transferring the data from the cache to the target memory. Non-limiting examples of standard protocols performed or supported by cache coherence process 152 may include snooping-based (or snoopy) protocols, write invalidate protocols, write update protocols, directory-based protocols, hardware-based protocols (e.g., Modified Exclusive Shared Invalidate (MESI) protocol), private memory-based protocols, and / or similar protocols. In some embodiments, cache coherence process 152 may include one or more standard cache coherence protocols for maintaining cache coherence for logical device 120 having attached logical device memory 124. In some embodiments, cache coherence process 150 may be implemented in hardware, software, or a combination thereof.

[0077] In some embodiments, coherence logic 150 may include coherence biasing processes, such as host biasing process or element 154 and device biasing process or element 156. Generally, coherence biasing processes may operate to maintain cache coherence related to requests, data streams, and / or other memory operations related to logical device memory 122. In some embodiments, at least a portion of the coherence logic, such as host biasing process 154, device biasing process 156, and / or biasing selection component 158, may be disposed external to processor 110, such as in one or more separate coherence logic 150 units. In some embodiments, host biasing process 154, device biasing process 156, and / or biasing selection component 158 may be implemented in hardware, software, or a combination thereof.

[0078] In some embodiments, the host biasing process 154 may include techniques, processes, data flows, data, algorithms, etc., for processing requests to the logical device memory 124 through the cache coherence process 152 of the processor 110, including requests from the logical device 120. In various embodiments, the device biasing process 156 may include techniques, processes, data flows, data, algorithms, etc., that allow the logical device 120 to directly access the logical device memory 124, e.g., without using the cache coherence process. In some embodiments, the biasing selection process 158 may include techniques, processes, data flows, data, algorithms, etc., for activating either the host biasing process 154 or the device biasing process 156 as the active biasing process for requests associated with the logical device memory. In various embodiments, the active biasing process may be based on biasing information 126, which may include data, data structures, and / or processes used by the biasing selection process to determine and / or set the active biasing process.

[0079] Figure 2a An example of a fully coherent operating environment 200A is shown. Figure 2a The operating environment 200A depicted may include a device 202 having a CPU 210, the CPU 210 including a plurality of processing cores 212a-n. As Figure 2a shown, the CPU may include various protocol agents, such as a cache agent 214, a host agent 216, a memory agent 218, and / or similar agents. Generally, the cache agent 214 may operate to initiate transactions into the coherent memory and maintain copies in its own cache structure. The cache agent 214 may be defined by the messages it can receive and send according to the behavior defined in the cache coherence protocol associated with the CPU. The cache agent 214 may also provide copies of the coherent memory content to other cache agents (e.g., the accelerator cache agent 224). The host agent 216 may be responsible for the protocol side of the memory interactions of the CPU 210, including coherent and non-coherent host agent protocols. For example, the host agent 216 may order memory reads / writes. The host agent 216 may be configured to service coherent transactions, including handshaking with the cache agent if necessary. The host agent 216 may operate to oversee a portion of the coherent memory of the CPU 210, e.g., maintaining the consistency of a given address space. The host agent 216 may be responsible for managing conflicts that may arise between different cache agents. For example, the host agent 216 may provide appropriate data and ownership responses based on the flow of a given transaction. The memory agent 218 may operate to manage access to the memory. For example, the memory agent 218 may facilitate the memory operations (e.g., load / store operations) and functions (e.g., swap and / or similar functions) of the CPU 210.

[0080] AsFigure 2a As shown, the apparatus 202 may include an accelerator 220 operatively coupled to the CPU 210. The accelerator 220 may include an accelerator engine 222 operable to perform functions (e.g., computations, etc.) offloaded from the CPU 210. The accelerator 220 may include an accelerator cache proxy 224 and a memory proxy 228.

[0081] The accelerator 220 and the CPU 210 may be configured according to various conventional hardware and / or memory access techniques, and / or the accelerator 220 and the CPU 210 may include various conventional hardware and / or memory access techniques. For example, as Figure 2a shown, all memory accesses, including those initiated by the accelerator 220, must pass through the path 230. The path 230 may include an incoherent link, such as a PCIe link. In the configuration of the apparatus 202, the accelerator engine 222 may be able to directly access the accelerator cache proxy 224 and the memory proxy 228, rather than the cache proxy 214, the host proxy 216, or the memory proxy 218. Similarly, the cores 212a-n will not be able to directly access the memory proxy 228. Thus, the memory behind the memory proxy 228 will not be part of the system address map seen by the cores 212a-n. Since the cores 212a-n cannot access the common memory proxy, data can only be exchanged through copies. In some implementations, a driver may be used to facilitate the copying of data back and forth between the memory proxies 218 and 228. For example, the driver may include runtime elements that create a shared memory abstraction that hides all copies from the programmer. Conversely, and as described in detail below, some embodiments may provide a configuration in which when the accelerator engine wants to access accelerator memory, for example, via the accelerator proxy 228, requests from the accelerator engine may be forced to cross the link between the accelerator and the CPU.

[0082] Figure 2b An example of an incoherent operating environment 200B is shown. Figure 2b The operating environment 200B depicted may include an accelerator 220 having an accelerator host proxy 226. The CPU 210 and the accelerator 220 may be operatively coupled via an incoherent path 232 (e.g., a UPI path or a CCIX path).

[0083] For operation of the apparatus 204, the accelerator engine 222 and the cores 212a-n can access both the memory agents 228 and 218. The cores 212a-n can access the memory 218 without crossing the link 232, and the accelerator agent 222 can access the memory 228 without crossing. The cost of these local accesses from 222 to 228 is that a host agent 226 needs to be built such that it can track the coherence of all accesses from the cores 212a-n to the memory 228. When the apparatus 204 includes multiple CPU 210 devices all connected by other instances of the link 232, this requirement leads to complexity and high resource usage. The host agent 226 needs to be able to track the coherence of all the cores 212a-n on all instances of the CPU 210. This can become quite costly in terms of performance, area, and power, especially for large configurations. Specifically, for accesses from the CPU 210, it adversely affects the performance efficiency of the accesses between the accelerator 222 and the memory 228, even if the accesses from the CPU 210 are expected to be relatively few.

[0084] Figure 2c An example of a coherence engine without a biased operating environment 200C is shown. As shown in FIG. 2, the apparatus 206 can include an accelerator 220 operatively coupled to the CPU 210 via coherence links 236 and 238. The accelerator 220 can include an accelerator engine 222 operable to perform functions offloaded from the CPU 210 (e.g., computations, etc.). The accelerator 220 can include an accelerator cache agent 224, an accelerator host agent 226, and a memory agent 228.

[0085] In the configuration of apparatus 206, accelerator 220 and CPU 210 may be configured and / or include various conventional hardware and / or memory access techniques according to various conventional hardware and / or memory access techniques, such as CCIX, GCM, standard coherence protocols (e.g., symmetric coherence protocol) and / or similar protocols. For example, as shown in FIG. 2, all memory accesses, including those initiated by accelerator 220, must pass through path 230. In this way, in order to access accelerator memory (e.g., via memory agent 228), accelerator 220 must pass through CPU 220 (and thus, the coherence protocol associated with the CPU). As a result, the apparatus may not be able to provide the ability to access certain memories, such as accelerator-attached memories associated with accelerator 220, as part of the system memory (e.g., as part of the system address map), which may allow host software to set operands and access the computation results of accelerator 220 without the overhead of, for example, IO direct memory access (DMA) data copies. Compared to memory access, such data copies may require driver calls, interrupts, and MMIO accesses, all of which are inefficient and complex. As Figure 2c shown, the inability to access accelerator-attached memory without cache coherence overhead may be detrimental to the execution time of computations offloaded to accelerator 220. For example, during a process involving a large amount of streaming write memory traffic, the cache coherence overhead can halve the effective write bandwidth seen by accelerator 220.

[0086] The efficiency of operand setting, result access, and accelerator computing plays a role in determining the effectiveness and benefits of offloading the work of CPU 210 to accelerator 220. If the cost of offloading the work is too high, then offloading may not be beneficial or may be limited to very large tasks. Accordingly, various developers have created accelerators that attempt to improve the efficiency of using an accelerator (e.g., accelerator 220) that has limited effectiveness compared to the techniques configured according to some embodiments. For example, certain conventional GPUs can operate without mapping the memory attached to the accelerator as part of the system address and may or may not use certain virtual memory configurations (e.g., SVM) to access the memory attached to the accelerator. Thus, in such systems, the memory attached to the accelerator is not visible to the host system software. Instead, the memory attached to the accelerator is accessed only through the runtime software layer provided by the GPU device driver. Data copy and page table manipulation systems are used to create the appearance of a system that enables virtual memory (e.g., SVM). Such systems are inefficient, especially compared to some embodiments, because the system requires memory copying, memory pinning, memory copying, and complex software, etc. These requirements result in a large amount of overhead at memory page translation points that are not required in systems configured according to some embodiments. In certain other systems, conventional hardware coherence mechanisms are used for memory operations associated with the memory attached to the accelerator, which limits the ability of the accelerator to access the memory attached to the accelerator at high bandwidth and / or limits the deployment options of a given accelerator (e.g., cannot support an accelerator attached through an on-package or off-package link without significant bandwidth loss).

[0087] Typically, traditional systems can access accelerator-attached memory using one of two methods: the full coherence (or full hardware coherence) method or the private memory model or method. The full coherence method requires that all memory accesses, including accesses to accelerator requests for accelerator-attached memory, must pass through the coherence protocol of the corresponding CPU. In this way, the accelerator must take a roundabout route to access the accelerator-attached memory because the request must be transmitted at least to the corresponding CPU, through the CPU coherence protocol, and then to the accelerator-attached memory. Thus, the full coherence method incurs a coherence overhead when the accelerator accesses its own memory that can significantly impair the data bandwidth that the accelerator can extract from its own attached memory. The private memory model requires significant resource and time costs, such as memory copying, page pinning requirements, page copy data bandwidth costs, and / or page translation costs (e.g., translation lookaside buffer (TLB) thrashing, page table operations, and / or the like). Accordingly, some embodiments can provide a coherence biasing process that is configured to provide multiple cache coherence processes that provide better memory utilization and improved performance, among other things, for a system that includes accelerator-attached memory as compared to traditional systems.

[0088] Figure 3 An example of an operating environment 300 that can represent various embodiments is shown. Figure 3 The operating environment 300 depicted can include a device 305 operable to provide a coherence biasing process according to some embodiments. In some embodiments, the device 305 can include a CPU 310 having a plurality of processing cores 312a-n and various protocol agents, such as a cache agent 314, a host agent 316, a memory agent 318, etc. The CPU 310 can be communicatively coupled to an accelerator 320 using various links 335, 340. The accelerator 320 can include an accelerator engine 312 and a memory agent 318, and can include or access biasing information 338.

[0089] As Figure 3As shown, the accelerator engine 322 can be directly communicatively coupled to the memory agent 328 via the biased coherent bypass 330. In various embodiments, the accelerator 320 can be configured to operate during device biasing, where the biased coherent bypass 330 can allow memory requests of the accelerator engine 322 to directly access the accelerator-attached memory (not shown) of the accelerator facilitated via the memory agent 328. In various embodiments, the accelerator 320 can be configured to operate during host biasing, where memory operations associated with the accelerator-attached memory can be processed via links 335, 340 using the cache coherence protocol of the CPU, e.g., via the cache agent 314 and the host agent 316. Thus, the accelerator 320 of the device 305 can utilize the coherence protocol of the CPU 310 when appropriate (e.g., when a non-accelerator entity requests the accelerator-attached memory), while allowing the accelerator 320 to directly access the accelerator-attached memory via the biased coherent bypass 330.

[0090] In some embodiments, the coherent biasing (e.g., whether device biasing or host biasing is in effect) can be stored in the biasing information 338. In various embodiments, the biasing information 338 can include and / or can be stored in various data structures, such as data tables (e.g., "biasing tables"). In some embodiments, the biasing information 338 can include a biasing indicator having a value indicating the effective biasing (e.g., 0 = host biasing, 1 = device biasing). In some embodiments, the biasing information 338 and / or the biasing indicator can be at various granularity levels, such as memory regions, page tables, address ranges, etc. For example, the biasing information 338 can specify that certain memory pages are set for device biasing while other memory pages are set for host biasing. In some embodiments, the biasing information 338 can include a biasing table configured to operate as a low-cost scalable snooping filter.

[0091] Figure 4 An example operating environment 400 is shown that can represent various embodiments. According to some embodiments, Figure 4The operating environment 400 depicted may include apparatus 405 operable to provide a coherent biasing process. The apparatus 405 may include an accelerator 410 communicatively coupled to a main processor 445 via a link (or multi-protocol link) 489. The accelerator 410 and the main processor 445 may communicate via the link using interconnect structures 415 and 450 respectively, which allows data and messages to be passed between them. In some embodiments, the link 489 may include a multi-protocol link that can be used to support multiple protocols. For example, the link 489 and the interconnect structures 415 and 450 may support various communication protocols, including but not limited to serial link protocols, device cache protocols, memory protocols, memory semantic protocols, directory bit support protocols, networking protocols, coherence protocols, accelerator protocols, data storage protocols, point-to-point protocols, fabric-based protocols, on-package (or on-chip) protocols, fabric-based on-package protocols, and / or similar protocols. Non-limiting examples of supported communication protocols may include PCI, PCIe, USB, SPI, SATA, QPI, UPI, OAP, IAL, IDI, IOSF, SMI, SMI3, etc. In some embodiments, the link 489 and the interconnect structures 415 and 450 may support in-device protocols (e.g., IDI) and memory interconnect protocols (e.g., SMI3). In various embodiments, the link 489 and the interconnect structures 415 and 450 may support in-device protocols (e.g., IDI), memory interconnect protocols (e.g., SMI3), and fabric-based protocols (e.g., IOSF).

[0092] In some embodiments, the accelerator 410 may include bus logic 435 having a device TLB 437. In some embodiments, the bus logic 435 may be or may include PCIe logic. In various embodiments, the bus logic 435 may communicate via the interconnect 480 using a fabric-based protocol (e.g., IOSF) and / or a peripheral component interconnect express (PCIe or PCI-E) protocol. In various embodiments, communication via the interconnect 480 may be used for various functions, including but not limited to discovery, register access (e.g., registers of the accelerator 410 (not shown)), configuration, initialization, interrupt, direct memory access, and / or address translation services (ATS).

[0093] The accelerator 410 may include a core 420 having a host memory cache 422 and an accelerator memory cache 424. The core 420 may communicate using an on-chip protocol (e.g., IDI) for various functions such as coherent requests and memory streams using the interconnect 481. In various embodiments, the accelerator 410 may include coherent logic 425 that includes or accesses bias mode information 427. The coherent logic 425 may communicate using, for example, a memory interconnect protocol (e.g., SMI3) using the interconnect 482. In some embodiments, communication over the interconnect 482 may be used for memory streams. The accelerator 410 may be operatively coupled to an accelerator memory 430 (e.g., as an accelerator-attached memory) that may store bias information 432.

[0094] In various embodiments, the host processor 445 may be operatively coupled to a host memory 440 and may include coherent logic (or coherent and cache logic) 455 having a last-level cache (LLC) 457. The coherent logic 455 may communicate using various interconnects, e.g., interconnects 484 and 485. In some embodiments, the interconnects 484 and 485 may include a memory interconnect protocol (e.g., SMI3) and / or an on-chip protocol (e.g., IDI). In some embodiments, the LLC 457 may include a combination of at least a portion of the host memory 440 and the accelerator memory 430.

[0095] The host processor 445 may include bus logic 460 having an input-output memory management unit (IOMMU) 462. In some embodiments, the bus logic 460 may be or may include PCIe logic. In various embodiments, the bus logic 460 may communicate using a fabric-based protocol (e.g., IOSF) and / or a peripheral component interconnect express (PCIe or PCI-E) protocol over the interconnects 486 and 488. In various embodiments, the host processor 445 may include multiple cores 465a-n, each core having a cache 467a-n. In some embodiments, the cores 465a-n may include an architecture (IA) core. Each of the cores 465a-n may communicate with the coherent logic 455 via interconnects 487a-n. In some embodiments, the interconnects 487a-n may support an on-chip protocol (e.g., IDI). In various embodiments, the host processor may include a device 470 operable to communicate with the bus logic 460 over the interconnect 488. In some embodiments, the device 470 may include an IO device, e.g., a PCIe IO device.

[0096] In some embodiments, the apparatus 405 is operable to perform a coherent biasing process applicable to various configurations, such as a system having an accelerator 410 and a host processor 445 (e.g., a computer processing complex including one or more computer processor chips), where the accelerator 410 is communicatively coupled to the host processor 445 via a multi-protocol link 489, and where memories are directly attached to the accelerator 410 and the host processor 445 (e.g., accelerator memory 430 and host memory 440, respectively). The coherent biasing process provided by the apparatus 405 can provide several technical advantages over conventional systems, such as providing the accelerator 410 and “host” software running on processing cores 465a-n with access to the accelerator memory 430. The coherent biasing process provided by the apparatus can include a host biasing process and a device biasing process (collectively, the biasing protocol flow) and multiple options for modulating and / or selecting the biasing protocol flow for a particular memory access.

[0097] In some embodiments, the biasing protocol flow can be implemented at least in part using protocol layers (e.g., “biasing protocol layer”) on the multi-protocol link 489. In some embodiments, the biasing protocol layer can include: an in-device protocol (e.g., IDI) and / or a memory interconnect protocol (e.g., SMI3). In some embodiments, the biasing protocol flow can be enabled by adding new information to the biasing protocol layer and / or adding support for the protocol by using various information of the biasing protocol layer. For example, the biasing protocol flow (e.g., a conventional multi-protocol link may only include an in-device protocol (e.g., IDI) and a fabric-based protocol (for, e.g., IOSF)) can be implemented by using existing opcodes for the in-device protocol (e.g., IDI), adding opcodes to the memory interconnect protocol (e.g., SMI3) standard and / or adding support for the memory interconnect protocol (e.g., SMI3) on the multi-protocol link 489.

[0098] In some embodiments, device 405 may be associated with at least one operating system (OS). The OS may be configured to not use accelerator memory 430 or certain portions of accelerator memory 430. Such an OS may include support for a "memory-only NUMA module" (e.g., without a CPU). Device 405 may execute drivers (e.g., including driver 118) to perform various accelerator memory services. Illustrative and non-limiting accelerator memory services implemented in the driver may include driver discovery and / or acquisition / assignment of accelerator memory 430, providing an allocation API and mapping pages via the OS page mapping service, providing procedures for managing multi-process memory oversubscription and job scheduling, providing an API to allow software applications to set and change the bias mode of the storage regions of accelerator memory 430, and / or a deallocation API to return pages to the driver's free page list and / or return pages to the default bias mode.

[0099] Figure 5a An example of an operating environment 500 that may represent various embodiments is shown. According to some embodiments, Figure 5a the operating environment 500 depicted may provide a host bias process flow. As Figure 5a shown, device 505 may include a CPU 510 communicatively coupled to an accelerator 520 via a link 540. In some embodiments, link 540 may include a multi-protocol link. CPU 510 may include a coherence controller 530 and may be communicatively coupled to host memory 512. In various embodiments, coherence controller 530 may be used to provide one or more standard cache coherence protocols. In some embodiments, coherence controller 530 may include and / or be associated with various agents, such as a host agent. In some embodiments, CPU 510 may include and / or be communicatively coupled to one or more IO devices. Accelerator 520 may be communicatively coupled to accelerator memory 522.

[0100] The host bias process flows 550 and 560 can include a set of data streams that funnel all requests, including requests from the accelerator 520, to the accelerator memory 522 through the coherence controller 530 in the CPU 510. In this way, the accelerator 522 takes a detour to access the accelerator memory 522, but allows access from both the accelerator 522 and the CPU 510 (including requests from the I / O device via the CPU 510) to be kept coherent using the standard cache coherence protocol of the coherence controller 530. In some embodiments, the host bias process flows 550 and 560 can use an in-device protocol (e.g., IDI). In some embodiments, the host bias process flows 550 and 560 can use the standard opcodes of an in-device protocol (e.g., IDI), e.g., by issuing requests to the coherence controller 530 over the multi-protocol link 540. In various embodiments, the coherence controller 530 can issue various coherence messages (e.g., snoops) generated by requests from the accelerator 520 to all peer processor chips and internal processor agents on behalf of the accelerator 520. In some embodiments, the various coherence messages can include point-to-point protocol (e.g., UPI) coherence messages and / or in-device protocol (e.g., IDI) messages.

[0101] In some embodiments, the coherence controller 530 can conditionally issue memory access messages to the accelerator memory controller (not shown) of the accelerator 520 over the multi-protocol link 540. Such memory access messages can be the same or substantially similar to the memory access messages that the coherence controller 530 can send to the CPU memory controller (not shown) and can include a new opcode that allows data to be returned directly to an agent inside the accelerator 520 instead of forcing the data to be returned to the coherence controller and then back to the accelerator 520 as an in-device protocol (e.g., IDI) response over the multi-protocol link 540.

[0102] The host bias process flow 550 can include a process generated by a request or memory operation for the accelerator memory 522 that originates from the accelerator. The host bias process path 560 can include a flow generated by a request or memory operation for the accelerator memory 522 that originates from the CPU 510 (or an I / O device or a software application associated with the CPU 510). When the device 505 is active in the host bias mode, the host bias process flows 550 and 560 can be used to access the accelerator memory 522, as Figure 5aAs shown. In various embodiments, in the host bias mode, all requests from the CPU 510 targeted at the accelerator memory 522 can be directly sent to the coherence controller 530. The coherence controller 530 can apply standard cache coherence protocols and send standard cache coherence messages. In some embodiments, the coherence controller 530 can send memory interconnect protocol (e.g., SMI3) commands for such requests through the multi-protocol link 540, where the memory interconnect protocol (e.g., SMI3) returns data through the multi-protocol link 540.

[0103] Figure 5b Another example of an operating environment 500 that can represent various embodiments is shown. According to some embodiments, Figure 5a The operating environment 500 depicted in FIG. 5 can provide a device bias process flow. As shown in FIG. 5, when the device 505 is active in the device bias mode, the device bias path 570 can be used to access the accelerator memory 522. For example, the device bias flow or path 570 can allow the accelerator 520 to directly access the accelerator memory 522 without consulting the coherence controller 530. More specifically, the device bias path 570 can allow the accelerator 520 to directly access the accelerator memory 522 without having to send requests through the multi-protocol link 540.

[0104] In the device bias mode, according to some embodiments, CPU 510 requests for the accelerator memory can be issued in the same or substantially similar manner as described for the host bias mode, but are different in the memory interconnect protocol (e.g., SMI3) portion of the path 580. In some embodiments, in the device bias mode, CPU 510 requests for attached memory can be completed as if they were issued as "uncached" requests. Generally, data for requests that are uncached during the device bias mode is not cached in the CPU cache hierarchy. In this way, the accelerator 520 is allowed to access data in the accelerator memory 522 during the device bias mode without consulting the coherence controller 530 of the CPU 510. In some embodiments, the uncached requests can be implemented on the device internal protocol (e.g., IDI) bus of the CPU 510. In various embodiments, a global observed once-use (GO-UO) protocol on the device internal protocol (e.g., IDI) bus of the CPU 510 can be used to implement the uncached requests. For example, the response to an uncached request can return a piece of data to the CPU 510 and indicate that the CPU 510 should use that piece of data only once, e.g., to prevent caching of that piece of data and support the use of an uncached data stream.

[0105] In some embodiments, the apparatus 505 and / or the CPU 510 may not support GO-UO. In such embodiments, an un-cached flow (e.g., path 580) can be implemented using a multi-message response sequence on the multi-protocol link 540 and the memory interconnect protocol (e.g., SMI3) of the CPU 510 internal device protocol (e.g., IDI) bus. For example, when the "device bias" page of the CPU 510 silver accelerator 520 is targeted, the accelerator 520 can set one or more states to block future requests to the target memory region (e.g., cache line) from the accelerator 520 and send a "device bias hit" response on the memory interconnect protocol (e.g., SMI3) line of the multi-protocol link 540. In response to the "device bias hit" message, the coherence controller 530 (or its agent) can return the data to the requesting processor core, followed by a snoop invalidate message. When the corresponding processor core acknowledges the completion of the snoop invalidate, the coherence controller 530 (or its agent) can send a "device bias block complete" message to the accelerator 520 on the memory interconnect protocol (e.g., SMI3) line of the multi-protocol link 540. In response to receiving the "device bias block complete" message, the accelerator can clear the corresponding blocking state.

[0106] Reference Figure 4 , the bias mode information 427 can include a bias indicator that is configured to indicate a valid bias mode (e.g., device bias mode or host bias mode). The selection of the valid bias mode can be determined by the bias information 432. In some embodiments, the bias information 432 can include a bias table. In various embodiments, the bias table can include bias information 432 for certain regions of the accelerator memory, such as pages, lines, etc. In some embodiments, the bias table can include bits (e.g., 1 or 3 bits) for each accelerator memory 430 memory page. In some embodiments, the bias table can be implemented using RAM, such as SRAM at the accelerator 410 and / or the stolen range of the accelerator memory 430, with or without a cache inside the accelerator 410.

[0107] In some embodiments, the bias information 432 may include bias table entries in a bias table. In various embodiments, the bias table entries associated with each access to the accelerator memory 430 may be accessed prior to the actual access to the accelerator memory 430. In some embodiments, local requests from the accelerator 410 that find pages in its device bias may be forwarded directly to the accelerator memory 430. In various embodiments, local requests from the accelerator 410 that find pages in its host bias may be forwarded to the host processor 445, e.g., as an in-device protocol (e.g., IDI) request on the multi-protocol link 489. In some embodiments, e.g., using a memory interconnect protocol (e.g., SMI3), host processor 445 requests that find pages in its device bias may use an uncached stream (e.g., Figure 5b of path 580) to complete the request. In some embodiments, e.g., using a memory interconnect protocol (e.g., SMI3), host processor 445 requests that find pages in its host bias may complete the request as a standard memory read of the accelerator memory (e.g., via Figure 5a of path 560).

[0108] The bias mode of the bias indicator of the bias mode information 427 for a region (e.g., a memory page) of the accelerator memory 430 may be changed by a software-based system, a hardware-assisted system, a hardware-based system, or a combination thereof. In some embodiments, the bias indicator may be changed via an application programming interface (API) call (e.g., OpenCL), which in turn may call the accelerator 410 device driver (e.g., driver 118). The accelerator 410 device driver may send a message (or enqueue a command descriptor) to the accelerator 410, instructing the accelerator 410 to change the bias indicator. In some embodiments, a change in the bias indicator may be accompanied by a cache flush operation in the main processor 445. In various embodiments, a transition from a host bias mode to a device bias mode may require a cache flush operation, but a cache flush operation may not be required for a transition from a device bias mode to a host bias mode. In various embodiments, software may change the bias mode of one or more memory regions of the accelerator memory 430 via a work request sent to the accelerator 430.

[0109] In some cases, software may not be able to or may not easily determine when to make a bias conversion API call and identify the memory regions that require bias conversion. In such cases, the accelerator 410 can provide a bias conversion hint process, where the accelerator 410 determines the need for bias conversion and sends a message indicating the need for bias conversion to the accelerator driver (e.g., driver 118). In various embodiments, the bias conversion hint process can be activated in response to a bias table lookup that triggers the accelerator 410 to access a host bias mode memory region or the host processor 445 to access a device bias mode memory region. In some embodiments, the bias conversion hint process can signal the need for bias conversion to the accelerator driver via an interrupt. In various embodiments, the bias table can include bias status bits for enabling bias conversion status values. The bias status bits can be used to allow access to memory regions during the process of changing the bias (e.g., when partially flushing the cache and having to suppress incremental cache pollution due to subsequent requests).

[0110] One or more logic flows are included herein that represent exemplary methods for performing novel aspects of the disclosed architecture. While, for purposes of simplifying the description, one or more of the methods shown herein are illustrated and described as a series of acts, those skilled in the art will understand and appreciate that these methods are not limited by the order of acts. Accordingly, some acts may occur in a different order and / or concurrently with other acts shown and described herein. For example, those skilled in the art will understand and appreciate that a method may alternatively be represented as a series of interrelated states or events, such as in a state diagram. Additionally, not all acts shown in the methods may be required for a novel implementation.

[0111] The logic flows can be implemented in software, firmware, hardware, or any combination thereof. In software and firmware embodiments, the logic flows can be implemented by computer-executable instructions stored on a non-transitory computer-readable medium or machine-readable medium (e.g., optical, magnetic, or semiconductor memory). The embodiments are not limited to this context.

[0112] Figure 6 An embodiment of a logic flow 600 is shown. The logic flow 600 can represent some or all of the operations performed by one or more of the embodiments described herein, such as apparatuses 105, 305, 405, and 505. In some embodiments, the logic flow 600 can represent some or all of the operations for a coherent bias process according to some embodiments.

[0113] As Figure 6As shown, at block 602, logic flow 600 may set the bias mode of the accelerator memory page to the host bias mode. For example, a host software application (e.g., software application 140) may set the bias mode of the accelerator device memory 430 to the host bias mode via a driver and / or API call. The host software application may use an API call (e.g., OpenCL API) to convert an allocated (or target) page of the accelerator memory 430 storing operands to host bias. Since the allocated page is being converted from the device bias mode to the host bias mode, a cache flush is not initiated. The device bias mode may be specified in the bias table of the bias information 432.

[0114] At block 604, logic flow 600 may push operands and / or data to the accelerator memory page. For example, the accelerator 420 may perform a function for the CPU that requires certain operands. The host software application may push the operands from a peer CPU core (e.g., core 465a) to the allocated page of the accelerator memory 430. The host processor 445 may generate operand data in the allocated page in the accelerator memory 430 (and anywhere in the host memory 440).

[0115] At block 606, logic flow 600 may convert the accelerator memory page to the device bias mode. For example, the host software application may use an API call to convert the operand memory page of the accelerator memory 430 to the device bias mode. When the device bias conversion is complete, the host software application may submit work to the accelerator 430. The accelerator 430 may perform the function associated with the submitted work without host-related coherence overhead.

[0116] At block 608, logic flow 600 may generate a result using the operands by the accelerator and store the result in the accelerator memory page. For example, the accelerator 420 may use the operands to perform a function (e.g., floating point operations, graphics calculations, FFT operations, and / or similar functions) to generate a result. The result may be stored in the accelerator memory 430. Additionally, the software application may use an API call to cause the work descriptor submission to flush the operand page from the host cache. In some embodiments, a cache (or cache line) flush routine (such as CLFLUSH) on an in-device protocol (e.g., IDI protocol) may be used to perform the cache flush. The result generated by this function may be stored in the allocated accelerator memory 430 page.

[0117] At block 610, the logic flow may set the bias mode of the accelerator memory page storing the result to the host bias mode. For example, a host software application may use an API call to convert the operand memory page of the accelerator memory 430 to the host bias mode without causing a coherence process and / or cache flush operation. The host CPU 445 may access, cache, and share the result. At block 612, the logic flow 600 may provide the result from the accelerator memory page to the host software. For example, a host software application may directly access the result from the accelerator memory page 430. In some embodiments, the allocated accelerator memory page may be freed by the logic flow. For example, a host software application may use a driver and / or an API call to free the allocated memory page of the accelerator memory 430.

[0118] Figure 7 is a block diagram showing a structure according to one or more examples of the present specification. In this case, a coherent accelerator structure 700 is provided. The coherent accelerator structure 700 is interconnected with an IAL endpoint 728, and the IAL endpoint 728 communicatively couples the coherent accelerator structure 700 to a host device, such as the host device disclosed in the foregoing figures.

[0119] The coherent accelerator structure 700 is provided to communicatively couple the accelerator 740 and its attached memory 722 to the host device. The memory 722 includes a plurality of memory controllers 720-1 to 720-n. In one example, 8 memory controllers 720 may serve 8 separate memory banks.

[0120] The structure controller 736 includes a set of controllers and interconnections to provide the coherent memory structure 700. In this example, the structure controller 736 is divided into n separate slices to serve the n memory banks of the memory 722. Each slice may be substantially independent of each other slice. As described above, the structure controller 736 includes both "vertical" interconnections 706 and "horizontal" interconnections 708. Vertical interconnections are generally understood to connect upstream or downstream devices to each other. For example, the last-level cache (LLC) 734 is vertically connected to the LLC controller 738, which is connected to the in-die interconnect (F2IDI) block that communicatively couples the structure controller 736 to the accelerator 740. The F2IDI 730 provides a downstream link to the structure stop 712 and may also provide a bypass interconnect 715. The bypass interconnect 715 directly connects the LLC controller 738 to the structure-to-memory interconnect 716, where signals are multiplexed to the memory controller 720. In a non-bypass route, requests from the F2IDI 730 travel along the horizontal interconnect to the host, then back to the structure stop 712, then to the structure coherence engine 704, and down to the F2MEM 716.

[0121] The horizontal bus includes a bus that interconnects the fabric stops 712 with each other and connects the LLC controllers to each other.

[0122] In one example, the IAL endpoint 728 can receive a packet from a host device that includes an instruction to perform an acceleration function, as well as a payload that includes a snoop for the accelerator to operate on. The IAL endpoint 728 passes these to the L2FAB 718, which acts as an interconnect for the fabric controllers 736 of the host device. The L2FAB 718 can act as a link controller for the fabric, including providing an IAL interface controller (although in some embodiments, additional IAL control elements may also be provided, and generally, any combination of elements that provide IAL interface control can be referred to as an "IAL interface controller"). The L2FAB 718 controls requests from the accelerator to the host and vice versa. The L2FAB 718 can also act as an IDI proxy and may be required to act as an ordering proxy between IDI requests from the accelerator and snoops from the host.

[0123] Then, the L2FAB 718 can operate the fabric stop 712-0 to fill values into the memory 722. The fabric stop L2FAB 718 can apply a load balancing algorithm, e.g., hashing based on a simple address, to tag the payload data for a particular destination bank. Once the banks in the memory 722 are filled with the appropriate data, the accelerator 740 operates the fabric controller 736 to fetch values from the memory to the LLC 734 via the LLC controller 738. The accelerator 740 performs its acceleration calculation and then writes the outputs to the LLC 734, where they are then passed downstream and written out to the memory 722.

[0124] In some examples, the fabric stops 712, the F2MEM controller 716, the multiplexer 710, and the F2IDI 730 can all be standard buses and interconnects that provide interconnections according to well-known principles. The aforementioned interconnections can provide virtual and physical channels, interconnections, buses, switching elements, and flow control mechanisms. They can also provide a conflict resolution mechanism related to the interaction between requests issued by the accelerator or device proxy and requests issued by the host. The fabric can include a physical bus in the horizontal direction, with server ring switching as the bus passes through the individual dies. The fabric can also include a specially optimized horizontal interconnect 739 between the LLC controllers 738.

[0125] Requests from the F2IDI 730 can be passed through hardware to split and multiplex traffic to the host between the horizontal fabric interconnect and each slice-optimized path between the LLC controller 738 and the memory 722. This includes multiplexing the traffic and directing it to the IDI block, where the traffic traverses the traditional route through the fabric stop 712 and the FCE 704, or using IDI stuffing to direct the traffic to the bypass interconnect 715. The F2IDI 730-1 can also include hardware for managing the ingress and egress to and from the horizontal fabric interconnect, such as by providing appropriate signals to the fabric stop 712.

[0126] The IAL interface controller 718 can be a suitable PCIe controller. The IAL interface controller provides an interface between the packetized IAL bus and the fabric interconnect. It is responsible for queuing and providing flow control for IAL messages and directing the IAL messages to the appropriate fabric physical and virtual channels. The L2FAB 718 can also provide arbitration between multiple classes of IAL messages. It can further enforce IAL ordering rules.

[0127] At least three control structures within the fabric controller 736 provide novel and advantageous features of the fabric controller 736 of this specification. These include the LLC controller 738, the FCE 704, and the power management module 750.

[0128] Advantageously, the LLC controller 738 can also provide a bias control function according to the IAL bias protocol. Thus, the LLC controller 738 can include hardware for performing cache lookups, hardware for checking the IAL base for cache miss requests, hardware for steering requests onto appropriate interconnect paths, and logic for responding to snoopings issued by the host processor or by the FCE 704.

[0129] When steering requests from the fabric stop 712 to the host through the L2FAB 718, the LLC controller 738 determines where the traffic should be directed through the fabric stop 712, directly to the F2MEM 716 through the bypass interconnect 715, or to another memory controller through the horizontal bus 739.

[0130] Note that in some embodiments, the LLC controller 738 is a device or block physically separate from the FCE 704. A single block providing the functions of both the LLC controller 738 and the FCE 704 can be provided. However, by separating the two blocks and providing the IAL bias logic in the LLC controller 738, the bypass interconnect 715 can be provided, thus accelerating certain memory operations. Advantageously, in some embodiments, separating the LLC controller 738 and the FCE 704 can also assist in selective power gating in parts of the fabric to use resources more efficiently.

[0131] The FCE 704 may include hardware for queuing, processing (e.g., issuing snoop to the LLC), and tracking SMI requests from the host. This provides consistency with the host device. The FCE 704 may also include hardware for queuing requests per chip, an optimized path to the banks within the memory 722. Embodiments of the FCE may also include hardware for arbitrating and multiplexing the above two request classes onto the CMI memory subsystem interface, and may include hardware or logic for resolving conflicts between the above two request classes. Other embodiments of the FCE may provide support for ordering requests from direct vertical interconnects and requests from the FCE 704.

[0132] The power management module (PMM) 750 also provides advantages for embodiments of this specification. For example, consider the case where each individual chip in the fabric controller 736 vertically supports a bandwidth of 1 GB per second. 1 GB per second is provided only as an illustrative example, and a real example of the fabric controller 736 may be much faster or much slower than 1 GB per second.

[0133] The LLC 734 may have a higher bandwidth, e.g., 10 times the bandwidth of the vertical bandwidth of one chip of the fabric controller 736. Thus, the LLC 734 may have a bandwidth of 10 GB per second, which may be bidirectional, such that the total bandwidth through the LLC 734 is 20 GB per second. Thus, with 8 chips of the fabric controller 736 supporting 20 GB per second bidirectional, the accelerator 740 can see a total bandwidth of 160 GB per second through the horizontal bus 739. Thus, running the LLC controller 738 and the horizontal bus 739 at full speed consumes a large amount of power.

[0134] However, as described above, the vertical bandwidth may be 1 GB per chip, and the total IAL bandwidth may be approximately 10 GB per second. Thus, the bandwidth provided by the horizontal bus 739 is approximately an order of magnitude higher than the bandwidth of the entire fabric controller 736. For example, the horizontal bus 739 may include thousands of physical lines, while the vertical interconnect may include hundreds of physical lines. The horizontal fabric 708 may support the full bandwidth of the IAL, i.e., 10 GB per second in each direction, 20 GB per second in total.

[0135] The accelerator 740 can perform computations at a much higher rate than the host device can consume data and operate on the LLC 734. Thus, data can burst into the accelerator 740 and then be consumed by the main processor as needed. Once the accelerator 740 has completed its computations and filled the appropriate values in the LLC 734, maintaining full bandwidth between the LLC controllers 738 consumes a large amount of power, which is essentially wasted because the LLC controllers 738 no longer need to communicate with each other when the accelerator 740 is idle. Thus, when the accelerator 740 is idle, the LLC controllers 738 can be powered down, thereby turning off the horizontal bus 739 while keeping the appropriate vertical buses active, such as from the fabric stop 712 to the FCE 704 to the F2MEM 716 to the memory controller 720, and also keeping the horizontal bus 708. Since the horizontal bus 739 operates at approximately an order of magnitude or more higher than the rest of the fabric 700, this can save approximately an order of magnitude of power when the accelerator 740 is idle.

[0136] Note that some embodiments of the coherent accelerator fabric 700 may also provide an isochronous controller, which can be used to provide isochronous traffic to latency or time-sensitive elements. For example, if the accelerator 740 is a display accelerator, an isochronous display path can be provided to a display generator (DG) such that the connected display receives isochronous data.

[0137] The overall combination of agents and interconnects in the coherent accelerator fabric 700 implements the IAL function in a high-performance, deadlock-free, and starvation-free manner. It provides increased efficiency while saving energy by bypassing the interconnect 715.

[0138] Figure 8 is a flowchart of a method 800 according to one or more examples of the present specification. The method 800 illustrates a method of power saving, such as can be provided by Figure 7 the PMM 750.

[0139] Inputs from the host device 804 can reach the coherent accelerator fabric, including instructions to perform computations and payloads for the computations. In block 808, if the horizontal interconnect between the LLC controllers is powered down, the PMM powers the interconnect to its full bandwidth.

[0140] In block 812, the accelerator computes results according to its normal function. In computing these results, it can operate the coherent accelerator fabric at its full available bandwidth, including the full bandwidth of the horizontal interconnect between the LLC controllers.

[0141] When the results are complete, in block 816, the accelerator fabric can flush the results to the local memory 820.

[0142] In decision block 824, the PMM determines whether there is new data available from the host that can be operated on. If any new data is available, control returns to block 812 and the accelerator continues to perform its acceleration function. At the same time, the host device can directly consume data from the local memory 820, which can be mapped to the host memory address space in a coherent manner.

[0143] Returning to block 824, if no new data is available from the host, then in block 828, the PMM reduces power, for example by turning off the LLC controller, thereby disabling the high-bandwidth level interconnection between the LLC controllers. As described above, since the local memory 820 is mapped to the host address memory space, the host can continue to consume data from the local memory 820 at the full IAL bandwidth, which in some embodiments is much lower than the full bandwidth between the LLC controllers.

[0144] In block 832, the controller waits for new input from the host device and, when new data is received, can power up the interconnection for standby.

[0145] Figures 9 - 11 An example of the IAL.mem tunnel via PCIe is shown. The described packet format includes standard PCIe packet fields, except for the fields highlighted in gray. The gray fields are the fields that provide the new tunnel area.

[0146] Figure 9 is a block diagram of an IAL.mem read operation via PCIe according to one or more examples of this specification. The new fields include:

[0147] · MemOpcode (4 bits) - Memory operation code. Contains information about the memory transaction that needs to be processed. For example, read, write, no operation, etc.

[0148] · MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together they specify which metadata field in the memory needs to be modified and what value to modify it to. Metadata fields in the memory typically contain information associated with the actual data. For example, QPI stores directory status in the metadata.

[0149] · TC (2 bits) - Traffic class. Used to distinguish traffic belonging to different quality of service classes.

[0150] · Snp type (3 bits) - Snoop type. Used to maintain coherence between the caches of the host and the device.

[0151] · R (5 bits) - Reserved

[0152] Figure 10It is a block diagram of IAL.mem writes via PCIe operations according to one or more examples of this specification. The new fields include:

[0153] · MemOpcode (4 bits) - Memory opcode. Contains information about the memory transaction to be processed. For example, read, write, no operation, etc.

[0154] · MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together they specify which metadata field in the memory needs to be modified and to what value. Metadata fields in memory usually contain information associated with the actual data. For example, QPI stores directory status in the metadata.

[0155] · TC (2 bits) - Traffic class. Used to distinguish traffic belonging to different quality of service classes.

[0156] · Snp type (3 bits) - Snoop type. Used to maintain coherence between the caches of the host and the device.

[0157] · R (5 bits) - Reserved

[0158] Figure 11 It is a block diagram of IAL.mem data completion via PCIe operations according to one or more examples of this specification. The new fields include:

[0159] · R (1 bit) - Reserved

[0160] · Opcode (3 bits) - IAL.io opcode

[0161] · MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together they specify which metadata field in the memory needs to be modified and to what value. Metadata fields in memory usually contain information associated with the actual data. For example, QPI stores directory status in the metadata.

[0162] · PCLS (4 bits) - Previous cache line state. Used to discern coherence transitions.

[0163] · PRE (7 bits) - Performance encoding. Used by the performance monitoring counters in the host.

[0164] Figure 12An embodiment of a structure composed of point - to - point links interconnecting a set of components in accordance with one or more examples of this specification is shown. System 1200 includes a processor 1205 and a system memory 1210 coupled to a controller hub 1215. The processor 1205 includes any processing element, such as a microprocessor, a main processor, an embedded processor, a coprocessor, or other processors. The processor 1205 is coupled to the controller hub 1215 via a front - side bus (FSB) 1206. In one embodiment, the FSB 1206 is a serial point - to - point interconnect as described below. In another embodiment, the link 1206 includes a serial differential interconnect architecture compliant with a differential interconnect standard.

[0165] The system memory 1210 includes any memory device, such as random access memory (RAM), non - volatile (NV) memory, or other memory accessible to the devices in the system 1200. The system memory 1210 is coupled to the controller hub 1215 via a memory interface 1216. Examples of memory interfaces include double data rate (DDR) memory interfaces, dual - channel DDR memory interfaces, and dynamic RAM (DRAM) memory interfaces.

[0166] In one embodiment, the controller hub 1215 is a root hub, root complex, or root controller in a Peripheral Component Interconnect Express (PCIe) interconnect hierarchy. Examples of the controller hub 1215 include a chipset, a memory controller hub (MCH), a north bridge, an interconnect controller hub (ICH), a south bridge, and a root controller / hub. Generally, the term chipset refers to two physically separate controller hubs, namely, a memory controller hub (MCH) coupled to an interconnect controller hub (ICH).

[0167] Note that current systems typically include an MCH integrated with the processor 1205, and the controller 1215 communicates with I / O devices in a manner similar to that described below. In some embodiments, peer - to - peer routing is optionally supported via the root complex 1215.

[0168] Here, the controller hub 1215 is coupled to a switch / bridge 1220 via a serial link 1219. Input / output modules 1217 and 1221 (which may also be referred to as interfaces / ports 1217 and 1221) include / implement a hierarchical protocol stack to provide communication between the controller hub 1215 and the switch 1220. In one embodiment, multiple devices can be coupled to the switch 1220.

[0169] Switch / bridge 1220 routes packets / messages from device 1225 upstream (i.e., towards the root complex) to controller hub 1215 and downstream (i.e., away from the root controller down the hierarchy) from processor 1205 or system memory 1210 to device 1225. In one embodiment, switch 1220 is referred to as a logical component of multiple virtual PCI-to-PCI bridge devices.

[0170] Device 1225 includes any internal or external device or component to be coupled to an electronic system, such as an I / O device, network interface controller (NIC), add-in card, audio processor, network processor, hard disk drive, storage device, CD / DVD ROM, monitor, printer, mouse, keyboard, router, portable storage device, Firewire device, universal serial bus (USB) device, scanner, and other input / output devices. Typically in PCIe parlance, such a device is referred to as an endpoint. Although not specifically shown, device 1225 may include a PCIe-to-PCI / PCI-X bridge to support legacy or other versions of PCI devices. Endpoint devices in PCIe are typically classified as legacy, PCIe, or root complex integrated endpoints.

[0171] Accelerator 1230 is also coupled to controller hub 1215 via serial link 1232. In one embodiment, graphics accelerator 1230 is coupled to the MCH, and the MCH is coupled to the ICH. Then, switch 1220 and thus I / O device 1225 are coupled to the ICH. I / O modules 1231 and 1218 are also used to implement a hierarchical protocol stack for communication between graphics accelerator 1230 and controller hub 1215. Similar to the MCH discussion above, the graphics controller or graphics accelerator 1230 itself may be integrated within processor 1205.

[0172] In some embodiments, accelerator 1230 may be an accelerator, such as Figure 7 accelerator 740, which provides coherent memory for processor 1205.

[0173] To support IAL over PCIe, controller hub 1215 (or another PCIe controller) may include extensions to the PCIe protocol, including by way of non-limiting example, mapping engine 1240, tunneling engine 1242, host-to-device bias flip engine 1244, and QoS engine 1246.

[0174] The mapping engine 1240 can be configured to provide opcode mapping between PCIe instructions and IAL.io (IOSF) opcodes. IOSF provides an incoherent ordered semantic protocol and can provide services such as device discovery, device configuration, error reporting, interrupt provision, interrupt handling, and DMA-style data transfer by way of non-limiting examples. Local PCIe can provide corresponding instructions, and thus in some cases, the mapping can be a one-to-one mapping.

[0175] The tunneling engine 1242 provides an IAL.mem (SMI) tunnel over PCIe. This tunnel enables a host (e.g., a processor) to map accelerator memory into the host memory address space and read from and write to the accelerator memory in a coherent manner. SMI is a transactional memory interface that can be used by a coherent engine on the host to transmit IAL transactions over the PCIe tunnel in a coherent manner. An example of a modified packet structure for such a tunnel is shown in Figures 9 - 11 In some cases, special fields for this tunnel can be allocated within one or more DVSEC fields of the PCIe packet.

[0176] The host bias to device bias flip engine 1244 provides the accelerator device with the ability to flush host cache lines (required for host-to-device bias flip). This can be done using cache line granularity un-allocated zero-length writes from the accelerator device over PCIe (i.e., writes with no byte enables set). Un-allocated semantics can be described using transactions and processing hints on transaction layer packets (TLPs). For example:

[0177] ·TH = 1, PH = 01

[0178] This enables the device to invalidate a given cache line so that it can access its own storage space without losing coherence. The device can issue a read after the page bias flip to ensure all lines are flushed. The device can also implement a CAM to ensure that while the flip is in progress, no new requests to the line are received from the host.

[0179] The QoS engine 1246 can divide IAL traffic into two or more virtual channels to optimize the interconnect. For example, these can include a first virtual channel (VC0) for MMIO and configuration operations, a second virtual channel (VC1) for host-to-device writes, and a third virtual channel (VC2) for host reads from the device.

[0180] Figure 13An embodiment of a hierarchical protocol stack in accordance with one or more embodiments of the present specification is shown. The hierarchical protocol stack 1300 includes any form of hierarchical communication stack, such as a QuickPath Interconnect (QPI) stack, a PCie stack, a next-generation high-performance computing interconnect stack, or other hierarchical stacks. Although the discussion below is given with reference to Figures 12 - 15 a PCIe stack, the same concepts can be applied to other interconnect stacks. In one embodiment, the protocol stack 1300 is a PCIe protocol stack, including a transaction layer 1305, a link layer 1310, and a physical layer 1320.

[0181] Interfaces such as Figure 12 interfaces 1217, 1218, 1221, 1222, 1226, and 1231 in can be represented as the communication protocol stack 1300. A representation as a communication protocol stack can also be referred to as a module or interface that implements / includes the protocol stack.

[0182] PCIe uses packets to transfer information between components. Packets are formed in the transaction layer 1305 and the data link layer 1310 to transfer information from a sending component to a receiving component.

[0183] Since the transmitted packets flow through other layers, they are extended with additional information required to process the packets at those layers. On the receiving side, the reverse process occurs, and the packets are transformed from their physical layer 1320 representation to the data link layer 1310 representation, and finally (for transaction layer packets) to a form that can be processed by the transaction layer 1305 of the receiving device.

[0184] Transaction Layer

[0185] In one embodiment, the transaction layer 1305 is used to provide an interface between the processing core of a device and the interconnect architecture, such as the data link layer 1310 and the physical layer 1320. In this regard, the main responsibility of the transaction layer 1305 is the assembly and disassembly of packets, i.e., transaction layer packets (TLPs). The transaction layer 1305 typically manages credit-based flow control for TLPs. PCIe implements split transactions, i.e., transactions with requests and responses separated in time, allowing the link to carry other traffic while the target device is collecting the data for the response.

[0186] In addition, PCIe utilizes credit-based flow control. In this scheme, a device advertises an initial credit amount for each receive buffer in the transaction layer 1305. At an external device at the opposite end of the link, such as Figure 1 the controller hub 115 in, the number of credits consumed by each TLP is calculated. If the transaction does not exceed the credit limit, the transaction can be sent. After receiving a reply, a certain amount of credit is restored. One advantage of the credit scheme is that if the credit limit is not reached, the delay in credit return does not affect performance.

[0187] In one embodiment, the four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Storage space transactions include one or more read requests and write requests to transfer data to or from a memory-mapped location. In one embodiment, memory space transactions can use two different address formats, e.g., a short address format such as a 32-bit address, or a long address format such as a 64-bit address. Configuration space transactions are used to access the configuration space of a PCIe device. Transactions in the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.

[0188] Thus, in one embodiment, the transaction layer 1305 assembles the packet header / payload 1306. The format of the current packet header / payload can be found in the PCIe specification on the PCIe specification website.

[0189] Figure 14 An embodiment of a PCIe transaction descriptor in accordance with one or more examples of this specification is shown. In one embodiment, the transaction descriptor 1400 is a mechanism for carrying transaction information. In this regard, the transaction descriptor 1400 supports the identification of transactions in the system. Other potential uses include tracking modifications to the default transaction ordering and the association of transactions with channels.

[0190] The transaction descriptor 1400 includes a global identifier field 1402, an attribute field 1404, and a channel identifier field 1406. In the example shown, the global identifier field 1402 is depicted as including a local transaction identifier field 1408 and a source identifier field 1410. In one embodiment, the global transaction identifier 1402 is unique for all outstanding requests.

[0191] According to one implementation, the local transaction identifier field 1408 is a field generated by the requesting agent and is unique for all outstanding requests that need to be completed by that requesting agent. Additionally, in this example, the source identifier 1410 uniquely identifies the requesting agent within the PCIe hierarchy. Thus, together with the source ID 1410, the local transaction identifier 1408 field provides a global identification of transactions within the hierarchical domain.

[0192] The attribute field 1404 specifies the characteristics and relationships of the transaction. In this regard, the attribute field 1404 can potentially be used to provide additional information that allows modification of the default handling of the transaction. In one embodiment, the attribute field 1404 includes a priority field 1412, a reserved field 1414, a sorting field 1416, and a non-snoop field 1418. Here, the priority sub-field 1412 can be modified by the initiator to assign a priority to the transaction. The reserved attribute field 1414 is reserved for future use or vendor-defined use. The reserved attribute field can be used to implement possible usage models using priorities or security attributes.

[0193] In this example, the sorting attribute field 1416 is used to provide optional information that conveys the sorting type that can modify the default sorting rule. According to one example implementation, the sorting attribute "0" indicates that the default sorting rule is to be applied, where the sorting attribute "1" indicates relaxed sorting, writes can be passed in the same direction, and read completions can be passed in the same direction. The snoop attribute field 1418 is used to determine whether the transaction has been snooped. As shown, the channel ID field 1406 identifies the channel associated with the transaction.

[0194] Link layer

[0195] The link layer 1310 (also referred to as the data link layer 1310) acts as an intermediate level between the transaction layer 1305 and the physical layer 1320. In one embodiment, the responsibility of the data link layer 1310 is to provide a reliable mechanism for exchanging TLPs between two linked components. One side of the data link layer 1310 accepts the TLP assembled by the transaction layer 1305, applies a packet sequence identifier 1311, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., CRC 1312, and submits the modified TLP to the physical layer 1320 for transmission across the physical layer to an external device.

[0196] Physical layer

[0197] In one embodiment, the physical layer 1320 includes a logic sub-block 1321 and an electronics sub-block 1322 to physically send the packet to an external device. Here, the logic sub-block 1321 is responsible for the "digital" functions of the physical layer 1321. In this regard, the logic sub-block includes a transmit portion for preparing outgoing information for transmission by the physical sub-block 1322, and a receive portion for identifying and preparing the received information before passing the received information to the link layer 1310.

[0198] Physical block 1322 includes a transmitter and a receiver. The transmitter is provided with symbols by logic sub-block 1321. The transmitter serializes the symbols and transmits them to an external device. The receiver is provided with the serialized symbols from the external device and converts the received signal into a bit stream. The bit stream is deserialized and provided to logic sub-block 1321. In one embodiment, an 8b / 10b transmission code is employed, where 10-bit symbols are transmitted / received. Here, special symbols are used to frame packets with frame 1323. Additionally, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.

[0199] As described above, although the transaction layer 1305, the link layer 1310, and the physical layer 1320 are discussed with reference to specific embodiments of the PCIe protocol stack, the layered protocol stack is not limited thereto. In fact, any layered protocol can be included / implemented. For example, the port / interface represented as a layered protocol includes: (1) a first layer for assembling packets, i.e., the transaction layer; a second layer for sorting packets, i.e., the link layer; and a third layer for transmitting packets, i.e., the physical layer. As a specific example, the Common Standard Interface (CSI) layered protocol is used.

[0200] Figure 15 An embodiment of a PCIe serial point-to-point structure according to one or more examples of the present specification is shown. Although an embodiment of a PCIe serial point-to-point link is shown, the serial point-to-point link is not limited thereto, as it includes any transmission path for transmitting serial data. In the shown embodiment, the basic PCIe link includes two low-voltage differential drive signal pairs: a transmit pair 1506 / 1511 and a receive pair 1512 / 1507. Thus, device 1505 includes transmit logic 1506 for sending data to device 1510 and receive logic 1507 for receiving data from device 1510. In other words, in the PCIe link, there are two transmit paths, i.e., paths 1516 and 1517, and two receive paths, i.e., paths 1518 and 1519.

[0201] A transmit path refers to any path for transmitting data, such as a transmission line, a copper wire, an optical fiber, a wireless communication channel, an infrared communication link, or other communication paths. The connection between two devices (e.g., device 1505 and device 1510) is referred to as a link, such as link 1515. A link can support one lane - each lane represents a set of differential signal pairs (one pair for transmission and one pair for reception). To expand the bandwidth, a link can aggregate multiple lanes represented by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.

[0202] A differential pair refers to two transmission paths, such as lines 1516 and 1517, for transmitting differential signals. As an example, when line 1516 switches from a low voltage level to a high voltage level, i.e., a rising edge, line 1517 is driven from a high logic level to a low logic level, i.e., a falling edge. Differential signals potentially exhibit better electrical characteristics, such as better signal integrity, i.e., cross-coupling, voltage overshoot / undershoot, ringing, etc. This allows for a better timing window, which enables a faster transmission frequency.

[0203] The features of one or more embodiments of the subject matter disclosed herein have been outlined above. These embodiments are provided to enable a person of ordinary skill in the art (PHOSITA) to better understand various aspects of the present disclosure. Certain terms that are readily understandable and the underlying technology and / or standards may be referred to without detailed description. It is expected that the PHOSITA will have or acquire the background knowledge or information in those technologies and standards sufficient to implement the teachings of this specification.

[0204] The PHOSITA will understand that they can readily use the present disclosure as a basis for designing or modifying other processes, structures, or variations to achieve the same purposes and / or realize the same advantages of the embodiments described herein. The PHOSITA will also recognize that such equivalent constructs do not depart from the spirit and scope of the present disclosure, and that they can make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.

[0205] In the foregoing description, certain aspects of some or all of the embodiments have been described in more detail than is strictly necessary to implement the appended claims. These details are provided for purposes of non-limiting examples only, to provide context and illustration of the disclosed embodiments. It should not be understood that these details are required, and the claims should not be "construed" as being limited. The phraseology may refer to "an embodiment" or "embodiments". Such phraseology and any other reference to embodiments should be understood broadly as referring to any combination of one or more embodiments. Additionally, several features disclosed in a particular "embodiment" may also be distributed among multiple embodiments. For example, if features 1 and 2 are disclosed in an "embodiment", embodiment A may have feature 1 but lack feature 2, while embodiment B may have feature 2 but lack feature 1.

[0206] This specification may provide descriptions in block diagram format, where certain features are disclosed in separate blocks. These should be broadly understood as disclosing how the various features interoperate, but do not imply that these features must necessarily be embodied in separate hardware or software. Additionally, in cases where more than one feature in the same block is disclosed, those features do not necessarily have to be embodied in the same hardware and / or software. For example, a computer "memory" can, in some cases, be distributed or mapped among multiple levels of cache or local memory, main memory, battery-backed volatile memory, and various forms of persistent memory (such as hard disks, storage servers, optical memory, magnetic disks, tape drives, or similar devices). In some embodiments, some components may be omitted or combined. In a general sense, the arrangements depicted in the figures may be more logical in their representation, while the physical architecture may include various arrangements, combinations, and / or mixtures of these elements. Countless possible design configurations can be used to achieve the operational objectives outlined herein. Thus, the associated infrastructure has countless alternative arrangements, design choices, device possibilities, hardware configurations, software implementations, and device options.

[0207] Reference may be made here to computer-readable media, which may be tangible and non-transitory computer-readable media. As used throughout this specification and the claims, "computer-readable media" should be understood to include one or more computer-readable media of the same or different types. By way of non-limiting example, computer-readable media may include optical drives (e.g., CD / DVD / Blue-ray), hard disk drives, solid state drives, flash memory, or other non-volatile media. Computer-readable media may also include such media as read-only memory (ROM), an FPGA or ASIC configured to execute desired instructions, stored instructions for programming an FPGA or ASIC to execute desired instructions, intellectual property (IP) blocks that can be integrated into other circuits in hardware, or instructions directly encoded into hardware or microcode on a processor such as a microprocessor, digital signal processor (DSP), microcontroller, or any other suitable component, device, element, or object as appropriate to specific requirements. The non-transitory storage media herein are expressly intended to include any non-transitory dedicated or programmable hardware configured to provide the disclosed operations or cause a processor to perform the disclosed operations.

[0208] In this specification and the claims, various elements may be "communicatively", "electrically", "mechanically", or otherwise "coupled" to each other. Such coupling may be direct point-to-point coupling, or may include intermediate devices. For example, two devices may be communicatively coupled to each other via a controller facilitating communication. Devices may be electrically coupled to each other via intermediate devices such as signal boosters, voltage dividers, or buffers. Mechanically coupled devices may be mechanically coupled indirectly.

[0209] Any "module" or "engine" disclosed herein may refer to or include software, a software stack, hardware, firmware, and / or a combination of software, circuitry configured to perform the functions of the engine or module, or any computer-readable medium as described above. Where appropriate, these modules or engines may be provided on or in combination with a hardware platform, which may include hardware computing resources such as processors, memory, storage, interconnects, networks and network interfaces, accelerators, or other suitable hardware. Such a hardware platform may be provided as a single monolithic device (e.g., in a PC form factor), or some or part of the functionality may be distributed (e.g., a "composite node" in a high-end data center, where computing, memory, storage, and other resources may be dynamically allocated and need not be local to each other).

[0210] Flowcharts, signal flow diagrams, or other illustrations showing operations performed in a particular order may be disclosed herein. Unless otherwise explicitly stated, or unless required in a particular context, this order should be understood as being merely a non-limiting example. Additionally, in cases where one operation is shown following another operation, other intervening operations may occur, which may or may not be relevant. Some operations may also be performed simultaneously or in parallel. In cases where an operation is said to be "based on" or "in accordance with" another item or operation, this should be understood to imply that the operation is at least partially based on or at least partially in accordance with the other item or operation. This should not be construed as implying that the operation is based solely or entirely on the item or operation, or in accordance solely or entirely with the item or operation.

[0211] All or part of any hardware element disclosed herein may be readily provided in a system-on-chip (SoC), including a central processing unit (CPU) package. An SoC represents an integrated circuit (IC) that integrates the components of a computer or other electronic system onto a single chip. Thus, for example, a client device or a server device may be provided in whole or in part in an SoC. An SoC may contain digital, analog, mixed-signal, and radio-frequency functionality, all of which may be provided on a single chip substrate. Other embodiments may include a multi-chip module (MCM), where multiple chips are located within a single electronic package and are configured to interact closely with each other through the electronic package.

[0212] In a general sense, any appropriately configured circuit or processor can execute any type of instructions associated with data to implement the operations detailed herein. Any processor disclosed herein can transform an element or article (e.g., data) from one state or thing to another. Additionally, based on specific requirements and implementations, the information that is tracked, sent, received, or stored in a processor can be provided in any database, register, table, cache, queue, control list, or storage structure, all of which can be referenced within any suitable time frame. Any memory or storage element disclosed herein should be considered to be appropriately included within the broad terms "memory" and "storage".

[0213] The computer program logic for implementing all or part of the functionality described herein is embodied in various forms, including but not limited to source code form, computer-executable form, machine instructions or microcode, programmable hardware, and various intermediate forms (e.g., forms generated by assemblers, compilers, linkers, or locators). In one example, the source code includes a series of computer program instructions implemented in various programming languages, such as object code, assembly language, or high-level languages such as OpenCL, FORTRAN, C, C++, JAVA, or HTML, for various operating systems or operating environments, or implemented in hardware description languages such as Spice, Verilog, and VHDL. The source code can define and use various data structures and communication messages. The source code can be in computer-executable form (e.g., via an interpreter), or the source code can be transformed (e.g., via a translator, assembler, or compiler) into computer-executable form, or into an intermediate form, such as bytecode. In appropriate cases, any of the foregoing can be used to construct or describe appropriate discrete or integrated circuits, whether sequential, combinational, state machine, or otherwise.

[0214] In one example embodiment, any number of the circuits of the figures can be implemented on a board of a related electronic device. The board can be a general-purpose circuit board that can hold various components of the internal electronic system of the electronic device and also provide connectors for other peripheral devices. Based on specific configuration requirements, processing requirements, and computing designs, any suitable processor and memory can be appropriately coupled to the board. Note that, using the numerous examples provided herein, interactions can be described in terms of two, three, four, or more electronic components. However, this is done only for purposes of clarity and example. It should be realized that the system can be combined or reconfigured in any suitable way. Along similar design alternatives, any of the shown components, modules, and elements in the figures can be combined in various possible configurations, all of which are within the broad scope of this specification.

[0215] Numerous other changes, substitutions, variations, alterations and modifications can be determined by those skilled in the art, and this disclosure is intended to cover all such changes, substitutions, variations, alterations and modifications that fall within the scope of the appended claims. To assist the United States Patent and Trademark Office (USPTO) and any readers of any patent issued in this application in interpreting the appended claims, the applicant wishes to note that the applicant: (a) does not intend for any of the appended claims, as they exist on their filing date, to invoke 35 USC section 112, paragraph 6 (pre-AIA) or paragraph (f) (post-AIA), unless the phrase "means for... " or "step for... " is specifically used in a particular claim; (b) does not intend to limit this disclosure in any way not expressly reflected in the appended claims by any statement in the specification.

[0216] Example implementation

[0217] In one example, a structure controller for providing a coherent accelerator structure is disclosed, including: a host interconnect for communicatively coupling to a host device; a memory interconnect communicatively coupled to an accelerator memory; an accelerator interconnect for communicatively coupling to an accelerator having a last-level cache (LLC); and an LLC controller configured to provide a bias check for memory access operations.

[0218] A structure controller is also disclosed, further including a fabric coherence engine (FCE) configured to be able to map the accelerator memory to a host fabric memory address space, wherein the structure controller is configured to direct host memory access operations to the accelerator memory via the FCE.

[0219] A structure controller is also disclosed, wherein the FCE is physically separated from the LLC controller.

[0220] A structure controller is also disclosed, further including a direct bypass bus for connecting the LLC to the memory interconnect and bypassing the FCE.

[0221] A structure controller is also disclosed, wherein the structure controller is configured to provide the structure in a plurality of n independent slices.

[0222] A structure controller is also disclosed, wherein n = 8.

[0223] A structure controller is also disclosed, wherein the n independent slices include n independent LLC controllers interconnected by a horizontal interconnect and communicatively coupled to respective memory controllers via respective vertical interconnects.

[0224] Also disclosed is a fabric controller, further comprising: a power manager configured to determine that the LLC controller is idle, and power down the horizontal interconnect and keep the corresponding vertical interconnect and host interconnect active.

[0225] Also disclosed is a fabric controller, wherein the LLC is a level-3 cache.

[0226] Also disclosed is a fabric controller, wherein the host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

[0227] Also disclosed is a fabric controller, wherein the host interconnect is a PCIe interconnect.

[0228] Also disclosed is a fabric controller, wherein the fabric controller is an integrated circuit.

[0229] Also disclosed is a fabric controller, wherein the fabric controller is an intellectual property (IP) block.

[0230] Also disclosed is an accelerator device, comprising: an accelerator including a last-level cache (LLC); and a fabric controller for providing a coherent accelerator fabric, including: a host interconnect for communicatively coupling the accelerator to a host device; a memory interconnect for communicatively coupling the accelerator and the host device to an accelerator memory; an accelerator interconnect for communicatively coupling the accelerator fabric to the LLC; and an LLC controller configured to provide a bias check for memory access operations.

[0231] Also disclosed is an accelerator device, wherein the fabric controller further includes a fabric coherence engine (FCE) configured to enable the accelerator memory to be mapped to a host fabric memory address space, and wherein the fabric controller is configured to direct host memory access operations to the accelerator memory via the FCE.

[0232] Also disclosed is an accelerator device, wherein the FCE is physically separated from the LLC controller.

[0233] Also disclosed is an accelerator device, wherein the fabric controller further includes a direct bypass bus for connecting the LLC to the memory interconnect and bypassing the FCE.

[0234] Also disclosed is an accelerator device, wherein the fabric controller is configured to provide the fabric in a plurality of n independent slices.

[0235] Also disclosed is an accelerator device, wherein n = 8.

[0236] Also disclosed is an accelerator device, wherein the n independent slices include n independent LLC controllers interconnected by a horizontal interconnect and communicatively coupled to corresponding memory controllers via corresponding vertical interconnects.

[0237] Further disclosed is an accelerator device, further comprising: a power manager configured to determine that the LLC controller is idle, and power down the horizontal interconnect and keep the corresponding vertical interconnect and host interconnect active.

[0238] Also disclosed is an accelerator device, wherein the LLC is a level 3 cache.

[0239] Also disclosed is an accelerator device, wherein the host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

[0240] Also disclosed is an accelerator device, wherein the host interconnect is a PCIe interconnect.

[0241] Also disclosed is one or more tangible non-transitory computer-readable media having stored thereon instructions for providing a fabric controller, including instructions for: providing a host interconnect to communicatively couple to a host device; providing a memory interconnect to communicatively couple to an accelerator memory; providing an accelerator interconnect to communicatively couple to an accelerator having a last level cache (LLC); and providing an LLC controller configured to provide a bias check for memory access operations.

[0242] Further disclosed is one or more tangible non-transitory computer-readable media, wherein the instructions further provide a fabric coherence engine (FCE) configured to be able to map the accelerator memory to the host fabric memory address space, wherein the fabric controller is configured to direct host memory access operations to the accelerator memory via the FCE.

[0243] Further disclosed is one or more tangible non-transitory computer-readable media, wherein the FCE is physically separated from the LLC controller.

[0244] Also disclosed is one or more tangible non-transitory computer-readable media, further comprising a direct bypass bus for connecting the LLC to the memory interconnect and bypassing the FCE.

[0245] Also disclosed is one or more tangible non-transitory computer-readable media, wherein the fabric controller is configured to provide the fabric in a plurality of n independent slices.

[0246] Also disclosed is one or more tangible non-transitory computer-readable media, wherein n = 8.

[0247] Also disclosed is one or more tangible non-transitory computer-readable media, wherein the n independent slices include n independent LLC controllers interconnected by a horizontal interconnect and communicatively coupled to corresponding memory controllers via corresponding vertical interconnects.

[0248] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions further provide a power manager configured to determine that the LLC controller is idle and power down the horizontal interconnect and keep the corresponding vertical interconnect and host interconnect active.

[0249] One or more tangible non-transitory computer-readable media are also disclosed, wherein the LLC is a level 3 cache.

[0250] One or more tangible non-transitory computer-readable media are also disclosed, wherein the host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

[0251] One or more tangible non-transitory computer-readable media are also disclosed, wherein the host interconnect is a PCIe interconnect.

[0252] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions include hardware instructions.

[0253] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions include field programmable gate array (FPGA) instructions.

[0254] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions include data for programming a field programmable gate array (FPGA).

[0255] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions include instructions for manufacturing a hardware device.

[0256] One or more tangible non-transitory computer-readable media are also disclosed, wherein the instructions include instructions for manufacturing an intellectual property (IP) block.

[0257] A method of providing a coherent accelerator fabric is also disclosed, including: communicatively coupling to a host device; communicatively coupling to an accelerator memory; communicatively coupling to an accelerator having a last level cache (LLC); and providing a bias check for memory access operations within the LLC controller.

[0258] A method is also disclosed, further including providing a fabric coherence engine (FCE) configured to be able to map the accelerator memory to the host fabric memory address space, wherein the fabric controller is configured to direct host memory access operations to the accelerator memory via the FCE.

[0259] A method is also disclosed, wherein the FCE is physically separated from the LLC controller.

[0260] Also disclosed is a method, further comprising providing a direct bypass path for connecting the LLC to the memory interconnect and bypassing the FCE.

[0261] Also disclosed is a method, further comprising providing the structure in a plurality of n independent slices.

[0262] Also disclosed is a method, wherein n = 8.

[0263] There is also a method, wherein the n independent slices include n independent LLC controllers interconnected by a horizontal interconnect and communicatively coupled to respective memory controllers via respective vertical interconnects.

[0264] Also disclosed is a method, further comprising: determining that the LLC controller is idle, and powering down the horizontal interconnect and maintaining the respective vertical interconnect and host interconnect in an active state.

[0265] Also disclosed is a method, wherein the LLC is a level-3 cache.

[0266] Also disclosed is a method, wherein the host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

[0267] Also disclosed is a method, wherein the host interconnect is a PCIe interconnect.

[0268] Also disclosed is an apparatus, comprising units for performing the method of any one of the above examples.

[0269] Also disclosed is an apparatus, wherein the unit comprises a structure controller.

[0270] Also disclosed is an accelerator device, comprising an accelerator, accelerator memory, and a structure controller.

[0271] Also disclosed is one or more tangible non-transitory computer-readable media having stored thereon instructions for providing a method or manufacturing the device or apparatus of any of the above examples.

Claims

1. A structure controller for providing a coherent accelerator structure, comprising: A host interconnect communicatively coupled to a host device; A memory interconnect communicatively coupled to an accelerator memory; An accelerator interconnect communicatively coupled to an accelerator having a last-level cache (LLC); And An LLC controller configured to provide a bias check for memory access operations; Wherein at least one of the host interconnect, the memory interconnect, or the accelerator interconnect supports transmitting data on a common physical channel according to multiple different protocols, and the multiple different protocols include an I / O protocol, a cache protocol, and a memory protocol.

2. The structure controller according to claim 1 further comprises: A fabric coherence engine (FCE) configured to enable mapping the accelerator memory to a host fabric memory address space, wherein the structure controller is configured to direct host memory access operations to the accelerator memory via the fabric coherence engine (FCE).

3. The structure controller according to claim 2, wherein The fabric coherence engine (FCE) is physically separated from the LLC controller.

4. The structure controller according to claim 3, further comprising a direct bypass bus for connecting the LLC to the memory interconnect and bypassing the fabric coherence engine (FCE).

5. The structure controller according to claim 1, wherein, The structure controller is configured to provide the structure in a plurality of n independent slices.

6. The structure controller according to claim 5, wherein, n=8。 7. The structure controller according to claim 5, wherein, The n independent slices include n independent LLC controllers interconnected via a horizontal interconnect and communicatively coupled to respective memory controllers via respective vertical interconnects.

8. The structure controller according to claim 7 further comprises: A power manager configured to determine that the LLC controller is idle and power down the horizontal interconnect and keep the respective vertical interconnects and the host interconnect active.

9. The structure controller according to claim 1, wherein the LLC is a level 3 cache.

10. The structure controller according to any one of claims 1 to 9, wherein, The host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

11. The structural controller according to any one of claims 1 to 9, wherein, The host interconnect is a PCIe interconnect.

12. The structure controller according to any one of claims 1 to 9, wherein, The structure controller is an integrated circuit.

13. The structure controller according to any one of claims 1 to 9, wherein, The structure controller is an intellectual property (IP) block.

14. An accelerator device, comprising: An accelerator including a last-level cache (LLC); And A structure controller for providing a coherent accelerator structure, the structure controller comprising: A host interconnect communicatively coupling the accelerator to a host device; A memory interconnect communicatively coupling the accelerator and the host device to an accelerator memory; An accelerator interconnect communicatively coupling the accelerator structure to the LLC; and An LLC controller configured to provide a bias check for memory access operations; Wherein at least one of the host interconnect, the memory interconnect, or the accelerator interconnect multiplexes data of multiple different protocols on a common physical channel, and the multiple different protocols include an I / O protocol, a cache protocol, and a memory protocol.

15. The accelerator device according to claim 14, wherein, The structure controller further includes a Fabric Coherence Engine (FCE) configured to enable mapping of the accelerator memory to the host fabric memory address space, wherein the structure controller is configured to direct host memory access operations to the accelerator memory via the Fabric Coherence Engine (FCE).

16. The accelerator device according to claim 15, wherein, The Fabric Coherence Engine (FCE) is physically separated from the LLC controller.

17. The accelerator device according to claim 16, wherein, The structure controller further includes a direct bypass bus for connecting the LLC to the memory interconnect and bypassing the Fabric Coherence Engine (FCE).

18. The accelerator device according to claim 14, wherein, The structure controller is configured to provide the structure in a plurality of n independent slices.

19. The accelerator device according to claim 18, wherein, n=8。 20. The accelerator device according to claim 18, wherein, The n independent slices include n independent LLC controllers interconnected via a horizontal interconnect and communicatively coupled to respective memory controllers via respective vertical interconnects.

21. The accelerator device according to claim 20, further comprising: A power manager configured to determine that the LLC controller is idle and power down the horizontal interconnect and keep the respective vertical interconnects and the host interconnect active.

22. The accelerator device according to claim 14, wherein, The LLC is a level-3 cache.

23. The accelerator device according to any one of claims 14 to 22, wherein, The host interconnect is an interconnect compliant with the Intel Accelerator Link (IAL).

24. The accelerator device according to any one of claims 14 to 22, wherein, The host interconnect is a PCIe interconnect.

25. A tangible non-transitory computer-readable medium having stored thereon instructions for providing a structure controller, including instructions for: Providing a host interconnect communicatively coupled to a host device; Providing a memory interconnect communicatively coupled to an accelerator memory; Providing an accelerator interconnect communicatively coupled to an accelerator having a last-level cache (LLC); and Providing an LLC controller configured to provide a bias check for memory access operations; Wherein at least one of the host interconnect, the memory interconnect, or the accelerator interconnect supports transmitting data according to a plurality of different protocols over a common physical channel, and the plurality of different protocols include an I / O protocol, a cache protocol, and a memory protocol.

26. The tangible non-transitory computer-readable medium as recited in claim 25, wherein the instructions further: provide a fabric coherent engine (FCE) configured to enable mapping of the accelerator memory to a host fabric memory address space, wherein, The structure controller is configured to direct host memory access operations to the accelerator memory via the Fabric Coherence Engine (FCE).

Citation Information

Patent Citations

  • Implementing vector memory operations

    US20070094477A1

  • Data processor

    US20100153656A1

  • Providing Hardware Support For Shared Virtual Memory Between Local And Remote Physical Memory

    US20110072234A1