Coherent memory devices via PCIe
By extending the consistent accelerator structure on the PCIe protocol, the consistency mapping between the accelerator memory and the host memory address space is achieved, which solves the problem of low accelerator memory management efficiency, improves access efficiency and performance, and reduces power consumption.
Patent Information
- Application Number
- CN201810995529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-08-29
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2038-08-29
AI Technical Summary
In the prior art, the memory management efficiency of the accelerator is inefficient, resulting in a cache coherence mechanism limiting the accelerator's ability to access local memory at high bandwidth and deployment options, and conventional systems cannot effectively manage the accelerator's local memory.
Using a consistent accelerator structure, the consistent mapping of accelerator memory and host memory address space is achieved through PCIe protocol expansion. It uses the combination of IOSF, IDI and SMI protocols to provide dynamic multiplexing, supports different types of interconnections to connect the accelerator and its cache to memory, and realizes consistent memory operation through the PCIe link.
It improves the memory access efficiency of the accelerator and the overall system performance, reduces power consumption, simplifies the entry burden of the device ecosystem, and supports the efficient deployment of multiple accelerator models.
Smart Images

Figure CN109582605B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to the field of interconnect devices and, more particularly, but not exclusively, to systems and methods for coherent memory devices over Peripheral Component Interconnect Express (PCIe). Background Art
[0002] A computing system includes various components for managing demand for processor resources. For example, a developer may include a hardware accelerator (or "accelerator") operably coupled to a central processing unit (CPU). Typically, an accelerator is an autonomous element configured to perform functions delegated to it by the CPU. An accelerator can be configured for specific functions and / or can be programmable. For example, an accelerator can be configured to perform specific calculations, graphics functions, and / or other functions. When the accelerator performs the specified function, the CPU is free to use resources for other needs. In conventional systems, an operating system (OS) can manage the physical memory available within the computing system (e.g., "system memory"); however, the OS does not manage or allocate memory local to the accelerator. As a result, memory protection mechanisms such as cache coherence introduce inefficiencies into accelerator-based configurations. For example, conventional cache coherence mechanisms limit the ability of an accelerator to access its attached local memory at very high bandwidth and / or limit deployment options for the accelerator. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings. It is emphasized that, in accordance with standard industry practice, various features are not necessarily drawn to scale and are shown for illustration purposes only. Where scale is shown, either explicitly or implicitly, it is merely provided as an illustrative example. In other embodiments, the dimensions of various features may be arbitrarily increased or decreased for clarity of discussion.
[0004] Figure 1 An example operating environment is shown that can be representative of various embodiments according to one or more examples of this specification.
[0005] Figure 2a An example of a fully consistent operating environment according to one or more examples of the present specification is shown.
[0006] Figure 2b An example of a non-uniform operating environment according to one or more examples of the present specification is shown.
[0007] Figure 2c An example of a coherence engine without a biased operating environment according to one or more examples of the present specification is shown.
[0008] Figure 3An example of an operating environment that can be representative of various embodiments according to one or more examples of this specification is shown.
[0009] Figure 4 Another example operating environment is shown that can be representative of various embodiments according to one or more examples of this specification.
[0010] Figure 5a and 5b Other example operating environments are shown that can represent various embodiments according to one or more examples of this specification.
[0011] Figure 6 An embodiment of a logic flow according to one or more examples of this specification is shown.
[0012] Figure 7 is a block diagram illustrating a structure according to one or more examples of this specification.
[0013] Figure 8 is a flowchart illustrating a method according to one or more examples of this specification.
[0014] Figure 9 is operated via PCIe according to one or more examples of this specification Block diagram of accelerator link memory (IAL.mem) read.
[0015] Figure 10 is a block diagram of an IAL.mem write over PCIe operation according to one or more examples of the present specification.
[0016] Figure 11 is a block diagram illustrating IAL.mem data completion via PCIe operations according to one or more examples of the present specification.
[0017] Figure 12 An embodiment of a structure composed of point-to-point links interconnecting a set of components according to one or more examples of the present specification is shown.
[0018] Figure 13 An embodiment of a layered protocol stack according to one or more embodiments of the present specification is shown.
[0019] Figure 14 An embodiment of a PCIe transaction descriptor according to one or more examples of the present specification is shown.
[0020] Figure 15 An embodiment of a PCIe serial point-to-point architecture according to one or more examples of the present specification is shown. DETAILED DESCRIPTION
[0021] This manual The Accelerator Link (IAL) is an extension of the Rosetta Link (R-Link) multi-chip package (MCP) interconnect link. IAL extends the R-Link protocol to enable support for accelerators and input / output (IO) devices that may not be adequately supported by the baseline R-Link or Peripheral Component Interconnect Express (PCIe) protocols.
[0022] The following disclosure provides many different embodiments or examples for realizing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present disclosure. Of course, these are merely examples and are not intended to be limiting. In addition, the present disclosure may repeat reference numerals and / or letters in various examples. This repetition is for simplicity and clarity purposes and does not in itself represent a relationship between the various embodiments and / or configurations being discussed. Different embodiments may have different advantages, and a particular advantage is not essential for any embodiment.
[0023] In the following description, numerous specific details are set forth, such as specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific measurements / heights, specific processor pipeline stages and examples of operations, etc., in order to provide a thorough understanding of the present invention. However, it will be apparent to one skilled in the art that these specific details are not required to practice the present invention.
[0024] In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / code for described algorithms, specific firmware code, specific interconnect operations, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific expression of algorithms in code, specific power-off and gating techniques / logic, and other specific operational details of computer systems have not been described in detail to avoid unnecessarily obscuring the present invention.
[0025] Although the following embodiments may be described with reference to energy conservation and energy efficiency in a particular integrated circuit, such as in a computing platform or microprocessor, other embodiments are also applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks. TM , and can also be used in other devices such as handheld devices, tablets, other thin notebooks, system-on-chip (SOC) devices and embedded applications.
[0026] Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and handheld personal computers (PCs). Embedded applications typically include microcontrollers, digital signal processors (DSPs), systems on chips (SoCs), network personal computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can perform the functions and operations taught below. In addition, the apparatus, methods, and systems described herein are not limited to physical computing devices, but may also involve software optimization for energy conservation and efficiency. As will become apparent in the following description, embodiments of the methods, apparatus, and systems described herein (whether with reference to hardware, firmware, software, or a combination thereof) are crucial for a "green technology" future balanced with performance considerations.
[0027] Various embodiments may generally relate to techniques for providing cache coherence between multiple components within a processing system. In some embodiments, the multiple components may include a processor, such as a central processing unit (CPU), and a logic device communicatively coupled to the processor. In various embodiments, the logic device may include locally attached memory. In some embodiments, the multiple components may include a processor communicatively coupled to an accelerator having locally attached memory (e.g., logic device memory).
[0028] In some embodiments, the processing system may operate a coherence biasing process that is configured to provide multiple cache coherence processes. In some embodiments, the multiple cache coherence processes may include a device biasing process and a host biasing process (collectively referred to as "bias protocol flows"). In some embodiments, the host biasing process may route requests to locally attached memory of a logical device, including requests from the logical device, through a coherence component of the processor. In some embodiments, the device biasing process may route logical device requests for logical device memory directly to the logical device memory, for example, without consulting the coherence component of the processor. In various embodiments, the cache coherence process may switch between the device biasing process and the host biasing process based on a bias indicator determined using application software, hardware hints, combinations thereof, and / or otherwise. The embodiments are not limited in this context.
[0029] The IAL described in this specification uses the Optimized Accelerator Protocol (OAP), which is a further extension of the R-Link MCP interconnect protocol. In one example, the IAL can be used to provide an interconnect fabric to an accelerator device (in some examples, the accelerator device can be a heavy accelerator that performs, for example, graphics processing, intensive computation, SmartNIC services, or similar processing. The accelerator can have its own attached accelerator memory, and an interconnect fabric such as the IAL or, in some embodiments, a PCIe-based fabric can be used to attach a processor to the accelerator. The interconnect fabric can be a coherent accelerator fabric, in which case the accelerator memory can be mapped to the memory address space of the host device. The coherent accelerator fabric can maintain consistency within the accelerator and between the accelerator and the host device. This can be used to achieve state-of-the-art memory and coherence support for these types of accelerators.
[0030] Advantageously, a coherent accelerator architecture according to the present specification can provide optimizations that increase efficiency and throughput. For example, an accelerator can have a number of n memory banks, where each of the corresponding n last-level caches (LLCs) is controlled by an LLC controller. The architecture can provide different types of interconnects to connect the accelerator and its caches to memory, and to connect the architecture to a host device.
[0031] As an illustration, throughout this specification, a bus or interconnect that connects devices of the same nature is referred to as a "horizontal" interconnect, while an interconnect or bus that connects different devices upstream and downstream may be referred to as a "vertical" interconnect. The terms "horizontal" and "vertical" are used herein for convenience only and are not intended to imply any necessary physical arrangement of the interconnects or buses, or to require that they be physically orthogonal to one another on the die.
[0032] For example, an accelerator may include 8 memory banks with corresponding 8 LLCs, which may be level 3 (L3) caches, each controlled by an LLC controller. The coherent accelerator structure may be divided into multiple independent "slices". Each slice serves a memory bank and its corresponding LLC and operates substantially independently of other slices. In an example, each slice may utilize biased operations provided by the IAL and provide parallel paths to the memory banks. Memory operations involving the host device may be routed through a fabric consistency engine (FCE), which provides consistency with the host device. However, the LLC of any individual slice may also have a parallel bypass path that writes directly to the memory, connecting the LLC directly to the memory bank, bypassing the FCE. For example, this may be achieved by providing bias logic (e.g., host bias or accelerator bias) in the LLC controller itself. The LLC controller may be physically separated from the FCE and may be upstream of the FCE in a vertical orientation, enabling accelerator biased memory operations to bypass the FCE and write directly to the memory bank.
[0033] Embodiments of the present specification can also achieve significant power savings by providing a power manager that selectively shuts down portions of the coherence structure when they are not in use. For example, an accelerator may be a very high bandwidth accelerator that can perform many operations per second. When the accelerator is performing its acceleration function, it is making heavy use of the structure and requires extremely high bandwidth so that the calculated values can be flushed to memory as soon as they are calculated. However, once the calculations are complete, the host device may not be ready to consume the data. In this case, portions of the interconnect, such as the vertical bus from the FCE to the LLC controller, and the horizontal bus between the LLC controller and the LLC itself, can be powered off. These can remain powered off until the accelerator receives new data to operate on.
[0034] The following table illustrates several categories of accelerators. Note that the baseline R-Link can only support the first two categories of accelerators, while IAL can support all five categories of accelerators.
[0035]
[0036] Note that in addition to producer-consumer, embodiments of these accelerators may require some degree of cache coherence to support usage models. Thus, the IAL is a coherent accelerator link.
[0037] IAL implements the accelerator model disclosed above using a combination of three protocols that are dynamically multiplexed onto a common link. These protocols include:
[0038] ● System-on-Chip Fabric (IOSF) - A reformatted PCIe-based interconnect that provides a non-uniform in-order semantics protocol. The IOSF can include an on-chip implementation of all or part of the PCIe standard. The IOSF packages PCIe traffic so it can be sent to a companion chip, such as a system-on-chip (SoC) or multi-chip module (MCM). The IOSF supports device discovery, device configuration, error reporting, interrupts, direct memory access (DMA)-style data transfers, and various services provided as part of the PCIe standard.
[0039] • Intra-die interconnect (IDI) - enables devices to issue coherent read and write requests to the processor.
[0040] • Scalable Memory Interconnect (SMI)—enables the processor to access memory attached to the accelerator.
[0041] These three protocols can be used in different combinations (eg, IOSF only, IOSF plus IDI, IOSF plus IDI plus SMI, IOSF plus SMI) to support the various models described in the table above.
[0042] As a baseline, IAL provides a single link or bus definition that can cover all five accelerator models through a combination of the aforementioned protocols. Note that producer-consumer accelerators are essentially PCIe accelerators. They only require the IOSF protocol, which is already a reformatted version of PCIe. IOSF supports some Accelerator Interface Architecture (AiA) operations, such as support for the enqueue (ENQ) instruction, which industry-standard PCIe devices may not support. Therefore, IOSF provides added value over PCIe for this type of accelerator. Producer-consumer plus accelerators are accelerators that can use only the IDI layer and IOSF layer of IAL.
[0043] In some embodiments, software-assisted device memory and autonomous device memory may require the SMI protocol over the IAL, including special operation codes (opcodes) over the SMI and special controller support for flows associated with those opcodes in the processor. These additions support the consistency bias model of the IAL. Usage can use all of the IOSF, IDI, and SMI.
[0044] The Big Cache usage also uses IOSF, IDI, and SMI, but may also add new qualifiers to the IDI and SMI protocols that are specifically designed for Big Cache accelerators (i.e., not used in the device memory model discussed above). Big Cache may add new special controller support in the processor that is not required for any other usage.
[0045] IAL refers to these three protocols as IAL.IO, IAL.cache, and IAL.mem. The combination of these three protocols provides the required performance benefits for the five accelerator models.
[0046] To achieve these benefits, IAL can use the R-Link (for MCP) or Flexbus (for discrete) physical layer to allow dynamic multiplexing of IO, cache and mem protocols.
[0047] However, some form factors do not natively support the R-Link or Flexbus physical layers. In particular, level 3 and level 4 device memory accelerators may not support R-Link or Flexbus. Existing examples of these may use standard PCIe, which restricts devices to a proprietary memory model rather than providing coherent memory that can be mapped into the host device's write-back memory address space. This model is limited because the device-attached memory cannot be directly addressed by software. This can result in suboptimal data marshaling between host and device memory over the bandwidth-limited PCIe link.
[0048] Thus, embodiments of the present specification provide consistency semantics that adhere to the same bias-based model defined by IAL, retaining the benefits of consistency without incurring the legacy overhead. All of this can be provided over existing PCIe physical links.
[0049] Thus, some of the benefits of IAL can be achieved at the physical layer, which does not provide the dynamic multiplexing between IO, cache, and mem protocols offered by R-Link and Flexbus. Advantageously, enabling the IAL protocol over PCIe for certain classes of devices reduces the burden of entry for the ecosystem of devices using physical PCIe links. It can leverage the existing PCIe infrastructure, including the use of off-the-shelf components such as switches, root ports, and endpoints. This also allows for easier cross-platform use of devices with attached memory using either conventional dedicated memory models or a consistent system-addressable memory model appropriate to the installation.
[0050] To support device types 3 and 4 (software assisted memory and autonomous device memory) as described above, the components of the IAL can be mapped as follows:
[0051] IOSF or IAL.io can use standard PCIe. This can be used for device discovery, enumeration, configuration, error reporting, interrupts, and DMA-style data transfers.
[0052] SMI or IAL.mem can use SMI tunneling on PCIe. The following describes the details of SMI tunneling on PCIe, including the following Figure 9 、 10 and the tunnel described in 11.
[0053] In some embodiments of this specification, IDI or IAL.cache are not supported. IDI enables a device to issue coherent read or write requests to host memory. Although IAL.cache may not be supported, the methods disclosed herein can be used to implement bias-based coherency for device-attached memory.
[0054] To achieve this result, the accelerator device can use one of its standard PCIe memory base address register (BAR) regions for the size of its attached memory. To do this, the device can implement a designated vendor-specific extension capability (DVSEC), similar to the standard IAL, to point to the BAR region that should be mapped into the consistent address space. In addition, the DVSEC can declare additional information such as memory type, latency, and other attributes that help the basic input / output system (BIOS) map this memory into the system address decoder in the consistent region. The BIOS can then program the memory base and limit the host physical address in the device.
[0055] This allows the host to read device-attached memory using standard PCIe Memory Read (MRd) opcodes.
[0056] However, for writes, non-post semantics may be required since access to metadata may be required upon completion. To achieve NP MWr on PCIe, the following reserved encoding can be used:
[0057] Fmt[2:0]-011b
[0058] Type[4:0]-11011b
[0059] Using the novel Non-Posted Memory Write (NP MWr) over PCIe has the added benefit of enabling the AiA ENQ instruction to efficiently submit work to the device.
[0060] To achieve optimal quality of service, embodiments of this specification may implement three different virtual channels (VC0, VC1, and VC2) to separate different service types, as follows:
[0061] VC0 → All memory-mapped input / output (MMIO) and configuration (CFG) traffic, both upstream and downstream
[0062] VC1 → IAL.mem write (from host to device)
[0063] VC2 → IAL.mem read (from host to device)
[0064] Note that because IAL.cache or IDI are not supported, embodiments of the present description may not allow the accelerator device to issue coherent reads or writes to host memory.
[0065] Embodiments of the present specification may also have the ability to flush cache lines from the host (required for host-to-device bias flipping). This can be accomplished at cache line granularity using non-allocate zero-length writes from the device over PCIe. Non-allocate semantics are described using transactions and transaction hints on transaction layer packets (TLPs).
[0066] TH=1, PH=01
[0067] This allows the host to invalidate a given row. The device can issue a read after the page bias flip to ensure that all rows are flushed. The device can also implement a content addressable memory (CAM) to ensure that no new requests for that row are received from the host while the flip is in progress.
[0068] Systems and methods for consistent memory devices over PCIe will now be described with more particular reference to the accompanying drawings. It should be noted that certain reference numerals may be repeated throughout the drawings to indicate that a particular device or block is completely or substantially consistent throughout the drawings. However, this is not intended to imply any particular relationship between the various disclosed embodiments. In some examples, a class of elements may be referred to by a particular reference numeral ("widget 10"), while individual species or examples of that class may be referred to by hyphenated numbers ("first particular widget 10-1" and "second particular widget 10-2").
[0069] Figure 1 An example operating environment 100 is shown that can be representative of various embodiments according to one or more examples of the present specification. Figure 1 The operating environment 100 depicted in FIG may include a device 105 having a processor 110, such as a central processing unit (CPU). The processor 110 may include any type of computing element, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a virtual processor, such as a virtual central processing unit (VCPU), or any other type of processor or processing circuit. In some embodiments, the processor 110 may be a processor available from Intel Corporation in Santa Clara, California. The company obtained One or more processors in a processor family. Although Figure 1Only one processor 110 is depicted in the drawings, but the apparatus may include multiple processors 110. Processor 110 may include processing elements 112, such as processing cores. In some embodiments, processor 110 may include a multi-core processor having multiple processing cores. In various embodiments, processor 110 may include processor memory 114, which may include, for example, a processor cache or local cache memory to facilitate efficient access to data processed by processor 110. In some embodiments, processor memory 114 may include random access memory (RAM); however, processor memory 114 may be implemented using other memory types, such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, and / or the like.
[0070] like Figure 1 As shown, processor 110 can be communicatively coupled to logic device 120 via link 115. In various embodiments, logic device 120 can include a hardware device. In various embodiments, logic device 120 can include an accelerator. In some embodiments, logic device 120 can include a hardware accelerator. In various embodiments, logic device 120 can include an accelerator implemented in hardware, software, or any combination thereof.
[0071] Although an accelerator may be used as an example logic device 120 in this specific embodiment, the embodiments are not limited in this regard, as the logic device 120 may include any type of device, processor (e.g., a graphics processing unit (GPU)), logic unit, circuit, integrated circuit, application specific integrated circuit (ASIC), field programmable gate array (FPGA), memory unit, computational unit, and / or the like capable of operating in accordance with some embodiments. In embodiments where the logic device 120 includes an accelerator, the logic device 120 may be configured to perform one or more functions of the processor 110. For example, the logic device 120 may include an accelerator operable to perform graphics functions (e.g., a GPU or graphics accelerator), floating point operations, fast Fourier transform (FFT) operations, and / or the like. In some embodiments, the logic device 120 may include an accelerator configured to operate using various hardware components, standards, protocols, and / or the like. Non-limiting examples of types of accelerators and / or accelerator technologies that can be used by the logic device may include OpenCAPI. TM , CCIX, GenZ, NVLink TM , Accelerator Interface Architecture (AiA), Cache Coherence Agent (CCA), Global Mapping and Coherent Device Memory (GCM), Graphics Media Accelerator (GMA) for directed input / output (IO) Virtualization technology (eg, VT-d, VT-x, and / or the like), shared virtual memory (SVM), and / or the like. The embodiments are not limited in this context.
[0072] Logic device 120 may include a processing element 122, such as a processing core. In some embodiments, logic device 120 may include multiple processing elements 122. Logic device 120 may include logic device memory 124, e.g., configured as locally attached memory for logic device 120. In some embodiments, logic device memory 124 may include local memory, cache memory, and / or the like. In various embodiments, logic device memory 124 may include random access memory (RAM); however, logic device memory 124 may be implemented using other memory types, such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, and / or the like. In some embodiments, at least a portion of logic device memory 124 may be visible or accessible to processor 110. In some embodiments, at least a portion of logic device memory 124 may be visible or accessible to processor 110 as system memory (e.g., as an accessible portion of system memory 130).
[0073] In various embodiments, the processor 110 may execute a driver 118. In some embodiments, the driver 118 may be used to control various functional aspects of the logic device 120 and / or manage communications with one or more applications that use the logic device 120 and / or computational results generated by the logic device 120. In various embodiments, the logic device 120 may include and / or have access to bias information 126. In some embodiments, the bias information 126 may include information associated with a coherence biasing process. For example, the bias information 126 may include information indicating which cache coherence process may be active for the logic device 120 and / or a particular process, application, thread, memory operation, and / or the like. In some embodiments, the bias information 126 may be read, written, or otherwise managed by the driver 118.
[0074] In some embodiments, the link 115 may include a bus component, such as a system bus. In various embodiments, the link 115 may include a communication link (e.g., a multi-protocol link) operable to support multiple communication protocols. The supported communication protocols may include standard load / store IO protocols for component communication, including serial link protocols, device cache protocols, memory protocols, memory semantic protocols, directory bit support protocols, networking protocols, consistency protocols, accelerator protocols, data storage protocols, point-to-point protocols, fabric-based protocols, on-package (or on-chip) protocols, fabric-based on-package protocols, and / or the like. Non-limiting examples of supported communication protocols may include a Peripheral Component Interconnect (PCI) protocol, a Peripheral Component Interconnect Express (PCIe or PCI-E) protocol, a Universal Serial Bus (USB) protocol, a Serial Peripheral Interface (SPI) protocol, a Serial AT Attachment (SATA) protocol, QuickPath Interconnect (QPI) protocol, UltraPath Interconnect (UPI) protocol, Optimized Accelerator Protocol (OAP), Accelerator Link (IAL), Intra-Device Interconnect (IDI) protocol (or IAL.cache), On-Chip Extensible Fabric (IOSF) protocol (or IAL.io), Scalable Memory Interconnect (SMI) protocol (or IAL.mem), SMI third generation (SMI3), and / or similar protocols. In some embodiments, link 115 may support intra-device protocols (e.g., IDI) and memory interconnect protocols (e.g., SMI3). In various embodiments, link 115 may support intra-device protocols (e.g., IDI), memory interconnect protocols (e.g., SMI3), and fabric-based protocols (e.g., IOSF).
[0075] In some embodiments, device 105 may include system memory 130. In various embodiments, system memory 130 may include main system memory for device 105. System memory 130 may store data and instruction sequences executed by processor 110 or any other device or component of device 105. In some embodiments, system memory 130 may be RAM; however, system memory 130 may be implemented using other memory types, such as dynamic DRAM, SDRAM, combinations thereof, and / or the like. In various embodiments, system memory 130 may store software applications 140 (e.g., "host software") that are executable by processor 110. In some embodiments, software applications 140 may use or otherwise be associated with logic device 120. For example, software applications 140 may be configured to use computational results generated by logic device 120.
[0076] The apparatus may include coherence logic 150 to provide a cache coherence process. In various embodiments, the coherence logic 150 may be implemented in hardware, software, or a combination thereof. In some embodiments, at least a portion of the coherence logic 150 may be disposed within the processor 110, partially disposed within the processor 110, or otherwise associated with the processor 110. For example, in some embodiments, coherence logic 150 for a cache coherence element or process 152 may be disposed within the processor 110. In some embodiments, the processor 110 may include a coherence controller 116 to perform various cache coherence processes, such as the cache coherence process 152. In some embodiments, the cache coherence process 152 may include one or more standard cache coherence techniques, functions, methods, processes, elements (including hardware or software elements), protocols, and / or the like, performed by the processor 110. Typically, the cache coherence process 152 may include a standard protocol for managing the system's cache so that no data is lost or no data is overwritten before transferring the data from the cache to the target memory. Non-limiting examples of standard protocols performed or supported by cache coherence process 152 may include a snoop-based (or snoop) protocol, a write-invalidate protocol, a write-update protocol, a directory-based protocol, a hardware-based protocol (e.g., a modified exclusive shared invalidate (MESI) protocol), a private memory-based protocol, and / or the like. In some embodiments, cache coherence process 152 may include one or more standard cache coherence protocols for maintaining cache coherence for logical device 120 having attached logical device memory 124. In some embodiments, cache coherence process 150 may be implemented in hardware, software, or a combination thereof.
[0077] In some embodiments, coherence logic 150 may include coherence biasing processes, such as a host biasing process or element 154 and a device biasing process or element 156. Generally, the coherence biasing process may operate to maintain cache coherence with respect to requests, data flows, and / or other memory operations associated with logical device memory 122. In some embodiments, at least a portion of the coherence logic, such as host biasing process 154, device biasing process 156, and / or bias selection component 158, may be disposed external to processor 110, such as in one or more separate coherence logic 150 units. In some embodiments, host biasing process 154, device biasing process 156, and / or bias selection component 158 may be implemented in hardware, software, or a combination thereof.
[0078] In some embodiments, host bias process 154 may include techniques, processes, data flows, data, algorithms, etc., for handling requests to logical device memory 124, including requests from logical device 120, through cache coherence process 152 of processor 110. In various embodiments, device bias process 156 may include techniques, processes, data flows, data, algorithms, etc., for allowing logical device 120 to directly access logical device memory 124, e.g., without using cache coherence process 152. In some embodiments, bias selection process 158 may include techniques, processes, data flows, data, algorithms, etc., for activating host bias process 154 or device bias process 156 as the active bias process for requests associated with logical device memory. In various embodiments, the active bias process may be based on bias information 126, which may include data, data structures, and / or processes used by the bias selection process to determine and / or set the active bias process.
[0079] Figure 2a An example of a fully consistent operating environment 200A is shown. Figure 2a The operating environment 200A depicted in FIG. 2 may include a device 202 having a CPU 210 including a plurality of processing cores 212a-n. Figure 2a As shown, the CPU may include various protocol agents, such as a cache agent 214, a home agent 216, a memory agent 218, and / or the like. Generally, the cache agent 214 may be operable to initiate transactions into coherent memory and retain copies in its own cache structure. The cache agent 214 may be defined by the messages it can sink and source based on behavior defined in the cache coherence protocol associated with the CPU. The cache agent 214 may also provide copies of coherent memory contents to other cache agents (e.g., accelerator cache agent 224). The home agent 216 may be responsible for the protocol side of memory interactions with the CPU 210, including both coherent and non-coherent home agent protocols. For example, the home agent 216 may command memory reads / writes. The home agent 216 may be configured to service coherent transactions, including handshaking with the cache agent when necessary. The home agent 216 may be operable to oversee a portion of the coherent memory of the CPU 210, for example, maintaining the coherence of a given address space. The home agent 216 may be responsible for managing conflicts that may arise between different cache agents. For example, home agent 216 can provide appropriate data and ownership responses as needed for the flow of a given transaction. Memory agent 218 can operate to manage access to memory. For example, memory agent 218 can facilitate memory operations (e.g., load / store operations) and functions (e.g., swap and / or the like) for CPU 210.
[0080] like Figure 2a As shown in FIG, the device 202 may include an accelerator 220 operably coupled to the CPU 210. The accelerator 220 may include an accelerator engine 222 operable to perform functions (e.g., computations and / or the like) offloaded from the CPU 210. The accelerator 220 may include an accelerator cache agent 224 and a memory agent 228.
[0081] The accelerator 220 and the CPU 210 may be configured according to and / or include various conventional hardware and / or memory access techniques. For example, Figure 2a As shown, all memory accesses, including those initiated by accelerator 220, must traverse path 230. Path 230 may include non-coherent links, such as PCIe links. In the configuration of device 202, accelerator engine 222 may be able to directly access accelerator cache agent 224 and memory agent 228, but not cache agent 214, home agent 216, or memory agent 218. Similarly, cores 212a-n will not be able to directly access memory agent 228. Therefore, the memory behind memory agent 228 will not be part of the system address map seen by cores 212a-n. Because cores 212a-n cannot access a common memory agent, data can only be exchanged through copies. In some implementations, a driver may be used to facilitate copying data back and forth between memory agents 218 and 228. For example, the driver may include runtime elements that create a shared memory abstraction that hides all copies from the programmer. In contrast, and as described in detail below, some embodiments may provide a configuration in which, when the accelerator engine wants to access accelerator memory, such as via the accelerator proxy 228, requests from the accelerator engine may be forced to traverse the link between the accelerator and the CPU.
[0082] Figure 2b An example of a non-conforming operating environment 200B is shown. Figure 2b The operating environment 200B depicted in FIG may include an accelerator 220 having an accelerator home agent 226. The CPU 210 and the accelerator 220 may be operably coupled via a non-coherent path 232 (eg, a UPI path or a CCIX path).
[0083] For operation of device 204, accelerator engine 222 and cores 212a-n can access memory agents 228 and 218. Cores 212a-n can access memory 218 without traversing link 232, and accelerator agent 222 can access memory 228 without traversing link 232. The cost of these local accesses from 222 to 228 is the need to construct a home agent 226 so that it can track the consistency of all accesses from cores 212a-n to memory 228. This requirement leads to complexity and high resource usage when device 204 includes multiple CPU 210 devices, all connected via other instances of link 232. Home agent 226 needs to be able to track the consistency of all cores 212a-n across all instances of CPU 210. This can become quite expensive in terms of performance, area, and power, especially for large configurations. Specifically, for accesses from CPU 210, it adversely affects the performance efficiency of accesses between accelerator 222 and memory 228, even if accesses from CPU 210 are expected to be relatively rare.
[0084] Figure 2c An example of a coherence engine without a biasing operating environment 200C is shown. As shown in FIG2 , device 206 may include an accelerator 220 operatively coupled to CPU 210 via coherence links 236 and 238. Accelerator 220 may include an accelerator engine 222 operable to perform functions (e.g., computations and / or the like) offloaded from CPU 210. Accelerator 220 may include an accelerator cache agent 224, an accelerator home agent 226, and a memory agent 228.
[0085] In the configuration of device 206, accelerator 220 and CPU 210 may be configured according to and / or include various conventional hardware and / or memory access technologies, such as CCIX, GCM, standard coherence protocols (e.g., symmetric coherence protocols), and / or the like. For example, as shown in FIG2 , all memory accesses, including those initiated by accelerator 220, must pass through path 230. In this manner, accelerator 220 must pass through CPU 220 (and, therefore, the coherence protocol associated with the CPU) to access accelerator memory (e.g., via memory agent 228). Consequently, the device may not provide the ability to access certain memory, such as accelerator-attached memory associated with accelerator 220, as part of system memory (e.g., as part of the system address map), which would allow host software to set operands and access computation results of accelerator 220 without the overhead of, for example, IO direct memory access (DMA) data copies. Compared to memory accesses, such data copies may require driver calls, interrupts, and MMIO accesses, which are all inefficient and complex. like Figure 2c As shown, the inability to access accelerator-attached memory without cache coherence overhead can be detrimental to the execution time of computations offloaded to the accelerator 220. For example, in a process involving a large number of streaming write-to-memory transactions, the cache coherence overhead can halve the effective write bandwidth seen by the accelerator 220.
[0086] The efficiency of operand setup, result access, and accelerator computations plays a role in determining the effectiveness and benefits of offloading CPU 210 work to accelerator 220. If the cost of offloading work is too high, offloading may not be beneficial or may be limited to very large tasks. Consequently, various developers have created accelerators that attempt to improve the efficiency of using accelerators (such as accelerator 220) with limited effectiveness compared to techniques configured according to some embodiments. For example, some conventional GPUs can operate without mapping accelerator-attached memory as part of the system address and may or may not use certain virtual memory configurations (e.g., SVM) to access accelerator-attached memory. Consequently, in such systems, accelerator-attached memory is invisible to host system software. Instead, accelerator-attached memory is accessed only through a runtime software layer provided by the GPU device driver. Data copies and page table manipulation systems are used to create the appearance of a system with virtual memory (e.g., SVM) enabled. Such systems are inefficient, particularly compared to some embodiments, because they require memory copying, memory pinning, memory replication, and complex software. These requirements result in significant overhead at memory page transition points that is not required in systems configured according to some embodiments. In certain other systems, conventional hardware coherency mechanisms are used for memory operations associated with accelerator-attached memory, which limits the accelerator's ability to access accelerator-attached memory at high bandwidth and / or limits deployment options for a given accelerator (e.g., accelerators attached via on-package or off-package links cannot be supported without significant bandwidth loss).
[0087] Typically, conventional systems can use one of two methods to access accelerator-attached memory: a full consistency (or full hardware consistency) method or a private memory model or method. The full consistency method requires that all memory accesses, including accelerator-requested accesses to accelerator-attached memory, must go through the consistency protocol of the corresponding CPU. In this way, the accelerator must take a circuitous route to access the accelerator-attached memory because the request must at least be transmitted to the corresponding CPU, through the CPU consistency protocol, and then transmitted to the accelerator-attached memory. Therefore, the full consistency method carries a consistency overhead when the accelerator accesses its own memory, which may significantly impair the date bandwidth that the accelerator can extract from its own attached memory. The private memory model requires a large amount of resource and time costs, such as memory copies, page pinning requirements, page copy data bandwidth costs and / or page translation costs (e.g., translation lookaside buffer (TLB) beats, page table operations, and / or the like). Thus, some embodiments may provide a coherency biasing process configured to provide multiple cache coherency processes that provide, among other things, better memory utilization and improved performance for systems including accelerator-attached memory compared to conventional systems.
[0088] Figure 3 An example of an operating environment 300 is shown that may be representative of various embodiments. Figure 3 The operating environment 300 depicted in FIG. 3 may include an apparatus 305 operable to provide a consistent biasing process according to some embodiments. In some embodiments, the apparatus 305 may include a CPU 310 having a plurality of processing cores 312a-n and various protocol agents, such as a cache agent 314, a home agent 316, a memory agent 318, and / or the like. The CPU 310 may be communicatively coupled to an accelerator 320 using various links 335, 340. The accelerator 320 may include an accelerator engine 312 and a memory agent 318, and may include or access bias information 338.
[0089] like Figure 3As shown, the accelerator engine 322 can be directly communicatively coupled to the memory agent 328 via a biased coherence bypass 330. In various embodiments, the accelerator 320 can be configured to operate in a device biased process, wherein the biased coherence bypass 330 can allow memory requests of the accelerator engine 322 to directly access the accelerator's accelerator-attached memory (not shown) facilitated by the memory agent 328. In various embodiments, the accelerator 320 can be configured to operate in a host biased process, wherein memory operations associated with the accelerator-attached memory can be processed using the CPU's cache coherence protocol via links 335, 340, for example, via the cache agent 314 and the home agent 316. Thus, the accelerator 320 of the device 305 can utilize the coherence protocol of the CPU 310 when appropriate (e.g., when a non-accelerator entity requests accelerator-attached memory), while allowing the accelerator 320 to directly access the accelerator-attached memory via the biased coherence bypass 330.
[0090] In some embodiments, the consistency bias (e.g., whether device bias or host bias is active) can be stored in bias information 338. In various embodiments, bias information 338 can include and / or can be stored in various data structures, such as a data table (e.g., a "bias table"). In some embodiments, bias information 338 can include a bias indicator having a value indicating the active bias (e.g., 0 = host bias, 1 = device bias). In some embodiments, bias information 338 and / or bias indicators can be at various levels of granularity, such as memory regions, page tables, address ranges, etc. For example, bias information 338 can specify that certain memory pages are set for device bias, while other memory pages are set for host bias. In some embodiments, bias information 338 can include a bias table that is configured to operate as a low-cost, scalable snoop filter.
[0091] Figure 4 An example operating environment 400 is shown that can represent various embodiments. According to some embodiments, Figure 4The operating environment 400 depicted in FIG may include an apparatus 405 operable to provide a coherent biasing process. Apparatus 405 may include an accelerator 410 communicatively coupled to a host processor 445 via a link (or multi-protocol link) 489. Accelerator 410 and host processor 445 may communicate via the link using interconnect structures 415 and 450, respectively, which allow data and messages to be passed between them. In some embodiments, link 489 may include a multi-protocol link operable to support multiple protocols. For example, link 489 and interconnect structures 415 and 450 may support various communication protocols, including but not limited to serial link protocols, device cache protocols, memory protocols, memory semantic protocols, directory bit support protocols, networking protocols, coherence protocols, accelerator protocols, data storage protocols, point-to-point protocols, fabric-based protocols, on-package (or on-chip) protocols, fabric-based on-package protocols, and / or the like. Non-limiting examples of supported communication protocols may include PCI, PCIe, USB, SPI, SATA, QPI, UPI, OAP, IAL, IDI, IOSF, SMI, SMI3, and / or the like. In some embodiments, link 489 and interconnect structures 415 and 450 may support intra-device protocols (e.g., IDI) and memory interconnect protocols (e.g., SMI3). In various embodiments, link 489 and interconnect structures 415 and 450 may support intra-device protocols (e.g., IDI), memory interconnect protocols (e.g., SMI3), and fabric-based protocols (e.g., IOSF).
[0092] In some embodiments, the accelerator 410 may include bus logic 435 having a device TLB 437. In some embodiments, the bus logic 435 may be or include PCIe logic. In various embodiments, the bus logic 435 may communicate over the interconnect 480 using a fabric-based protocol (e.g., IOSF) and / or a Peripheral Component Interconnect Express (PCIe or PCI-E) protocol. In various embodiments, communication over the interconnect 480 may be used for various functions, including, but not limited to, discovery, register access (e.g., registers (not shown) of the accelerator 410), configuration, initialization, interrupts, direct memory access, and / or address translation services (ATS).
[0093] The accelerator 410 may include a core 420 having a host memory cache 422 and an accelerator memory cache 424. The core 420 may communicate using an interconnect 481 using, for example, an intra-device protocol (e.g., IDI) for various functions such as coherence requests and memory streaming. In various embodiments, the accelerator 410 may include coherence logic 425 that includes or accesses bias mode information 427. The coherence logic 425 may communicate using an interconnect 482 using, for example, a memory interconnect protocol (e.g., SMI3). In some embodiments, communication over the interconnect 482 may be used for memory streaming. The accelerator 410 may be operably coupled to an accelerator memory 430 (e.g., as an accelerator-attached memory) that may store bias information 432.
[0094] In various embodiments, host processor 445 may be operably coupled to host memory 440 and may include coherence logic (or coherence and cache logic) 455 with a last level cache (LLC) 457. Coherence logic 455 may communicate using various interconnects, such as interconnects 484 and 485. In some embodiments, interconnects 484 and 485 may include a memory interconnect protocol (e.g., SMI3) and / or an intra-device protocol (e.g., IDI). In some embodiments, LLC 457 may include a combination of at least a portion of host memory 440 and accelerator memory 430.
[0095] Host processor 445 may include bus logic 460 having an input-output memory management unit (IOMMU) 462. In some embodiments, bus logic 460 may be or include PCIe logic. In various embodiments, bus logic 460 may communicate via interconnects 486 and 488 using a fabric-based protocol (e.g., IOSF) and / or a Peripheral Component Interconnect Express (PCIe or PCI-E) protocol. In various embodiments, host processor 445 may include multiple cores 465a-n, each with a cache 467a-n. In some embodiments, cores 465a-n may include Architecture (IA) core. Each of the cores 465a-n can communicate with the coherence logic 455 via interconnects 487a-n. In some embodiments, the interconnects 487a-n can support an intra-device protocol (e.g., IDI). In various embodiments, the host processor can include a device 470 operable to communicate with the bus logic 460 via interconnect 488. In some embodiments, the device 470 can include an IO device, such as a PCIe IO device.
[0096] In some embodiments, the apparatus 405 is operable to perform a consistent biasing process applicable to various configurations, such as a system having an accelerator 410 and a host processor 445 (e.g., a computer processing complex including one or more computer processor chips), wherein the accelerator 410 is communicatively coupled to the host processor 445 via a multi-protocol link 489, and wherein memory is directly attached to the accelerator 410 and the host processor 445 (e.g., accelerator memory 430 and host memory 440, respectively). The consistent biasing process provided by the apparatus 405 can provide a number of technical advantages over conventional systems, such as providing the accelerator 410 and "host" software running on the processing cores 465a-n with access to the accelerator memory 430. The consistent biasing process provided by the apparatus can include a host biasing process and a device biasing process (collectively, biasing protocol streams) and multiple options for modulating and / or selecting the biasing protocol stream for a particular memory access.
[0097] In some embodiments, a bias protocol flow may be implemented, at least in part, using a protocol layer (e.g., a "bias protocol layer") on the multi-protocol link 489. In some embodiments, the bias protocol layer may include: an intra-device protocol (e.g., IDI) and / or a memory interconnect protocol (e.g., SMI3). In some embodiments, the bias protocol flow may be enabled by using various information of the bias protocol layer, adding new information to the bias protocol layer, and / or adding support for the protocol. For example, the bias protocol flow may be implemented using existing opcodes for the intra-device protocol (e.g., IDI), adding opcodes to the memory interconnect protocol (e.g., SMI3) standard, and / or adding support for the memory interconnect protocol (e.g., SMI3) on the multi-protocol link 489 (e.g., a conventional multi-protocol link may only include an intra-device protocol (e.g., IDI) and a fabric-based protocol (for example, IOSF)).
[0098] In some embodiments, the device 405 can be associated with at least one operating system (OS). The OS can be configured not to use the accelerator memory 430 or not to use certain portions of the accelerator memory 430. Such an OS can include support for a "memory-only NUMA module" (e.g., no CPU). The device 405 can execute a driver (e.g., including driver 118) to perform various accelerator memory services. Illustrative and non-limiting accelerator memory services implemented in the driver can include driver discovery and / or acquisition / allocation of accelerator memory 430, providing allocation APIs and mapping pages through OS page mapping services, providing processes for managing multi-process memory oversubscription and work scheduling, providing APIs to allow software applications to set and change the bias mode of memory regions of the accelerator memory 430, and / or deallocation APIs that return pages to the driver's free page list and / or return pages to the default bias mode.
[0099] Figure 5a An example of an operating environment 500 is shown that can represent various embodiments. According to some embodiments, Figure 5a The operating environment 500 depicted in FIG. 5 can provide a host bias process flow. Figure 5a As shown, the device 505 may include a CPU 510 communicatively coupled to an accelerator 520 via a link 540. In some embodiments, the link 540 may include a multi-protocol link. The CPU 510 may include a coherence controller 530 and may be communicatively coupled to a host memory 512. In various embodiments, the coherence controller 530 may be operable to provide one or more standard cache coherence protocols. In some embodiments, the coherence controller 530 may include and / or be associated with various agents, such as a home agent. In some embodiments, the CPU 510 may include and / or be communicatively coupled to one or more IO devices. The accelerator 520 may be communicatively coupled to an accelerator memory 522.
[0100] Host-biased process flows 550 and 560 may include a set of data flows that funnel all requests, including requests from accelerator 520, to accelerator memory 522 via coherence controller 530 in CPU 510. In this manner, accelerator 520 takes a circuitous route to access accelerator memory 522, but allows accesses from accelerator 520 and CPU 510 (including requests from IO devices via CPU 510) to remain coherent using the standard cache coherence protocol of coherence controller 530. In some embodiments, host-biased process flows 550 and 560 may utilize an intra-device protocol (e.g., IDI). In some embodiments, host-biased process flows 550 and 560 may utilize standard opcodes of the intra-device protocol (e.g., IDI), for example, to issue requests to coherence controller 530 via multi-protocol link 540. In various embodiments, coherence controller 530 may issue various coherence messages (e.g., snoops) generated by requests from accelerator 520 to all peer processor chips and internal processor agents on behalf of accelerator 520. In some embodiments, the various consistency messages may include point-to-point protocol (eg, UPI) consistency messages and / or intra-device protocol (eg, IDI) messages.
[0101] In some embodiments, the coherence controller 530 may conditionally issue a memory access message to an accelerator memory controller (not shown) of the accelerator 520 over the multi-protocol link 540. Such a memory access message may be the same or substantially similar to a memory access message that the coherence controller 530 may send to a CPU memory controller (not shown), and may include a new opcode that allows data to be returned directly to an agent inside the accelerator 520, rather than forcing the data to be returned to the coherence controller and then returned again to the accelerator 520 over the multi-protocol link 540 as an intra-device protocol (e.g., IDI) response.
[0102] The host biased process flow 550 may include flows resulting from requests or memory operations to the accelerator memory 522 originating from the accelerator. The host biased process path 560 may include flows resulting from requests or memory operations to the accelerator memory 522 originating from the CPU 510 (or an IO device or a software application associated with the CPU 510). When the device 505 is active in the host biased mode, the host biased process flows 550 and 560 may be used to access the accelerator memory 522, as shown in FIG. Figure 5aAs shown. In various embodiments, in host bias mode, all requests from CPU 510 targeting accelerator memory 522 can be sent directly to coherence controller 530. Coherence controller 530 can apply a standard cache coherence protocol and send standard cache coherence messages. In some embodiments, coherence controller 530 can send a memory interconnect protocol (e.g., SMI3) command for such a request over multi-protocol link 540, where the memory interconnect protocol (e.g., SMI3) stream returns data across multi-protocol link 540.
[0103] Figure 5b Another example of an operating environment 500 that can represent various embodiments is shown. According to some embodiments, Figure 5a The operating environment 500 depicted in FIG5 can provide a device bias process flow. As shown in FIG5 , when the apparatus 505 is active in the device bias mode, a device bias path 570 can be used to access the accelerator memory 522. For example, the device bias flow or path 570 can allow the accelerator 520 to directly access the accelerator memory 522 without consulting the coherence controller 530. More specifically, the device bias path 570 can allow the accelerator 520 to directly access the accelerator memory 522 without having to send a request over the multi-protocol link 540.
[0104] In device bias mode, according to some embodiments, CPU 510 requests to accelerator memory may be issued in the same or substantially similar manner as described for host bias mode, but differ in the memory interconnect protocol (e.g., SMI3) portion of path 580. In some embodiments, in device bias mode, CPU 510 requests to attached memory may be completed as if they were issued as "uncached" requests. Typically, data for uncached requests during device bias mode is not cached in the CPU cache hierarchy. In this manner, the accelerator 520 is allowed to access data in accelerator memory 522 during device bias mode without consulting the coherence controller 530 of the CPU 510. In some embodiments, uncached requests may be implemented on the CPU 510's intra-device protocol (e.g., IDI) bus. In various embodiments, uncached requests may be implemented using a globally observed, once used (GO-UO) protocol on the CPU 510's intra-device protocol (e.g., IDI) bus. For example, a response to an uncached request may return a piece of data to CPU 510 and instruct CPU 510 to use the piece of data only once, e.g., to prevent caching of the piece of data and support use of an uncached data stream.
[0105] In some embodiments, the device 505 and / or the CPU 510 may not support GO-UO. In such embodiments, an uncached flow (e.g., path 580) may be implemented using a multi-message response sequence over a memory interconnect protocol (e.g., SMI3) of a multi-protocol link 540 and a CPU 510 intra-device protocol (e.g., IDI) bus. For example, when the CPU 510 targets the "device bias" page of the accelerator 520, the accelerator 520 may set one or more states to block future requests from the accelerator 520 for the target memory region (e.g., cache line) and send a "device bias hit" response over the memory interconnect protocol (e.g., SMI3) line of the multi-protocol link 540. In response to the "device bias hit" message, the coherence controller 530 (or its agent) may return the data to the requesting processor core, followed by a snoop invalidation message. When the corresponding processor core confirms that the snoop invalidation is complete, the coherence controller 530 (or its agent) may send a "device bias block complete" message to the accelerator 520 over the memory interconnect protocol (e.g., SMI3) line of the multi-protocol link 540. In response to receiving the "device bias block complete" message, the accelerator may clear the corresponding blocking state.
[0106] refer to Figure 4 , the bias mode information 427 may include a bias indicator that is configured to indicate an active bias mode (e.g., a device bias mode or a host bias mode). The selection of the active bias mode may be determined by the bias information 432. In some embodiments, the bias information 432 may include a bias table. In various embodiments, the bias table may include bias information 432 for certain areas of the accelerator memory, such as pages, lines, and / or the like. In some embodiments, the bias table may include a bit (e.g., 1 or 3 bits) for each accelerator memory 430 memory page. In some embodiments, the bias table may be implemented using RAM, such as SRAM at the accelerator 410 and / or a stolen range of the accelerator memory 430, with or without a cache inside the accelerator 410.
[0107] In some embodiments, the bias information 432 may include a bias table entry in a bias table. In various embodiments, the bias table entry associated with each access to the accelerator memory 430 may be accessed prior to the actual access to the accelerator memory 430. In some embodiments, a local request from the accelerator 410 that finds its page in the device bias may be forwarded directly to the accelerator memory 430. In various embodiments, a local request from the accelerator 410 that finds its page in the host bias may be forwarded to the host processor 445, for example, as an intra-device protocol (e.g., IDI) request over the multi-protocol link 489. In some embodiments, the host processor 445 request, for example, using a memory interconnect protocol (e.g., SMI3), that finds its page in the device bias may use an uncached stream (e.g., Figure 5b In some embodiments, the host processor 445 requests, for example, using a memory interconnect protocol (e.g., SMI3), finds that its page is in the host offset, and can complete the request as a standard memory read of the accelerator memory (e.g., via Figure 5a path 560).
[0108] The bias mode of the bias indicator of the bias mode information 427 for a region (e.g., a memory page) of the accelerator memory 430 can be changed by a software-based system, a hardware-assisted system, a hardware-based system, or a combination thereof. In some embodiments, the bias indicator can be changed via an application programming interface (API) call (e.g., OpenCL), which in turn can call an accelerator 410 device driver (e.g., driver 118). The accelerator 410 device driver can send a message to the accelerator 410 (or enqueue a command descriptor) instructing the accelerator 410 to change the bias indicator. In some embodiments, the change in the bias indicator can be accompanied by a cache flush operation in the host processor 445. In various embodiments, a cache flush operation may be required for a transition from host bias mode to device bias mode, but a cache flush operation may not be required for a transition from device bias mode to host bias mode. In various embodiments, software can change the bias mode of one or more memory regions of the accelerator memory 430 via a work request sent to the accelerator 430.
[0109] In some cases, the software may not be able or cannot easily determine when to make a bias conversion API call and identify the memory area that requires bias conversion. In this case, the accelerator 410 can provide a bias conversion prompt process, in which the accelerator 410 determines the need for a bias conversion and sends a message to the accelerator driver (e.g., driver 118) indicating that a bias conversion is required. In various embodiments, the bias conversion prompt process can be activated in response to a bias table lookup, which triggers the accelerator 410 to access the host bias mode memory area or the host processor 445 to access the device bias mode memory area. In some embodiments, the bias conversion prompt process can signal the need for bias conversion to the accelerator driver via an interrupt. In various embodiments, the bias table may include a bias status bit for enabling a bias conversion state value. The bias status bit can be used to allow access to a memory area during the bias change process (e.g., when partially flushing the cache and it is necessary to suppress incremental cache pollution caused by subsequent requests).
[0110] Included herein are one or more logical flows representing exemplary methods for performing novel aspects of the disclosed architecture. Although one or more methods shown herein are shown and described as a series of actions for purposes of simplicity of illustration, those skilled in the art will understand and appreciate that these methods are not limited by the order of the actions. Accordingly, some actions may occur in different orders and / or simultaneously with other actions shown and described herein. For example, those skilled in the art will understand and appreciate that a method may alternatively be represented as a series of interrelated states or events, such as in a state diagram. Furthermore, not all actions shown in a method are required for a novel implementation.
[0111] The logic flow may be implemented in software, firmware, hardware, or any combination thereof. In software and firmware embodiments, the logic flow may be implemented by computer-executable instructions stored on a non-transitory computer-readable medium or machine-readable medium (e.g., optical, magnetic, or semiconductor storage). The embodiments are not limited in this context.
[0112] Figure 6 An embodiment of a logic flow 600 is shown. Logic flow 600 may represent some or all operations performed by one or more embodiments described herein, such as apparatus 105, 305, 405, and 505. In some embodiments, logic flow 600 may represent some or all operations for a consistent biasing process according to some embodiments.
[0113] like Figure 6As shown in FIG, logic flow 600 may set the bias mode of the accelerator memory page to the host bias mode at block 602. For example, a host software application (e.g., software application 140) may set the bias mode of the accelerator device memory 430 to the host bias mode via a driver and / or API call. The host software application may use an API call (e.g., an OpenCL API) to convert the allocated (or target) page of the accelerator memory 430 storing the operand to the host bias mode. Since the allocated page is being converted from device bias mode to host bias mode, a cache flush is not initiated. The device bias mode may be specified in the bias table of bias information 432.
[0114] At block 604, the logic flow 600 may push operands and / or data to the accelerator memory pages. For example, the accelerator 420 may execute a function for the CPU that requires certain operands. The host software application may push operands from a peer CPU core (e.g., core 465a) to the allocated pages of the accelerator memory 430. The host processor 445 may generate operand data in the allocated pages in the accelerator memory 430 (and anywhere in the host memory 440).
[0115] At block 606, the logic flow 600 may convert the accelerator memory page to device bias mode. For example, the host software application may use an API call to convert the operand memory page of the accelerator memory 430 to device bias mode. When the device bias conversion is complete, the host software application may submit work to the accelerator 430. The accelerator 430 may perform the functions associated with the submitted work without the coherence overhead associated with the host.
[0116] At block 608, the logic flow 600 may generate a result using the operand by the accelerator and store the result in the accelerator memory page. For example, the accelerator 420 may use the operand to perform a function (e.g., floating point operation, graphics calculation, FFT operation, and / or similar function) to generate a result. The result may be stored in the accelerator memory 430. In addition, the software application may use an API call to cause the work descriptor to submit a flush of the operand page from the host cache. In some embodiments, a cache (or cache line) flush routine (such as CLFLUSH) over an intra-device protocol (e.g., IDI) protocol may be used to perform a cache flush. The result generated by the function may be stored in the allocated accelerator memory 430 page.
[0117] At block 610, the logic flow may set the bias mode of the accelerator memory page storing the result to the host bias mode. For example, the host software application may use an API call to convert the operand memory page of the accelerator memory 430 to the host bias mode without causing a consistency process and / or cache flush action. The host CPU 445 may access, cache, and share the result. At block 612, the logic flow 600 may provide the result from the accelerator memory page to the host software. For example, the host software application may access the result directly from the accelerator memory page 430. In some embodiments, the allocated accelerator memory page may be released by the logic flow. For example, the host software application may use a driver and / or API call to release the allocated memory page of the accelerator memory 430.
[0118] Figure 7 is a block diagram illustrating an architecture according to one or more examples of the present specification. In this case, a coherent accelerator architecture 700 is provided. The coherent accelerator architecture 700 is interconnected with an IAL endpoint 728, which communicatively couples the coherent accelerator architecture 700 to a host device, such as the host devices disclosed in the aforementioned figures.
[0119] A coherent accelerator fabric 700 is provided to communicatively couple an accelerator 740 and its attached memory 722 to a host device. The memory 722 includes a plurality of memory controllers 720-1 through 720-n. In one example, eight memory controllers 720 can service eight separate memory banks.
[0120] The fabric controller 736 includes a set of controllers and interconnects to provide a coherent memory fabric 700. In this example, the fabric controller 736 is divided into n separate slices to serve the n banks of memory 722. Each slice can be substantially independent of each other slice. As described above, the fabric controller 736 includes both a "vertical" interconnect 706 and a "horizontal" interconnect 708. Vertical interconnects can generally be understood as connecting upstream or downstream devices to each other. For example, the last level cache (LLC) 734 is vertically connected to the LLC controller 738, connected to the fabric, and connected to the intra-die interconnect (F2IDI) block that communicatively couples the fabric controller 736 to the accelerator 740. The F2IDI 730 provides a downstream link to the fabric stop 712 and can also provide a bypass interconnect 715. The bypass interconnect 715 connects the LLC controller 738 directly to the fabric-to-memory interconnect 716, where the signal is multiplexed out to the memory controller 720. In the non-bypass route, the request from F2IDI 730 travels along the horizontal interconnect to the host and then back to fabric stop 712 , then to fabric coherency engine 704 , and down to F2MEM 716 .
[0121] The horizontal bus includes a bus that interconnects the fabric stops 712 to each other and connects the LLC controllers to each other.
[0122] In one example, the IAL endpoint 728 may receive packets from a host device that include instructions to perform acceleration functions, as well as payloads that include snoops for accelerator operations. The IAL endpoint 728 passes these to the L2FAB 718, which acts as a host device interconnect for the fabric controller 736. The L2FAB 718 may act as a link controller for the fabric, including providing an IAL interface controller (although in some embodiments, additional IAL control elements may also be provided, and generally, any combination of elements that provide IAL interface control may be referred to as an "IAL interface controller"). The L2FAB 718 controls requests from the accelerator to the host and vice versa. The L2FAB 718 may also be an IDI agent and may need to act as a sorting agent between IDI requests from the accelerator and snoops from the host.
[0123] L2FAB 718 may then operate fabric stop 712-0 to populate the values into memory 722. Fabric stop L2FAB 718 may apply a load balancing algorithm, such as, for example, a simple address-based hash, to tag payload data for a particular destination bank. Once the banks in memory 722 are populated with the appropriate data, accelerator 740 operates fabric controller 736 to fetch the values from memory to LLC 734 via LLC controller 738. Accelerator 740 performs its accelerated computations and then writes the outputs to LLC 734, where they are then passed downstream and written out to memory 722.
[0124] In some examples, fabric bus 712, F2MEM controller 716, multiplexer 710, and F2IDI 730 can all be standard buses and interconnects that provide interconnectivity based on well-known principles. The aforementioned interconnects can provide virtual and physical channels, interconnects, buses, switching elements, and flow control mechanisms. They can also provide conflict resolution mechanisms related to the interaction between requests issued by the accelerator or device agent and requests issued by the host. The fabric can include a physical bus in the horizontal direction, with the server ring-shaped switching as the bus passes through each slice. The fabric can also include a specially optimized horizontal interconnect 739 between LLC controllers 738.
[0125] Requests from the F2IDI 730 can be passed through hardware to split and multiplex traffic to the host between the horizontal fabric interconnect and the optimized path per slice between the LLC controller 738 and the memory 722. This includes multiplexing the traffic and directing it to the IDI block, where it traverses the conventional route through the fabric stop 712 and FCE 704, or using IDI prime to direct the traffic to the bypass interconnect 715. The F2IDI 730-1 can also include hardware to manage ingress and egress to and from the horizontal fabric interconnect, for example by providing appropriate signaling to the fabric stop 712.
[0126] The IAL interface controller 718 can be a suitable PCIe controller. The IAL interface controller provides the interface between the packetized IAL bus and the fabric interconnect. It is responsible for queuing and providing flow control for IAL messages, directing IAL messages to the appropriate fabric physical and virtual channels. L2FAB 718 can also provide arbitration between multiple types of IAL messages and can further enforce IAL ordering rules.
[0127] At least three control structures within the fabric controller 736 provide the novel and advantageous features of the fabric controller 736 of the present description. These include the LLC controller 738, the FCE 704, and the power management module 750.
[0128] Advantageously, LLC controller 738 may also provide bias control functionality according to the IAL bias protocol. Thus, LLC controller 738 may include hardware for performing cache lookups, hardware for checking the IAL basis for cache miss requests, hardware for directing requests onto the appropriate interconnect path, and logic for responding to snoops issued by the host processor or by FCE 704.
[0129] When a request is directed from fabric stop 712 to the host through L2FAB 718 , LLC controller 738 determines where the traffic should be directed through fabric stop 712 , directly to F2MEM 716 through bypass interconnect 715 , or to another memory controller through horizontal bus 739 .
[0130] Note that in some embodiments, LLC controller 738 is a physically separate device or block from FCE 704. It is possible to provide a single block that provides the functionality of both LLC controller 738 and FCE 704. However, by separating the two blocks and providing IAL bias logic in LLC controller 738, it is possible to provide bypass interconnect 715, thereby accelerating certain memory operations. Advantageously, in some embodiments, separating LLC controller 738 and FCE 704 can also facilitate selective power gating in portions of the fabric to more efficiently use resources.
[0131] The FCE 704 may include hardware for queuing, processing (e.g., issuing snoops to the LLC), and tracking SMI requests from the host. This provides consistency with the host device. The FCE 704 may also include hardware for queuing requests on the per-chip optimized path to the memory banks within the memory 722. Embodiments of the FCE may also include hardware for arbitrating and multiplexing the two request classes described above onto the CMI memory subsystem interface, and may include hardware or logic for resolving conflicts between the two request classes described above. Other embodiments of the FCE may provide support for sequencing requests from the direct vertical interconnect and requests from the FCE 704.
[0132] The power management module (PMM) 750 also provides advantages to embodiments of the present specification. For example, consider a scenario where each individual slice in the fabric controller 736 vertically supports 1 GB per second of bandwidth. 1 GB per second is provided only as an illustrative example, and a real-world implementation of the fabric controller 736 may be much faster or much slower than 1 GB per second.
[0133] LLC 734 can have a higher bandwidth, for example, 10 times the vertical bandwidth of the slices of fabric controller 736. Thus, LLC 734 can have a bandwidth of 10 GB per second, which can be bidirectional, making the total bandwidth through LLC 734 20 GB per second. Thus, with 8 slices of fabric controller 736, each supporting 20 GB per second bidirectionally, accelerator 740 can see a total bandwidth of 160 GB per second through horizontal bus 739. Therefore, running LLC controller 738 and horizontal bus 739 at full speed consumes a lot of power.
[0134] However, as described above, the vertical bandwidth can be 1 GB per slice, and the total IAL bandwidth can be approximately 10 GB per second. Thus, the bandwidth provided by the horizontal bus 739 is approximately an order of magnitude higher than the bandwidth of the entire fabric controller 736. For example, the horizontal bus 739 can include thousands of physical lines, while the vertical interconnect can include hundreds of physical lines. The horizontal fabric 708 can support the full bandwidth of the IAL, i.e., 10 GB per second in each direction, for a total of 20 GB per second.
[0135] Accelerator 740 can perform computations and operate on LLC 734 at a much higher rate than the host device can consume data. Therefore, data can burst into accelerator 740 and then be consumed by the host processor as needed. Once accelerator 740 completes its computations and populates LLC 734 with appropriate values, maintaining full bandwidth between LLC controllers 738 consumes significant power, which is essentially wasted because LLC controllers 738 no longer need to communicate with each other when accelerator 740 is idle. Therefore, when accelerator 740 is idle, LLC controller 738 can be powered down, thereby shutting down horizontal bus 739, while keeping the appropriate vertical buses active—for example, from fabric stop 712 to FCE 704 to F2MEM 716 to memory controller 720—while also maintaining horizontal bus 708. Because horizontal bus 739 operates at a speed approximately an order of magnitude higher than the rest of fabric 700, this can save approximately an order of magnitude of power when accelerator 740 is idle.
[0136] Note that some embodiments of the coherent accelerator architecture 700 may also provide an isochronous controller that can be used to provide isochronous services to latency- or time-sensitive components. For example, if the accelerator 740 is a display accelerator, an isochronous display path can be provided to the display generator (DG) so that a connected display receives isochronous data.
[0137] The overall combination of agents and interconnects in the coherent accelerator fabric 700 implements the IAL functionality in a high-performance, deadlock-free, and starvation-free manner. It does so while saving energy while providing increased efficiency by bypassing the interconnect 715.
[0138] Figure 8 is a flow chart of a method 800 according to one or more examples of the present specification. The method 800 illustrates a method of power saving, such as that which can be performed by Figure 7 The PMM 750 offers
[0139] Inputs from the host device 804 may arrive at the coherent accelerator fabric, including instructions to perform computations and payloads for computations. In block 808, if the horizontal interconnect between the LLC controllers is powered down, the PMM powers up the interconnect to its full bandwidth.
[0140] The accelerator computes results according to its normal functions in block 812. In computing these results, it may operate the coherent accelerator fabric at its full available bandwidth, including the full bandwidth of the horizontal interconnect between LLC controllers.
[0141] When the results are complete, in block 816 , the accelerator fabric may flush the results to local memory 820 .
[0142] In decision block 824, the PMM determines whether there is new data available from the host that can be manipulated. If any new data is available, control returns to block 812 and the accelerator continues to perform its acceleration function. Meanwhile, the host device can consume data directly from local memory 820, which can be mapped to the host memory address space in a coherent manner.
[0143] Returning to block 824, if no new data is available from the host, the PMM reduces power, for example, by shutting down the LLC controllers, thereby disabling the high-bandwidth horizontal interconnect between the LLC controllers, in block 828. As described above, because the local memory 820 is mapped to the host address memory space, the host can continue to consume data from the local memory 820 at the full IAL bandwidth, which in some embodiments is much lower than the full bandwidth between the LLC controllers.
[0144] In block 832 , the controller waits for new input from the host device and when new data is received, the interconnect may be powered back up.
[0145] Figure 9-11 An example of an IAL.mem tunnel over PCIe is shown. The packet format depicted includes standard PCIe packet fields, except for the fields highlighted in gray. The gray fields are those that provide the new tunnel fields.
[0146] Figure 9 is a block diagram of an IAL.mem read via PCIe operation according to one or more examples of the present specification. The new fields include:
[0147] MemOpcode (4 bits) - Memory operation code. Contains information about the memory transaction that needs to be processed, such as read, write, no operation, etc.
[0148] MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together, they specify which metadata field in memory needs to be modified and to what value. Metadata fields in memory typically contain information associated with the actual data. For example, QPI stores directory status in metadata.
[0149] ●TC (2 bits) - Service Class. Used to distinguish services belonging to different service quality categories.
[0150] Snp Type (3 bits) - snoop type. Used to maintain consistency between the host's cache and the device's cache.
[0151] R (5 bits) - reserved
[0152] Figure 10is a block diagram of IAL.mem writes over PCIe operations according to one or more examples of the present specification. The new fields include:
[0153] MemOpcode (4 bits) - Memory operation code. Contains information about the memory transaction that needs to be processed. For example, read, write, no operation, etc.
[0154] MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together, they specify which metadata field in memory needs to be modified and to what value. Metadata fields in memory typically contain information associated with the actual data. For example, QPI stores directory status in metadata.
[0155] ●TC (2 bits) - Service Class. Used to distinguish services belonging to different service quality categories.
[0156] Snp Type (3 bits) - snoop type. Used to maintain consistency between the host's cache and the device's cache.
[0157] R (5 bits) - reserved
[0158] Figure 11 Figure 1 is a block diagram of IAL.mem data for PCIe operations according to one or more examples of this specification. The new fields include:
[0159] R (1 bit) - Reserved
[0160] Opcode (3 digits) - IAL.io opcode
[0161] MetaField and MetaValue (2 bits) - Metadata field and metadata value. Together, they specify which metadata field in memory needs to be modified and to what value. Metadata fields in memory typically contain information associated with the actual data. For example, QPI stores directory status in metadata.
[0162] PCLS (4 bits) - Previous cache line state. Used to identify coherency transitions.
[0163] ● PRE (7 bits) - Performance code. Used by performance monitoring counters in the host.
[0164] Figure 12An embodiment of a structure consisting of point-to-point links interconnecting a set of components according to one or more examples of the present specification is shown. System 1200 includes a processor 1205 and system memory 1210 coupled to a controller hub 1215. Processor 1205 includes any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or other processor. Processor 1205 is coupled to controller hub 1215 via a front-side bus (FSB) 1206. In one embodiment, FSB 1206 is a serial point-to-point interconnect as described below. In another embodiment, link 1206 includes a serial differential interconnect architecture that complies with a differential interconnect standard.
[0165] System memory 1210 includes any memory device, such as random access memory (RAM), nonvolatile (NV) memory, or other memory accessible to devices in system 1200. System memory 1210 is coupled to controller hub 1215 through memory interface 1216. Examples of memory interfaces include a double data rate (DDR) memory interface, a dual-channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.
[0166] In one embodiment, the controller hub 1215 is a root hub, root complex, or root controller in the Peripheral Component Interconnect Express (PCIe) interconnect hierarchy. Examples of controller hubs 1215 include a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. Typically, the term chipset refers to two physically separate controller hubs: a memory controller hub (MCH) coupled to an interconnect controller hub (ICH).
[0167] Note that current systems typically include an MCH integrated with the processor 1205, while the controller 1215 communicates with the I / O devices in a similar manner as described below. In some embodiments, peer-to-peer routing is optionally supported through the root complex 1215.
[0168] Here, controller hub 1215 is coupled to switch / bridge 1220 via serial link 1219. Input / output modules 1217 and 1221 (also referred to as interfaces / ports 1217 and 1221) include / implement a layered protocol stack to provide communication between controller hub 1215 and switch 1220. In one embodiment, multiple devices can be coupled to switch 1220.
[0169] The switch / bridge 1220 routes packets / messages from the device 1225 upstream (i.e., up the hierarchy toward the root complex) to the controller hub 1215 and downstream (i.e., down the hierarchy away from the root controller) from the processor 1205 or system memory 1210 to the device 1225. In one embodiment, the switch 1220 is referred to as a logical assembly of multiple virtual PCI-to-PCI bridge devices.
[0170] Device 1225 includes any internal or external device or component to be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-on card, an audio processor, a network processor, a hard drive, a storage device, a CD / DVDROM, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a Firewire device, a Universal Serial Bus (USB) device, a scanner, and other input / output devices. Typically in PCIe native language, for example, a device is referred to as an endpoint. Although not specifically shown, device 1225 can include a PCIe to PCI / PCI-X bridge to support traditional or other versions of PCI devices. Endpoint devices in PCIe are typically classified as traditional, PCIe, or root complex integrated endpoints.
[0171] The accelerator 1230 is also coupled to the controller hub 1215 via a serial link 1232. In one embodiment, the graphics accelerator 1230 is coupled to the MCH, which is coupled to the ICH. The switch 1220, and therefore the I / O devices 1225, are then coupled to the ICH. The I / O modules 1231 and 1218 are also used to implement a layered protocol stack for communication between the graphics accelerator 1230 and the controller hub 1215. Similar to the MCH discussion above, the graphics controller or graphics accelerator 1230 itself can be integrated into the processor 1205.
[0172] In some embodiments, the accelerator 1230 may be an accelerator, e.g. Figure 7 accelerator 740 that provides consistent memory for processor 1205.
[0173] To support IAL over PCIe, the controller hub 1215 (or another PCIe controller) may include extensions to the PCIe protocol, including, as non-limiting examples, a mapping engine 1240, a tunneling engine 1242, a host bias to device bias flipping engine 1244, and a QoS engine 1246.
[0174] The mapping engine 1240 can be configured to provide opcode mapping between PCIe instructions and IAL.io (IOSF) opcodes. IOSF provides a non-uniform in-order semantics protocol and can provide services such as device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and DMA-style data transfers, by way of non-limiting example. Native PCIe can provide corresponding instructions, so in some cases, the mapping can be a one-to-one mapping.
[0175] The tunnel engine 1242 provides IAL.mem (SMI) tunneling over PCIe. This tunnel enables the host (e.g., a processor) to map accelerator memory into the host memory address space and read from and write to the accelerator memory in a coherent manner. SMI is a transactional memory interface that can be used by the coherence engine on the host to transmit IAL transactions over the PCIe tunnel in a coherent manner. An example of a modified packet structure for such a tunnel is shown in Figure 9-11 In some cases, a special field for the tunnel may be allocated within one or more DVSEC fields of the PCIe packet.
[0176] The host bias to device bias flip engine 1244 provides the accelerator device with the ability to flush the host cache line (required for host to device bias flipping). This can be accomplished using non-allocating zero-length writes (i.e., writes without byte enables set) at the cache line granularity from the accelerator device over PCIe. Non-allocating semantics can be described using transactions and transaction hints on the transaction layer packet (TLP). For example:
[0177] TH=1, PH=01
[0178] This enables the device to invalidate a given cache line, thus enabling it to access its own memory space without losing coherency. The device can issue a read after the page bias flip to ensure that all lines are flushed. The device can also implement a CAM to ensure that no new requests for that line are received from the host while the flip is in progress.
[0179] The QoS engine 1246 can divide the IAL traffic into two or more virtual channels to optimize the interconnect. For example, these can include a first virtual channel (VC0) for MMIO and configuration operations, a second virtual channel (VC1) for host-to-device writes, and a third virtual channel (VC2) for host-to-device reads.
[0180] Figure 13An embodiment of a layered protocol stack according to one or more embodiments of the present specification is shown. The layered protocol stack 1300 includes any form of layered communication stack, such as a Quick Path Interconnect (QPI) stack, a PCIe stack, a next generation high performance computing interconnect stack, or other layered stack. Although immediately below reference is made to Figure 12-15 The discussion is given with respect to the PCIe stack, but the same concepts can be applied to other interconnect stacks. In one embodiment, the protocol stack 1300 is a PCIe protocol stack, including a transaction layer 1305, a link layer 1310, and a physical layer 1320.
[0181] Such as Figure 12 The interfaces 1217, 1218, 1221, 1222, 1226, and 1231 in the example may be represented as a communication protocol stack 1300. The representation as a communication protocol stack may also be referred to as a module or interface that implements / includes the protocol stack.
[0182] PCIe uses packets to transfer information between components. Packets are formed in the transaction layer 1305 and the data link layer 1310 to transfer information from a sending component to a receiving component.
[0183] As the transmitted packets flow through the other layers, they are expanded with additional information necessary to process the packets at those layers. On the receiving side, the reverse process occurs and the packet is transformed from its Physical Layer 1320 representation to the Data Link Layer 1310 representation and finally (for Transaction Layer packets) into a form that can be processed by the Transaction Layer 1305 of the receiving device.
[0184] Transaction Layer
[0185] In one embodiment, the transaction layer 1305 is used to provide an interface between the device's processing core and the interconnect architecture, such as the data link layer 1310 and the physical layer 1320. In this regard, the primary responsibility of the transaction layer 1305 is the assembly and disassembly of packets, namely transaction layer packets (TLPs). The transaction layer 1305 generally manages credit-based flow control of TLPs. PCIe implements split transactions, i.e., transactions with request and response separated by time, allowing the link to carry other traffic while the target device collects the response data.
[0186] Additionally, PCIe utilizes credit-based flow control. In this scheme, a device advertises an initial amount of credit for each receive buffer in the transaction layer 1305. External devices at the opposite end of the link, such as Figure 1 The controller hub 115 in the Transaction Processor calculates the number of credits consumed by each TLP. If the transaction does not exceed the credit limit, the transaction can be sent. Upon receiving a response, a certain number of credits are restored. One advantage of the credit scheme is that if the credit limit is not reached, the delay in credit return does not affect performance.
[0187] In one embodiment, the four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more read requests and write requests to transfer data to or from a memory mapped location. In one embodiment, memory space transactions can use two different address formats, for example, a short address format, such as a 32-bit address, or a long address format, such as a 64-bit address. Configuration space transactions are used to access the configuration space of a PCIe device. Transactions to the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.
[0188] Thus, in one embodiment, transaction layer 1305 assembles packet header / payload 1306. The current format of packet header / payload can be found in the PCIe specification at the PCIe specification website.
[0189] Figure 14 An embodiment of a PCIe transaction descriptor according to one or more examples of the present specification is shown. In one embodiment, transaction descriptor 1400 is a mechanism for carrying transaction information. In this regard, transaction descriptor 1400 supports the identification of transactions in the system. Other potential uses include tracking modifications to default transaction ordering and associating transactions with channels.
[0190] Transaction descriptor 1400 includes a global identifier field 1402, an attribute field 1404, and a channel identifier field 1406. In the example shown, global identifier field 1402 is depicted as including a local transaction identifier field 1408 and a source identifier field 1410. In one embodiment, global transaction identifier 1402 is unique for all pending requests.
[0191] According to one implementation, the local transaction identifier field 1408 is a field generated by the requesting agent and is unique for all pending requests that need to be completed for that requesting agent. Furthermore, in this example, the source identifier 1410 uniquely identifies the requester agent within the PCIe hierarchy. Thus, together with the source ID 1410, the local transaction identifier 1408 field provides a global identification of transactions within the hierarchy domain.
[0192] Attribute field 1404 specifies the characteristics and relationships of the transaction. In this regard, attribute field 1404 may be used to provide additional information that allows modification of the default handling of the transaction. In one embodiment, attribute field 1404 includes a priority field 1412, a reserved field 1414, a sort field 1416, and a non-snooping field 1418. Here, priority subfield 1412 can be modified by the initiator to assign a priority to the transaction. Reserved attribute field 1414 is reserved for future or vendor-defined use. Reserved attribute fields can be used to implement possible usage models that utilize priority or security attributes.
[0193] In this example, the ordering attribute field 1416 is used to provide optional information conveying the type of ordering that can modify the default ordering rules. According to one example implementation, an ordering attribute of "0" indicates that the default ordering rules are to be applied, while an ordering attribute of "1" indicates that the ordering is relaxed, and writes can be passed in the same direction as writes, and read completions can be passed in the same direction as writes. The snoop attribute field 1418 is used to determine whether the transaction is snooped. As shown, the channel ID field 1406 identifies the channel associated with the transaction.
[0194] Link Layer
[0195] The link layer 1310 (also known as the data link layer 1310) acts as an intermediate level between the transaction layer 1305 and the physical layer 1320. In one embodiment, the responsibility of the data link layer 1310 is to provide a reliable mechanism for exchanging TLPs between the two link components. One side of the data link layer 1310 accepts the TLPs assembled by the transaction layer 1305, applies a packet sequence identifier 1311, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., a CRC 1312, and submits the modified TLP to the physical layer 1320 for transmission across the physical layer to an external device.
[0196] Physical layer
[0197] In one embodiment, the physical layer 1320 includes a logical sub-block 1321 and an electronic block 1322 to physically transmit packets to an external device. Here, the logical sub-block 1321 is responsible for the "digital" functions of the physical layer 1321. In this regard, the logical sub-block includes a transmit portion for preparing outgoing information for transmission by the physical sub-block 1322, and a receive portion for identifying and preparing received information before passing it to the link layer 1310.
[0198] Physical block 1322 includes a transmitter and a receiver. The transmitter is provided with symbols by logic sub-block 1321, which serializes the symbols and transmits them to an external device. The receiver is provided with serialized symbols from the external device and converts the received signal into a bit stream. The bit stream is deserialized and provided to logic sub-block 1321. In one embodiment, 8b / 10b transmission code is used, in which 10-bit symbols are transmitted and received. Here, special symbols are used to frame packets using frames 1323. In addition, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.
[0199] As described above, although the transaction layer 1305, link layer 1310, and physical layer 1320 are discussed with reference to a specific embodiment of the PCIe protocol stack, the layered protocol stack is not limited thereto. In fact, any layered protocol may be included / implemented. For example, a port / interface represented as a layered protocol includes: (1) a first layer for assembling packets, i.e., a transaction layer; a second layer for sequencing packets, i.e., a link layer; and a third layer for transmitting packets, i.e., a physical layer. As a specific example, the Common Standard Interface (CSI) layered protocol is used.
[0200] Figure 15 An embodiment of a PCIe serial point-to-point structure according to one or more examples of the present specification is shown. Although an embodiment of a PCIe serial point-to-point link is shown, the serial point-to-point link is not limited thereto as it includes any transmission path for transmitting serial data. In the embodiment shown, the basic PCIe link includes two low-voltage differential drive signal pairs: a transmit pair 1506 / 1511 and a receive pair 1512 / 1507. Therefore, device 1505 includes transmit logic 1506 for transmitting data to device 1510 and receive logic 1507 for receiving data from device 1510. In other words, two transmit paths, namely paths 1516 and 1517, and two receive paths, namely paths 1518 and 1519, are included in the PCIe link.
[0201] A transmission path is any path used to transmit data, such as a transmission line, copper wire, fiber, wireless communication channel, infrared communication link, or other communication path. A connection between two devices (e.g., device 1505 and device 1510) is called a link, such as link 1515. A link can support one lane—each lane represents a set of differential signal pairs (one pair for transmit and one pair for receive). To expand bandwidth, a link can aggregate multiple lanes, represented by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.
[0202] A differential pair refers to two transmission paths, such as lines 1516 and 1517, for transmitting differential signals. For example, when line 1516 switches from a low voltage level to a high voltage level (i.e., a rising edge), line 1517 switches from a high logic level to a low logic level (i.e., a falling edge). Differential signals potentially exhibit better electrical characteristics, such as improved signal integrity, cross-coupling, voltage overshoot / undershoot, and ringing. This allows for better timing windows, which enables faster transmission frequencies.
[0203] The foregoing summarizes features of one or more embodiments of the subject matter disclosed herein. These embodiments are provided to enable a person of ordinary skill in the art (PHOSITA) to better understand the various aspects of the present disclosure. Reference may be made to certain readily understood terms and underlying technologies and / or standards without detailed description. It is expected that PHOSITA will possess or obtain access to background knowledge or information in those technologies and standards sufficient to implement the teachings of this specification.
[0204] PHOSITA will understand that they can readily use this disclosure as a basis for designing or modifying other processes, structures or variations to achieve the same purposes and / or achieve the same advantages of the embodiments described herein. PHOSITA will also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they can make various changes, substitutions and alterations herein without departing from the spirit and scope of the present disclosure.
[0205] In the foregoing description, certain aspects of some or all embodiments have been described in more detail than is strictly necessary to implement the appended claims. These details are provided for non-limiting illustrative purposes only, to provide context and illustration of the disclosed embodiments. These details should not be understood as being required, and the claims should not be "understood" as limiting. The wording may refer to "one embodiment" or "an embodiment." These words and any other references to an embodiment should be broadly understood to refer to any combination of one or more embodiments. In addition, several features disclosed in a particular "embodiment" may also be distributed across multiple embodiments. For example, if features 1 and 2 are disclosed in an "embodiment," embodiment A may have feature 1 but lack feature 2, while embodiment B may have feature 2 but lack feature 1.
[0206] This specification may provide an explanation in a block diagram format, in which certain features are disclosed in separate boxes. These should be broadly understood to disclose how various features interoperate, but are not meant to imply that these features must necessarily be embodied in separate hardware or software. In addition, where a single block discloses more than one feature in the same block, those features do not necessarily have to be embodied in the same hardware and / or software. For example, computer "memory" may in some cases be distributed or mapped between multiple levels of cache or local memory, main memory, battery-backed volatile memory, and various forms of persistent memory (e.g., hard disks, storage servers, optical disks, tape drives, or similar devices). In certain embodiments, some components may be omitted or merged. In general, the arrangements depicted in the figures may be more logical in their representation, while the physical architecture may include various permutations, combinations, and / or hybrids of these elements. Countless possible design configurations may be used to achieve the operational objectives outlined herein. Therefore, the associated infrastructure has countless alternative arrangements, design choices, device possibilities, hardware configurations, software implementations, and device options.
[0207] Reference may be made here to computer-readable media, which may be tangible and non-transitory computer-readable media. As used in this specification and throughout the claims, "computer-readable media" should be understood to include one or more computer-readable media of the same or different types. As non-limiting examples, computer-readable media may include optical drives (e.g., CD / DVD / Blu-ray), hard drives, solid-state drives, flash memory, or other non-volatile media. Computer-readable media may also include media such as read-only memories (ROMs), FPGAs or ASICs configured to execute desired instructions, stored instructions for programming FPGAs or ASICs to execute desired instructions, intellectual property (IP) blocks that can be integrated into other circuits in hardware, or instructions encoded directly into hardware or microcode on a processor such as a microprocessor, digital signal processor (DSP), microcontroller, or any other suitable component, device, element, or object, as appropriate, according to specific needs. Non-transitory storage media herein are explicitly intended to include any non-transitory dedicated or programmable hardware configured to provide the disclosed operations or to cause a processor to perform the disclosed operations.
[0208] Throughout this specification and claims, various elements may be "communicatively," "electrically," "mechanically," or otherwise "coupled" to one another. This coupling may be direct, point-to-point, or may involve intermediary devices. For example, two devices may be communicatively coupled to one another via a controller that facilitates communication. Devices may be electrically coupled to one another via an intermediary device such as a signal booster, voltage divider, or buffer. Mechanically coupled devices may be indirectly mechanically coupled.
[0209] Any "module" or "engine" disclosed herein may refer to or include software, a software stack, hardware, a combination of firmware and / or software, circuitry configured to perform the functions of the engine or module, or any computer-readable medium as described above. Where appropriate, these modules or engines may be provided on or in conjunction with a hardware platform that may include hardware computing resources such as processors, memory, storage, interconnects, networks and network interfaces, accelerators, or other suitable hardware. Such a hardware platform may be provided as a single monolithic device (e.g., in a PC form factor), or with some or part of the functionality distributed (e.g., a "compound node" in a high-end data center, where computing, memory, storage, and other resources may be dynamically allocated and need not be local to each other).
[0210] Flowcharts, signal flow charts or other diagrams showing operations performed in a particular order may be disclosed herein. Unless expressly stated otherwise, or unless required in a particular context, the order should be understood to be merely non-limiting examples. In addition, where one operation is shown following another operation, other intermediate operations may also occur, which may be related or unrelated. Some operations may also be performed simultaneously or in parallel. Where an operation is referred to as being "based on" or "according to" another item or operation, this should be understood to imply that the operation is at least partially based on or at least partially based on other items or operations. This should not be interpreted as implying that an operation is based only or exclusively on the item or operation, or only or exclusively on the item or operation.
[0211] All or part of any hardware element disclosed herein can be readily provided in a system on chip (SoC), including a central processing unit (CPU) package. SoC refers to an integrated circuit (IC) that integrates the components of a computer or other electronic system into a single chip. Thus, for example, a client device or a server device can be provided in whole or in part in a SoC. A SoC can contain digital, analog, mixed-signal, and radio frequency functions, all of which can be provided on a single chip substrate. Other embodiments may include a multi-chip module (MCM), in which multiple chips are located within a single electronic package and are configured to interact closely with each other through the electronic package.
[0212] In a general sense, any appropriately configured circuit or processor can execute any type of instruction associated with data to implement the operations detailed herein. Any processor disclosed herein can transform an element or article (e.g., data) from one state or thing to another state or thing. In addition, based on specific needs and implementations, the information tracked, sent, received, or stored in the processor can be provided in any database, register, table, cache, queue, control list, or storage structure, all of which can be referenced within any appropriate timeframe. Any memory or storage element disclosed herein should be considered to be appropriately included in the broad terms "memory" and "storage device."
[0213] The computer program logic that implements all or part of the functions described herein is embodied in various forms, including but not limited to source code form, computer executable form, machine instructions or microcode, programmable hardware and various intermediate forms (e.g., forms generated by assemblers, compilers, linkers or locators). In one example, the source code includes a series of computer program instructions implemented in various programming languages, such as object code, assembly language or high-level languages such as OpenCL, FORTRAN, C, C++, JAVA or HTML, for various operating systems or operating environments, or implemented in hardware description languages such as Spice, Verilog and VHDL. The source code can define and use various data structures and communication messages. The source code can be in computer executable form (e.g., by an interpreter), or the source code can be converted (e.g., by a translator, assembler or compiler) into a computer executable form, or converted into an intermediate form, such as bytecode. Where appropriate, any of the foregoing can be used to construct or describe appropriate discrete or integrated circuits, whether sequential, combinatorial, state machines or other.
[0214] In an example embodiment, any number of circuits in the accompanying drawings can be implemented on a board of a related electronic device. The board can be a general-purpose circuit board that can hold various components of the internal electronic system of the electronic device and also provide connectors for other peripheral devices. Based on specific configuration requirements, processing requirements, and computing design, any suitable processor and memory can be appropriately coupled to the board. Note that, using the numerous examples provided herein, interactions can be described based on two, three, four, or more electronic components. However, this is done for clarity and example purposes only. It should be appreciated that the system can be merged or reconfigured in any suitable manner. Along similar design alternatives, any of the components, modules, and elements shown in the accompanying drawings can be combined in various possible configurations, all of which are within the broad scope of this specification.
[0215] Numerous other changes, substitutions, variations, alterations, and modifications may be ascertained by those skilled in the art, and the present disclosure is intended to encompass all such changes, substitutions, variations, replacements, and modifications that fall within the scope of the appended claims. To assist the United States Patent and Trademark Office (USPTO), and in addition, any reader of any patent issued in this application, in interpreting the appended claims, applicants wish to note that applicants: (a) do not intend any of the appended claims to invoke 35 USC section 112 (pre-AIA) paragraph 6(6) or paragraph (f) of the same section (post-AIA) as it existed on the filing date thereof, unless the phrase "means for" or "step for" is specifically used in a particular claim; and (b) do not intend, by any statement in the specification, to limit the present disclosure in any manner not expressly reflected in the appended claims.
[0216] Example Implementation
[0217] In one example, a peripheral component interconnect express (PCIe) controller is disclosed for providing consistent memory mapping between an accelerator memory and a host memory address space, including: a PCIe controller hub including an extension for providing a consistent accelerator interconnect (CAI) to provide bias-based coherency tracking between the accelerator memory and the host memory address space; wherein the extension includes: a mapping engine for providing opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions for the CAI; and a tunneling engine for providing a scalable memory interconnect (SMI) tunnel of host memory operations to the accelerator memory via the CAI.
[0218] An example is further disclosed wherein the opcode mapping is a one-to-one mapping.
[0219] An example is further disclosed wherein the CAI is Accelerator Link (IAL)-compatible interconnect.
[0220] An example is further disclosed wherein the OSF instructions include instructions for performing operations selected from the group consisting of device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and direct memory access (DMA)-style data transfer.
[0221] An example is further disclosed that also includes a host bias to device bias (HBDB) flip engine to enable the accelerator to flush host cache lines.
[0222] An example is further disclosed wherein the HBDB rollover engine is operable to provide a transaction layer packet (TLP) hint including TH=1, PH=01.
[0223] An example is further disclosed, further comprising a QoS engine comprising a plurality of virtual channels.
[0224] An example is further disclosed in which the virtual channel includes VC0 for memory mapped input / output (MMIO) and configuration transactions, VC1 for host to accelerator writes, and VC2 for host to read from the accelerator.
[0225] An example is further disclosed wherein the extension is configured to provide a non-post write (NP Wr) opcode.
[0226] An example is further disclosed wherein the NP Wr opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.
[0227] An example is further disclosed wherein the extension is configured to provide for reading over a PCIe packet format comprising a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0228] An example is further disclosed wherein the extension is configured to provide writes over a PCIe packet format including a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0229] An example is further disclosed wherein the extension is configured to provide data read completion via a PCIe packet format including a 3-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 4-bit PCLS, and a 7-bit PRE.
[0230] Further disclosed are examples of interconnects that include a PCIe controller.
[0231] Examples of systems including the interconnect are further disclosed.
[0232] Examples of the system are further disclosed, including a system on a chip.
[0233] Examples of the system are further disclosed, including a multi-chip module.
[0234] Further disclosed are examples of the system wherein the accelerator is a software-assisted device memory.
[0235] Further disclosed are examples of the system wherein the accelerator is an autonomous device memory.
[0236] Further disclosed are examples of one or more tangible, non-transitory computer-readable media having instructions stored thereon to provide a Peripheral Component Interconnect Express (PCIe) controller on a host platform to provide a consistent memory mapping between an accelerator memory and a host memory address space, including instructions to provide a PCIe controller hub including an extension to provide a Coherent Accelerator Interconnect (CAI) to provide bias-based coherency tracking between the accelerator memory and the host memory address space; wherein the extension includes: a mapping engine to provide opcode mapping between PCIe instructions and System-on-Chip Fabric (OSF) instructions for the CAI; and a tunneling engine to provide Scalable Memory Interconnect (SMI) tunneling of host memory operations to the accelerator memory via the CAI.
[0237] An example is also disclosed where the opcode mapping is a one-to-one mapping.
[0238] An example is further disclosed wherein the CAI is Accelerator Link (IAL)-compatible interconnect.
[0239] An example is further disclosed wherein the OSF instructions include instructions for performing operations selected from the group consisting of device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and direct memory access (DMA)-style data transfer.
[0240] An example is further disclosed wherein the instructions also provide a host bias to device bias (HBDB) flip engine to enable the accelerator to flush host cache lines.
[0241] An example is further disclosed wherein the HBDB rollover engine is operable to provide a transaction layer packet (TLP) hint including TH=1, PH=01.
[0242] An example is further disclosed, further comprising a QoS engine comprising a plurality of virtual channels.
[0243] An example is further disclosed in which the virtual channel includes VC0 for memory mapped input / output (MMIO) and configuration traffic, VC1 for host to accelerator writes, and VC2 for host to read from the accelerator.
[0244] An example is further disclosed wherein the extension is configured to provide a non-post write (NP Wr) opcode.
[0245] An example is further disclosed wherein the NP Wr opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.
[0246] An example is further disclosed wherein the extension is configured to provide for reading over a PCIe packet format comprising a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0247] An example is further disclosed wherein the extension is configured to provide writes over a PCIe packet format including a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0248] An example is further disclosed wherein the extension is configured to provide data read completion via a PCIe packet format including a 3-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 4-bit PCLS, and a 7-bit PRE.
[0249] Further disclosed is an example of a computer-implemented method of providing a Peripheral Component Interconnect Express (PCIe) controller to a host platform to provide consistent memory mapping between accelerator memory and host memory address space, comprising: providing PCIe controller hub services; providing extensions to the PCIe controller hub to provide a Coherent Accelerator Interconnect (CAI) to provide bias-based coherency tracking between the accelerator memory and the host memory address space; wherein providing the extensions comprises: providing an opcode mapping between PCIe instructions and System-on-Chip Fabric (OSF) instructions for the CAI; and providing a Scalable Memory Interconnect (SMI) tunnel of host memory operations to the accelerator memory via the CAI.
[0250] An example is further disclosed wherein the opcode mapping is a one-to-one mapping.
[0251] An example is further disclosed wherein the CAI is Accelerator Link (IAL)-compatible interconnect.
[0252] An example is further disclosed wherein the OSF instructions include instructions for performing operations selected from the group consisting of device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and direct memory access (DMA)-style data transfer.
[0253] An example is further disclosed that also includes providing a host bias to device bias (HBDB) flip service to enable an accelerator to flush a host cache line.
[0254] An example is further disclosed wherein providing the HBDB rollover service includes providing a transaction layer packet (TLP) hint including TH=1, PH=01.
[0255] An example is also disclosed, further comprising providing QoS services, including providing multiple virtual channels.
[0256] An example is further disclosed in which the virtual channel includes VC0 for memory mapped input / output (MMIO) and configuration traffic, VC1 for host to accelerator writes, and VC2 for host to read from the accelerator.
[0257] An example is further disclosed wherein providing the extension includes providing a non-post write (NP Wr) opcode.
[0258] An example is further disclosed wherein the NP Wr opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.
[0259] An example is further disclosed wherein providing the extension includes providing a read through a PCIe packet format including a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0260] An example is further disclosed wherein providing the extension includes providing writing via a PCIe packet format including a 4-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 2-bit timecode, and a 3-bit snp type.
[0261] An example is further disclosed wherein providing the extension includes providing a data read completion via a PCIe packet format including a 3-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 4-bit PCLS, and a 7-bit PRE.
[0262] An example of an apparatus comprising means for performing the method is further disclosed.
[0263] An example is further disclosed wherein means for performing the method includes a processor and a memory.
[0264] An example is further disclosed wherein the memory includes machine-readable instructions that, when executed, cause the apparatus to perform the method.
[0265] An example is also disclosed wherein the apparatus is a computing system.
[0266] Further disclosed are examples of at least one computer-readable medium comprising instructions that, when executed, implement a method or realize an apparatus as described in the foregoing examples.
Claims
1. A Peripheral Component Interconnect Express (PCIe) controller for providing a coherent memory mapping between an accelerator memory and a host memory address space, comprising: a peripheral component interconnect express controller hub comprising circuitry for providing a coherent accelerator interconnect (CAI) to provide bias-based coherency tracking between the accelerator memory and the host memory address space; Wherein, the circuit is used for: providing an opcode mapping between PCI-Fast instructions and on-chip system fabric (OSF) instructions for the coherent accelerator interconnect; and A Scalable Memory Interconnect (SMI) tunnel of host memory operations to the accelerator memory is provided via the coherent accelerator interconnect.
2. The PCI Express controller according to claim 1, wherein: The opcode mapping is a one-to-one mapping.
3. The PCI Express controller according to claim 1, wherein: The Coherent Accelerator Interconnect is an Intel® Accelerator Link (IAL)-compatible interconnect.
4. The PCI Express controller according to claim 1, wherein: The system-on-chip architecture instructions include instructions for performing operations selected from the group consisting of device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and direct memory access (DMA)-style data transfers.
5. The PCI Express controller according to claim 1, wherein: The circuitry is further configured to enable the accelerator to flush a host cache line.
6. The PCI Express controller according to claim 5, wherein: The circuit is for enabling the accelerator to flush a host cache line based on a transaction layer packet (TLP) hint including TH=1, PH=01.
7. The PCI Express controller of claim 1, further comprising a QoS engine, the QoS engine comprising a plurality of virtual channels.
8. The PCI Express controller according to claim 7, wherein: The virtual channels include VC0 for memory mapped input / output (MMIO) and configuration transactions, VC1 for host to accelerator writes, and VC2 for host to read from the accelerator.
9. The PCI Express controller according to claim 1, wherein: The circuit is used to provide a non-issue write (NP Wr) opcode.
10. The PCI Express controller according to claim 9, wherein: The non-issued write opcode includes the reserved Peripheral Component Interconnect Fast fields Fmt[2:0]=011b and Type[4:0]=1101b.
11. The PCI Express controller according to any one of claims 1 to 10, wherein: The circuit is used to provide a read through a PCI Express packet format, the PCI Express packet format including a 4-bit opcode, a 2-bit meta field, a 2-bit meta value, a 2-bit service class, and a 3-bit snp type.
12. The PCI Express controller according to any one of claims 1 to 10, wherein: The circuit is used to provide writing through a PCI Express packet format, the PCI Express packet format including a 4-bit opcode, a 2-bit meta field, a 2-bit meta value, a 2-bit service class, and a 3-bit snp type.
13. The PCI Express controller according to any one of claims 1 to 10, wherein: The circuit is used to provide data read completion via a PCI Express packet format, wherein the PCI Express packet format includes a 3-bit opcode, a 2-bit metafield, a 2-bit metavalue, a 4-bit PCLS, and a 7-bit PRE.
14. An interconnect for providing accelerator coherence, comprising the PCI Express controller according to any one of claims 1 to 13.
15. A system comprising the interconnect of claim 14.
16. The system of claim 15 comprising a system on a chip.
17. The system of claim 15 comprising a multi-chip module.
18. The system of any one of claims 15 to 17, wherein: The accelerator is a software-assisted device memory.
19. The system of any one of claims 15 to 17, wherein: The accelerator is autonomous device memory.
20. One or more tangible, non-transitory computer-readable media having stored thereon instructions for providing a Peripheral Component Interconnect Express (PCIe) controller on a host platform to provide a coherent memory mapping between accelerator memory and host memory address spaces, comprising instructions for: providing a peripheral component interconnect express controller hub including circuitry for providing a coherent accelerator interconnect (CAI) to provide bias-based coherency tracking between the accelerator memory and the host memory address space; in, The circuit is used to: providing an opcode mapping between PCI-Fast instructions and on-chip system fabric (OSF) instructions for the coherent accelerator interconnect; and A Scalable Memory Interconnect (SMI) tunnel of host memory operations to the accelerator memory is provided via the coherent accelerator interconnect.
21. The one or more tangible, non-transitory media of claim 20, wherein: The opcode mapping is a one-to-one mapping.
22. The one or more tangible, non-transitory media of claim 20, wherein: The Coherent Accelerator Interconnect is an Intel® Accelerator Link (IAL)-compatible interconnect.
23. The one or more tangible, non-transitory media of claim 20, wherein: The system-on-chip architecture instructions include instructions for performing operations selected from the group consisting of device discovery, device configuration, error reporting, interrupt provisioning, interrupt handling, and direct memory access (DMA)-style data transfers.
24. The one or more tangible, non-transitory media of claim 20, wherein: The circuitry is further configured to enable the accelerator to flush a host cache line.
25. The one or more tangible, non-transitory media of claim 24, wherein: The circuit is for enabling the accelerator to flush a host cache line based on a transaction layer packet (TLP) hint including TH=1, PH=01.
Citation Information
Patent Citations
Data processor
US20100153656A1
Providing Hardware Support For Shared Virtual Memory Between Local And Remote Physical Memory
US20110072234A1