Coherent storage devices via PCIe
Patent Information
- Application Number
- DE102018006797
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-08-28
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2038-08-28
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of SpecificationThis disclosure relates generally to the field of interconnect devices, and more particularly, but not exclusively, to a system and method for coherent memory devices via peripheral component interconnect express (PCIe).Prior ArtComputing systems include various components to manage requests for processor resources. For example, designers may provide a hardware accelerator (or "accelerator") operatively coupled to a central processing unit (CPU). Generally, an accelerator is an autonomous element configured to perform functions delegated to it by the CPU. An accelerator may be designed for specific functions and / or programmable. For example, an accelerator may be configured to perform specific computations, graphical functions, and / or the like. When an accelerator performs an assigned function, the CPU is idle to dispatch resources for other requests. In conventional systems, the operating system (OS) may manage physical memory available in the computing system (e.g., "system memory"); however, the OS does not manage or allocate memory that is local to an accelerator. As a result, memory protection mechanisms such as cache coherency introduce inefficiencies in accelerator-based configurations. For example, conventional cache coherency mechanisms limit the ability of an accelerator to access its associated local memory at very high bandwidth, and / or limit deployment options for the accelerator.Document US 2011 / 0072234 A1 discloses providing hardware support for shared virtual storage between local and remote physical storage. Document US 2010 / 0153656 A1 discloses data processors.Brief Description of the DrawingsThe present disclosure will be best understood from the following detailed description when read with the accompanying drawings. It is to be appreciated that, in accordance with common practice in this field, various features are not necessarily drawn to scale and are used for purposes of illustration only. Where a scale is explicitly or implicitly shown, this provides only one illustrative example. In other embodiments, the dimensions of the various functions may be increased or decreased arbitrarily in the interest of clearer discussion. FIG. 1 illustrates an example of an operating environment that may be representative of various embodiments according to one or more examples of the present specification. FIG. 2 a illustrates an example of a full coherence operating environment, in accordance with one or more examples of the present specification. FIG. 2 b illustrates an example of a full coherence operating environment, in accordance with one or more examples of the present specification. FIG. 2 c illustrates an example coherence engine without impacting operating environment, according to one or more examples of the present specification. FIG. 3 illustrates an example of an operating environment that may be representative of various embodiments according to one or more examples of the present specification. FIG. 4 illustrates another example of an operating environment that may be representative of various embodiments according to one or more examples of the present specification. FIGS. 5 aand 5 b illustrate further examples of operating environments that may be representative of various embodiments according to one or more examples of the present specification. FIG. 6 illustrates an example of a logic flow according to one or more examples of the present specification. FIG. 7 is a block diagram illustrating a structure according to one or more examples of the present specification. FIG. 8 is a block diagram illustrating a method according to one or more examples of the present specification. FIG. 9 is a block diagram of a read operation of a Intel® accelerator link (IAL.mem) memory under PCIe operation, in accordance with one or more examples of the present specification. FIG. 10 is a block diagram of a write operation of an IAL.mem under PCIe operation, in accordance with one or more examples of the present specification. FIG. 11 is a block diagram of an IAL.mem termination with data under PCIe operation, in accordance with one or more examples of the present specification. FIG. 12 illustrates an embodiment of a structure composed of point-to-point connections connecting a group of components, according to one or more examples of the present specification. FIG. 13 illustrates an example of an embodiment of a layered protocol stack according to one or more examples of the present specification. FIG. 14 illustrates an embodiment of a PCIe transaction descriptor according to one or more examples of the present specification. FIG. 15 illustrates an embodiment of a PCIe serial point-to-point structure according to one or more examples of the present specification.Embodiments of the DisclosureThe Intel® Accelerator Link (IAL) of the present specification is an extension of the connection link of the Roseta Link (R-link) Multi-chip Packet (MCP). IAL extends the R-link protocol to enable it to support accelerators and input / output (IO) devices that may not be adequately supported by the basic R-link or Peripheral Component Interconnect Express (PCIe) protocols.The following disclosure provides many different embodiments or examples for implementing different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present disclosure. Of course, these are mere examples and are not intended to be limiting. Further, the present disclosure may repeat numerals and / or letters in the various examples. This repetition is for purposes of simplicity and clarity and, as such, does not establish a relationship between the various embodiments and / or configurations discussed. Different embodiments have different advantages, and no particular advantage is necessarily claimed by any of the embodiments.In the following description, numerous specific details are set forth, such as examples of specific processor types and system configurations, specific hardware structures, specific architectural and micro architectural details, specific register configurations, specific instruction types, specific system components, specific measurements / elevations, specific processor pipeline stages and operation(s), etc., in order to provide a thorough understanding of the present invention. It will be apparent, however, to one skilled in the art that these specific details need not be employed to practice the present invention.In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuitry / code for described algorithms, specific firmware code, specific interconnect operation, specific logic configurations, specific fabrication techniques and materials, specific compiler implementations, specific terms of algorithms in code, specific shutdown and gate techniques / logic, and other specific operational details of computer systems have not been described in detail to avoid unnecessarily obscuring the present invention.Although the following embodiments may be described with reference to power conservation and power efficiency in specific integrated circuits, such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of embodiments described herein may be applied to other types of circuits or semiconductor devices that may also benefit from better energy efficiency and conservation of energy. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks™ and may also be used in other devices such as handheld devices, tablets, other flat notebooks, system on a chip (SOC) devices, and embedded applications.Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld personal computers (PCs). Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip (SoC), network personal computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations taught below. Moreover, the apparatuses, methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimizations for energy conservation and efficiency. As will be readily apparent in the following description, the embodiments of methods, apparatus, and systems described herein (whether in terms of hardware, firmware, software, or a combination thereof) are essential for a "green technology" future balanced with performance considerations.Various embodiments may be generally directed to techniques for providing cache coherency between a plurality of components within a processing system. In some embodiments, the plurality of components may include a processor such as a central processing unit (CPU) and a logic device communicatively coupled to the processor. In various embodiments, the logic device may include locally attached memory. In some embodiments, the plurality of components may include a processor communicatively coupled to an accelerator, including locally-coupled memory (e.g., memory of the logic device).In some embodiments, the processing system may perform a coherency bias process configured to provide a plurality of cache coherency processes. In some embodiments, the plurality of cache coherency processes may include a device bias process and a host bias process (collectively, "bias protocol flows"). In some embodiments, the host bias process may direct requests to the locally attached memory of the logic device through a coherency component of the processor, including requests from the logic device. In some embodiments, the device bias process may direct logic device memory requests to logic device memory directly to the logic device memory, for example without incorporating the coherence component of the processor. In various embodiments, the cache coherency process may switch between the device bias process and the host bias processes based on a bias indicator established using application software, hardware alerts, a combination thereof, and / or the like. The embodiments are not limited in this context.The IAL described in this specification uses an optimized accelerator protocol (UHD), which is another extension of the R-link MCP connection protocol. The IAAF may be used to provide a connection structure to an accelerator device in one example (the accelerator device may be a high performance accelerator that performs graphics processing, dense computation, smartNIC services, or the like in some examples). The accelerator may have its own attached accelerator memory, and an interconnect structure such as an IAL or a PCIe-based structure may be used to connect the processor to the accelerator in some embodiments. The interconnect structure may be a coherent accelerator structure, in which case the accelerator memory may be mapped into the memory address space of the host device. The coherent accelerator structure may maintain coherency within the accelerator and between the accelerator and the host device. This can be used to implement prior art memory and coherency support in these accelerator types.Advantageously, coherent accelerator structure according to the present specification may provide optimizations that increase efficiency and throughput. For example, an accelerator may include a certain number of n memory banks, with corresponding n last level caches (LLCs) each being controlled by an LLC controller. The fabric may provide different types of connections for connecting the accelerator and its caches to the memory and for connecting the fabric to the host device.For purposes of illustration, in the course of this specification, buses or links connecting devices of the same nature are referred to as "horizontal" links, while links or buses connecting different devices in the upstream and downstream directions may be referred to as "vertical" links. The terms "horizontal" and "vertical" are used herein alone for ease of understanding and are not intended to imply or imply a required physical arrangement of the connections or buses that they must be physically perpendicular to each other on a die.For example, an accelerator may include an 8 memory banks, with corresponding 8 LLCs, which may be level 3 (L3) caches, each being controlled by an LLC controller. The coherent accelerator structure may be divided into a number of independent "slices.". Each slice serves a database and its associated LLC and operates substantially independently of the other slices. In one example, each slice may exploit the bias operations provided by the IAL and provide parallel paths to the memory bank. Memory operations involving the host device may be directed through a fabric coherence engine (FCE) that provides coherence with the host device. However, the LLC of each individual slice may also have a parallel bypass path that writes directly to the memory bank that connects the LLC to the memory bank directly bypassing the FCE. This may be performed, for example, by providing the bias logic (e.g., host bias or accelerator bias) in the LLC controller itself. The LLC controller may be physically separate from the FCE and located in the vertical orientation before the FCE, allowing an accelerator bias store operation to transition the FCE and write it directly to a memory bank.Embodiments of the present specification may also achieve substantial power savings by providing a power manager that shuts down portions of the coherent fabric when not in use. For example, the accelerator may be a very wide bandwidth accelerator that can perform many operations per second. While the accelerator performs its accelerated function, it takes up much of the structure and requires an extremely large bandwidth, so that calculated values can be stored in the memory in good time after their calculation. However, after completion of a calculation, the host device may not be ready for use of the data. In this case, portions of the link, such as vertical buses from the FCE to the LLC controller, as well as horizontal buses between the LLC controllers and the LLCs themselves, may be disabled. These may remain turned off until the accelerator receives new data with which it is to operate.The following table illustrates several classes of accelerators. Note that the base R link can support only the two first accelerator classes, while the IAL can support all five.1. Producer-Consumer• PCIe Basic Setup•Network Controllers•Crypto Crypto Crypto•Compression2. Producer-Consumer plus•PCIe Devices Having Special Purpose•Current Lake Data Center FabricRequirements (e.g., special ordering requirements)•Infiniband HBA3. SW-•Accelerator Accelerator•Discrete FPGAssupported device storesm memories connected to them•Graphic Graphic•Uses where software "data placement" is practicable4. Autonomous Setup Method•Accelerators with•Offload of Dense ComputationThe calibration of the calibration of the calibration system is also more easily adjustablem memories connected thereto•GPGPU•Uses where software "data placement" is impractical5. Giant Cache•Accelerator with m memory connected•Offload of Dense Computation•GPGPU•Uses wherein the data footprint is larger than the attached memoryNote that embodiments of these accelerators may require some level of cache coherency to support usage models, except producer consumers. Thus, IAL is a coherent accelerator link.IAL uses a combination of three protocols that are dynamically bundled into a common link to enable the accelerator models disclosed above. These protocols include:• Intel® On-chip System Fabric (IOSF) - A reformatted PCIe-based link that provides a non-coherent ordered semantic protocol. IOSF may include an on-chip implementation of all or part of the PCIe standard. IOSF packetizes PCIe traffic so that it can be sent to an associated die, such as a system-on-a-chip (SoC) or multi-chip module (MCM). IOSF enables device discovery, device configuration, error message, interrupts, direct memory access (DMA) type data transfers, and various services provided as part of the PCIe standard.• In-die Interconnect (IDI) - One device enables the output of coherent read and write requests to a processor.• Scalable Memory Interconnect (SMI) - allows a processor to access memory coupled to an accelerator.These three protocols may be used in different combinations (e.g., IOSF alone, IOSF plus IDI, IOSF plus IDI plus SMI, IOSF plus SMI) to support various ones of the models described in the table above.As a basis, IAL provides a single link or bus definition that can cover all five accelerator models through the combination of the protocols mentioned above. Note that producer-consumer accelerators are essentially PCIe accelerators. They require only the IOSF protocol, which is already a reformatted version of PCIe. IOSF supports some operations of the accelerator interface architecture (AiA), such as Enqueue (ENQ) instruction support, which may not be supported by industry standard PCIe devices. IOSF therefore provides added value to this class of accelerator as compared to PCIe. "producer-consumer plus" accelerators are accelerators that can only utilize the IDI and IOSF layers of the IAL.Software-assisted device memory and autonomous device memory may, in some embodiments, require the SMI protocol at the IAL, including special operation codes (opcodes) at SMI and special control support for operations associated with these opcodes in the processor. These additions support the coherence bias model of IAL. The use may all support IOSF, IDI, and SMI.The giant cache use also employs IOSF, IDI, and SMI, but may also add new flags to the IDI and SMI protocols specifically designed for use with giant cache accelerators (i.e., not employed in the device memory models discussed above). Giant cache can add new special control support in the processor that is not required for any of the other uses.IAL refers to these three protocols as IAL.IO, IAL.cache, and IAL.mem. The combination of these three protocols provides the desired performance benefit for the five accelerator models.To achieve these advantages, IAL may provide for physical layers according to R-link (for MCP) or Flexbus (discrete) to enable dynamic multiplexing of the 10, cache and mem protocols.However, some form factors do not natively support the physical layers according to R-link or Flexbus. In particular, class 3 and 4 device memory accelerators may not support R-link or Flexbus. Present examples of this may use standard PCIe that limits the devices to a private memory model, rather than providing coherent memory that may be mapped into the write back memory address space of the host device. This model is limited because the memory associated with the device is thus not directly addressable by software. This may result in suboptimal data partitioning between the host and device memories via a bandwidth-limited PCIe link.Thus, embodiments of the present specification provide coherence semantics that follow the same bias model-based definition defined by IAL that maintains the benefit of coherence without traditional overhead incurred. All this can be provided via an existing PCIe physical link.Thus, some of the advantages of the IAL may be realized via a physical layer that does not provide dynamic multiplexing between the IO, cache, and mem protocols provided by R-link and Flexbus. Advantageously, enabling an IAL protocol over PCIe for certain device classes reduces the workload on the ecosystem of devices that use PCIe physical links. It enables the advancement of existing PCIe infrastructure, including the use of off-the-shelf components such as switches, root ports, and endpoints. This also allows a connected memory device to be used more easily across the platform, using the traditional private memory model or the coherent system addressable memory model according to suitability for installation.To support Classes 3 and 4 devices as described above (software-supported memory and autonomous device memory), the IAL components may be mapped as follows:IOSF or IAL.io may use standard PCIe. This can be used for device discovery, listing, configuration, error message, interrupts, and DMA-type data transfers.SMI or IAL.mem may use SMI tunneling over PCIe. Details of SMI tunneling over PCIe are described below, including tunneling described in FIGS. 9, 10, and 11 below.IDI or IAL.cache are not supported in certain embodiments of this specification. IDI allows the device to issue coherent read and write requests to a host memory. Although IAL.cache may not be supported, the methods disclosed herein may be used to enable bias-based coherency for device-connected memories.To achieve this result, the accelerator device may use one of its standard PCIe memory base address register (BAR) regions in the size of its attached memory. To do so, the device may use an intended Vendor Specific Enhanced Capability (DVSEC) similar to the standard IAL to point to the BAR region to be mapped to the coherent address space. Further, the DVSEC may explain additional information such as memory type, latency, and other attributes that are helpful for the basic input / output system (BIOS) to map this memory to system address decoders in the coherent region. The BIOS can then program the memory base and limit the host physical address in the device.This allows the host to read attached device memory using standard memory read (MRd) opcode.However, non-posted semantics may be required for writes, as access to metadata may be needed at completion. To obtain NP MWr on PCIe, the following reserved encodings may be used:• Fmt[2:0] - 011b• Type[4:0] - 11011bThe use of a novel non-posted memory write (NP MWr) to PCIe has the additional benefit of enabling AiA-ENQ instructions for efficient jobs to the device.To achieve the best quality of service, embodiments of the present specification may implement three different virtual channels (VC0, VC1, and VC2) to separate different traffic types as follows:• VC0 → Total Memory-Mapped Input / Output (MMIO) and Configuration (CFG) traffic both upstream and downstream.• VC1 → IAL.mem writes (from host to device)• VC2 → IAL.mem reads (from host to device)Note that since IAL.cache or IDI are not supported, executions of this specification may not allow the accelerator device to issue coherent reads or writes to the host memory.Embodiments of this specification may also have the capability to flush cache lines from the host (required for host-to-device bias reversal). This may be done using a zero length non-mapping write from the device on the PCIe at cache line granularity. Non-mapping semantics are described using transaction and processing hints at the transaction layer packets (TLPs).• TH=1, PH=01This allows the host to invalidate a given line. The device may issue a read operation following the bias reversal of a page to ensure that all rows have been emptied. The device may also implement a content addressable memory (CAM) to ensure that no new requests to the row are received from the host during the execution of a flip-flip.A system and method for coherent memory devices over PCIe is described below with reference to the accompanying FIGURES. Note that during all the figures, certain indices may be repeated to indicate that a particular device or block in the figures is complete or substantially consistent. However, this is not intended to imply any particular relationship between the various embodiments disclosed. In certain examples, a genus of elements may be designated by a particular reference code ("widget 10"), while individual species or examples of the genus may be designated by a code separated by -- ("first specific widget 10-1" and "second specific widget 10-2").FIG. 1 illustrates an example of an operating environment 100, which may be representative of various embodiments according to one or more examples of the present specification. The operating environment 100 illustrated in FIG. 1 may include a device 105 including a processor 110, such as a central processing unit (CPU). Processor 110 may include any type of computing element, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computer (CISC) microprocessor, a reduced instruction set microprocessor (RISC), a very long instruction word (VLIW) microprocessor, a virtual processor such as a virtual central processing unit (VCPU), or any other type of processor or processing circuit. In some embodiments, processor 110 may be one or more processors of the Intel@ family of processors available from Intel® Corporation of Santa Clara, U.S. Federal California. Although only one processor 110 is shown in FIG. 1, an apparatus may include a plurality of processors 110. The processor 110 may include a processing element 112, e.g., a processing core. In some embodiments, processor 110 may include a multi-core processor including a plurality of processing cores. In various embodiments, processor 110 may include processor memory 114 including, for example, a processor cache or local cache memory to enable efficient access to data processed by processor 110. In some embodiments, processor memory 114 may include random access memory (RAM); however, processor memory 114 may be implemented using other types of memory, such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, and / or the like.As shown in FIG. 1, processor 110 may be communicatively coupled to a logic device 120 via a link 115. In various embodiments, the logic device 120 may include a hardware device. In various embodiments, the logic device 120 may include an accelerator. In various embodiments, the logic device 120 may include a hardware accelerator. In various embodiments, the logic device 120 may include an accelerator implemented in hardware, software, or a combination thereof.Although an accelerator may be used as example logic device 120 in this detailed description, embodiments are not limited in this respect, as logic device 120 includes any types of devices, processors (e.g., a graphics processing unit (GPU)), logic units, circuits, integrated circuits, application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), memory units, compute units, and / or the like suitable for operation in accordance with some embodiments. In an embodiment where the logic device 120 includes an accelerator, the logic device 120 may be configured to perform one or more functions for the processor 110. For example, the logic device 120 may include an accelerator operable to perform graphics functions (e.g., a GPU or a graphics accelerator), floating point operations, fast Fourier transform (FFT) operations, and / or the like. In various embodiments, the logic device 120 may include an accelerator configured to operate using various hardware components, standards, protocols, and / or the like. Non-limiting examples of types of accelerator and / or accelerator technology suitable for use by the logic device may include OpenCAPI™ CCIX, GenZ, NVIDIA® NVLink™ accelerator interface architecture (AiA), cache coherent agent (CCA), global mapped and coherent device memory (GCM), Intel® graphical media accelerators (GMA), Intel® controlled input / output virtualization technology (IO) (e.g., VT-d, VT-x, and / or the like), shared virtual memory (SVM), and / or the like. The embodiments are not limited in this context.The logic device 120 may include a processing element 122, e.g., a processing core. In some embodiments, the logic device 120 may include a plurality of processing elements 122. The logic device 120 may comprise logic device memory 124, for example designed as a locally connected memory for the logic device 120. In some embodiments, the logic device memory 124 may include local memory, cache memory, and / or the like. In various embodiments, the logic device memory 124 may include random access memory (RAM); however, the logic device memory 124 may be implemented using other types of memory such as dynamic RAM (DRAM), synchronous DRAM (SDRAM), combinations thereof, and / or the like. In some embodiments, at least a portion of the logic device memory 124 may be visible or accessible to the processor 110. In some embodiments, at least a portion of logic device memory 124 may be visible or accessible to processor 110 as system memory (e.g., as an accessible portion of system memory 130).In various embodiments, processor 110 may execute a driver 118. In some embodiments, driver 118 may be operable to control various functional aspects of logic device 120 and / or manage communication with one or more applications that use logic device 120 and / or computational results generated by logic device 120. In various embodiments, the logic device 120 may include and / or access bias information 126. In some embodiments, bias information 126 may include information associated with a coherency bias process. For example, the bias information 126 may include information indicating which cache coherency process may be active for the logic device 120 and / or a particular process, application, thread, storage operation, and / or the like. In some embodiments, the bias information 126 may be read, written, or otherwise managed by the driver 118.In some embodiments, link 115 may include a bus component, such as a system bus. In various embodiments, link 115 may include a communication link operable to support multiple communication protocols (e.g., a multi-protocol link). Supported communication protocols may include standard load / store IO protocols for component communication, including serial link protocols, device caching protocols, storage protocols, memory semantics protocols, directory bit support protocols, network protocols, coherency protocols, accelerator protocols, data storage protocols, point-to-point protocols, fabric-based protocols, packet-internal (or on-chip) protocols, fabric-based packet-internal protocols, and / or the like. Non-limiting examples of supported communication protocols may include Peripheral Component Interconnect (PCI) protocol, Peripheral Component Interconnect Express (PCIe or PCI-E) protocol, Universal Serial Bus (USB) protocol, Serial Peripheral Interface (SPI) protocol, Serial AT Attachment (SATA) protocol, Intel® QuickPath Interconnect (QPI) protocol, Intel® UltraPath Interconnect (UPI) protocol, Intel® Optimized Accelerator Protocol (OAP) protocol, "Intel® accelerator link" (IAL) link, "Intra-device interconnect" (IDI) protocol (or IAL.cache), "Intel® On-chip scalable fabric" (IOSF) protocol (or IAL.io), "scalable memory interconnect" (SMI) protocol (or IAL.mem), "SMI 3rd Generation" (SMI 3), and / or the like. In some embodiments, link 115 may support an in-device protocol (e.g., IDI) and a storage connection protocol (e.g., SMI3). In various embodiments, link 115 may support an in-device protocol (e.g., IDI), a storage connection protocol (e.g., SMI3), and a structure-based protocol (e.g., IOSF).In some embodiments, device 105 may include system memory 130. In various embodiments, system memory 130 may include main memory for device 105. System memory 130 may store data and sequences of instructions executed by processor 110 or by any other device or component of device 105. In some embodiments, system memory 130 may include RAM; however, system memory 130 may be implemented using other types of memory, such as dynamic DRAM, SDRAM, combinations thereof, and / or the like. In various embodiments, system memory 130 may store a software application 140 (e.g., "host software") executable by processor 110. In some embodiments, software application 140 may use or otherwise be associated with logic device 120. For example, the software application 140 may be configured to use computational results generated by the logic device 120.The apparatus may include coherency logic 150 to provide cache coherency processes. In various embodiments, coherence logic 150 may be implemented in hardware, software, or a combination thereof. In some embodiments, at least a portion of coherence logic 150 may be wholly or partially located in or otherwise associated with processor 110. For example, in some embodiments, coherency logic 150 for a cache coherency element or process 152 may be disposed within processor 110. In some embodiments, processor 110 may include a coherency controller 116 for performing various cache coherency processes, such as a cache coherency process 152. In some embodiments, cache coherency process 152 may include one or more standard cache coherency techniques, functions, methods, processes, elements (including hardware or software elements), protocols, and / or the like executed by processor 110. In general, cache coherency process 152 may include a standard protocol for managing caches of a system so that data is not lost or overwritten before the data is transferred from a cache to a target memory. Non-limiting examples of standard protocols executed or supported by cache coherency process 152 may include snooping-based (or snoopy) protocols, write invalidation protocols, write update protocols, directory-based protocols, hardware-based protocols (e.g., a modified exclusive shared invalid (MESI) protocol), private storage-based protocols, and / or the like. In some embodiments, cache coherency process 152 may include one or more standard cache coherency protocols to maintain cache coherency for a logic device 120 having a connected logic device memory 124. In some embodiments, cache coherency process 150 may be implemented in hardware, software, or a combination thereof.In some embodiments, coherency logic 150 may include coherency bias processes such as a host bias process or element 154 and a device bias process or element 156. Generally, coherency bias processes may operate to maintain cache coherency with respect to requests, data flows, and / or other memory operations with respect to logic device memory 122. In some embodiments, at least a portion of coherency logic such as host bias process 154, device bias process 156, and / or a bias selection component 158 may be external to processor 110, for example, in one or more individual coherency logic units 150. In some embodiments, the host bias process 154, the device bias process 156, and / or the bias selection component 158 may be implemented in hardware, software, or a combination thereof.In some embodiments, the host bias process 154 may include techniques, processes, data flows, data, algorithms, and / or the like that process requests for the logic device memory 124 by the cache coherency process 152 of the processor 110, including requests from the logic device 120. In various embodiments, the device bias process 156 may include techniques, processes, data flows, data, algorithms, and / or the like that enable the logic device 120 to directly access the logic device memory 124, for example, without using the cache coherency process 152. In some embodiments, the bias selection process 158 may include techniques, processes, data flows, data, algorithms, and / or the like for enabling the host bias process 154 or the device bias process 156 as an active bias process for requests associated with the logic device memory. In various embodiments, the active bias process may be based on bias information 126, which may include data, data structures, and / or processes that may be used by the bias selection process to set the active bias process and / or set the active bias process.FIG. 2 a illustrates an example of a full coherence operating environment 200A. The operating environment 200A illustrated in FIG. 2 amay include a device 202 including a CPU 210 including a plurality of processing cores 212 a- n. As shown in FIG. 2a, the CPU may include various protocol agents, such as caching agent 214, home agent 216, storage agent 218, and / or the like. Generally, caching agent 214 may operate to initiate transactions in coherent memory and maintain copies in its own cache structure. Caching agent 214 may be defined by the messages it can forward and access according to behaviors defined in a cache coherency protocol associated with the CPU. Caching agent 214 may also provide copies of the coherent memory content to other caching agents (e.g., accelerator caching agent 224). Home agent 216 may be responsible for the protocol side of memory interactions for CPU 210, including coherent and non-coherent home agent protocols. For example, home agent 216 may instruct memory read / write operations. The home agent 216 may be configured to service coherent transactions including handshaking with caching agents as needed. The space agent 216 may operate to monitor a portion of the coherent memory of the CPU 210, for example, maintaining coherence for a given address space. The home agent 216 may be responsible for managing conflicts that may arise among the different caching agents. For example, home agent 216 may provide the appropriate data and ownership responses as needed through a given transaction flow. The storage agent 218 may operate to manage access to memory. For example, the storage agent 218 may enable storage operations (e.g., load / store operations) and functions (e.g., swaps and / or the like) for the CPU 210.As shown in FIG. 2 a, the device 202 may include an accelerator 220 operatively connected to the CPU 210. The accelerator 220 may include an accelerator engine 222 operable to perform functions (e.g., computations and / or the like) offloaded from the CPU 210. Accelerator 220 may include an accelerator caching agent 224 and a storage agent 228.The accelerator 220 and the CPU 210 may be configured in accordance with and / or include various conventional hardware and / or memory access techniques. According to the illustration in FIG. 2 a, for example, all memory accesses including those initiated by the accelerator 220 must also pass through the path 230. Path 230 may include a non-coherent link, such as a PCIe link. In the configuration of device 202, accelerator engine 222 may be capable of directly accessing accelerator caching agent 224 and storage agent 228, but not caching agent 214, home agent 216, or storage agent 218. Similarly, cores 212a-n would not be able to directly access storage agent 228. Accordingly, the memory behind the memory agent 228 would not be part of the system address map recognized by the cores 212 a- n. Since cores 212a-n cannot access a common storage agent, data exchange is possible only via copies. In certain implementations, a driver may be used to enable copying back and forth between the storage agents 218 and 228. For example, drivers may have a runtime element that establishes shared memory abstraction that conceals all copies from the programmer. In contrast, as detailed description proceeds, some embodiments may provide configurations in which requests from an accelerator engine may be forced to traverse a link between the accelerator and the CPU when the accelerator engine wants to access accelerator memory, such as via an accelerator agent 228.FIG. 2 b illustrates an example of a non-coherent operating environment 200B. The operating environment 200B illustrated in FIG. 2 bmay include an accelerator 220 including an accelerator home agent 226. The CPU 210 and accelerator 220 may be operatively connected via a non-coherent path 232 such as a UPI path or a CCIX path.For operation of device 204, accelerator engine 222 and cores 212 a- nmay access both storage agents 228 and 218. Cores 212 a- nmay access memory 218 without traversing link 232, and accelerator agent 222 may access memory 228 without traversing link 232. The price of these local accesses from 222 to 228 is that home agent 226 must be set up to track coherency for all accesses from cores 212a-n to memory 228. This requirement results in complexity and high resource usage when the device 204 includes multiple CPU devices 210, all interconnected via other instances of the link 232. Home agent 226 must be able to track coherency for all cores 212a-n at all instances of CPU 210. This can become comparatively complicated in view of performance, area and performance, in particular in the case of large configurations. In particular, this negatively affects performance efficiency of accesses between accelerator 222 and memory 228 in favor of accesses from CPUs 210, even though the accesses from CPUs 210 are likely to be relatively rare.FIG. 2 crepresents an example of a coherency engine without bias operating environment 200C. As shown in FIG. 2, device 206 may include accelerator 220 operatively connected to CPU 210 via coherent links 236 and 238. The accelerator 220 may include an accelerator engine 222 operable to perform functions (e.g., computations and / or the like) offloaded from the CPU 210. Accelerator 220 may include an accelerator caching agent 224, an accelerator home agent 226, and a storage agent 228.In the configuration of the device 206, the accelerator 220 and the CPU 210 may be configured and / or include various conventional hardware and / or memory access techniques such as CCIX, GCM, standard coherency protocols (e.g., symmetric coherency protocols). According to the illustration in FIG. 2, for example, all memory accesses, including those initiated by the accelerator 220, must pass through the path 230. In this manner, the accelerator 220 must traverse the CPU 220 (and therefore the coherency protocols associated with the CPU) to access the accelerator memory (e.g., through the memory agent 228). Accordingly, the apparatus may not provide the ability to access particular memory, such as accelerator-attached memory associated with the accelerator 220, as part of system memory (e.g., as part of a system address map), thereby allowing host software to establish whether the address and access computational results of the accelerator 220 without the overhead of, for example, IO random access memory (DMA) copies of data. Such copies of data may require driver calls, interrupts, and MMIO access, all of which are inefficient and complex compared to memory accesses. The inability to access accelerator-attached memories without cache coherency overhead as shown in FIG. 2c may affect the execution time of a computation offloaded to accelerator 220. In a process that involves streaming of memory write traffic to a large extent, cache coherency overhead may bisect the effective write bandwidth detected by accelerator 220.The efficiency of operand setup, access to results, and accelerator computation plays a role in determining efficiency and exploiting the offloading of operations from the CPU 210 to the accelerator 220. If the price for the removal of operations is too high, the removal may not benefit or may be limited to only very extensive operations. Accordingly, various designers have configured accelerators that attempt to increase the efficiency of using an accelerator, such as accelerator 220, with limited efficiency as compared to technology configured in accordance with some embodiments. For example, certain conventional GPUs may operate without mapping the accelerator-connected memory as part of the system address map or without using certain virtual memory configurations (e.g., SVMs) to access the accelerator-connected memory. Accordingly, in such systems, memory associated with the accelerator is not visible to host system software. Instead, memory coupled to the accelerator is accessed only over a runtime layer of software provided by the device driver of the GPUs. A system of data copies and page table manipulations is used to generate the appearance of a virtual memory (e.g., SVM) activated system. Such a system is inefficient, particularly compared to some embodiments, because, among other things, the system requires memory replication, memory pinning, storage copies, and complex software. Such requirements result in great overhead at memory page transition points, which are not required in systems designed in accordance with some embodiments. In certain other systems, conventional hardware coherency mechanisms are employed for memory operations associated with memory associated with the accelerator, limiting the ability of an accelerator to access the memory associated with the accelerator at high bandwidth, and / or limiting the deployment options for a given accelerator (e.g., accelerators connected via an intra-packet or an out-of-packet link may not be supported without substantial bandwidth loss).Generally, conventional systems may use one or two methods for accessing accelerator-connected memory: a full coherency method (or full hardware coherency) or a private memory model or method. The full coherency method requires that all memory accesses comprising accesses requested by an accelerator in accelerator-connected memory must traverse the coherency protocol of the associated CPU. In this way, the accelerator must take a cumbersome way to access accelerator-associated memory, as the request must be transmitted at least to the associated CPU through the CPU coherency protocol and then to the accelerator-associated memory. Accordingly, the full coherence method carries overhead when an accelerator accesses its own memory, which can substantially degrade the data bandwidth that an accelerator can extract from its own connected memory. The private memory model requires significant resources and time overhead, such as memory replication, page pinning requirements, bandwidth cost for page copy data, and / or page transition cost (e.g., translation buffer (TLB) eliminations, page table manipulation, and / or the like). Accordingly, some embodiments may provide a coherency bias process configured to provide a plurality of cache coherency processes that provide, among other things, better memory usage and improved performance for systems having accelerator-connected memory compared to conventional systems.FIG. 3 illustrates an example of an operating environment 300 that may be representative of various embodiments. The operating environment 300 illustrated in FIG. 3 may include a device 305 operative to provide a coherence bias process, in accordance with some embodiments. In some embodiments, device 305 may include a CPU 310 that includes a plurality of processing cores 312 a- nand various protocol agents such as caching agent 314, home agent 316, storage agent 318, and / or the like. The CPU 310 may be communicatively coupled to the accelerator 320 using various links 335, 340. Accelerator 320 may include an engine 312 and a storage agent 318, and may include or access bias information 338.As shown in FIG. 3, the accelerator engine 322 may be communicatively coupled directly to the storage agent 328 via a biased coherency bypass 330. In various embodiments, the accelerator 320 may be configured to operate in a device bias process, wherein the biased coherency bypass 330 enables the storage requests of the accelerator engine 322 to directly access accelerator-connected storage (not shown) using the storage agent 328. In various embodiments, accelerator 320 may be configured to operate in a host bias process in which memory operations associated with memory associated with accelerator may be processed via links 335, 340 using cache coherency protocols of the CPU, for example, via caching agent 314 and home agent 316. Accordingly, the accelerator 320 of the device 305 may promote coherency protocols of the CPU 310 as appropriate (e.g., when a non-accelerator unit requests accelerator-connected memory) while allowing the accelerator 320 to directly access accelerator-connected memory via a biased coherency bypass 330.In some embodiments, coherency bias (e.g., regardless of whether device bias or host bias is active) may be stored in bias information 338. In various embodiments, bias information 338 may include and / or be stored in various data structures, such as a data table (e.g., a "bias table"). In various embodiments, bias information 338 may include a bias indicator having a value indicative of the active bias (e.g., 0=Host bias, 1=Geräte bias). In some embodiments, the bias information 338 and / or the bias indicator may be at various levels of granularity, such as memory regions, page tables, address ranges, and / or the like. For example, bias information 338 may indicate that certain memory pages are set for device bias while other memory pages are set for host bias. In some embodiments, bias information 338 may include a bias table configured to operate as a low cost scalable snoop filter.FIG. 4 illustrates an example of an operating environment 400 that may be representative of various embodiments. The operating environment 400 illustrated in FIG. 4 may include a device 405 operative to provide a coherence bias process, in accordance with some embodiments. The device 405 may include an accelerator 410 communicatively coupled to a host processor 445 via a link (or multi-protocol link) 489. The accelerator 410 and the host processor 445 may communicate via a link using connection structures 415 and 450, respectively, that facilitate the exchange of data and messages therebetween. In some embodiments, link 489 may include a multi-protocol link operable to support multiple protocols. For example, the link 489 and connection structures 415 and 450 may support various communication protocols, including, without limitation, serial link protocols, device caching protocols, storage protocols, storage semantics protocols, directory bit support protocols, network protocols, coherency protocols, accelerator protocols, data storage protocols, point-to-point protocols, structure-based protocols, intra-packet (or on-chip) protocols, structure-based intra-packet protocols, and / or the like. Non-limiting examples of supported communication protocols may include PCI, PCIe, USB, SPI, SATA, QPI, UPI, OAP, IAL, IDI, IOSF, SMI, SMI3, and / or the like. In some embodiments, link 489 and connection structures 415 and 450 may support an in-device protocol (e.g., IDI) and a storage connection protocol (e.g., SMI3). In various embodiments, link 489 and connection structures 415 and 450 may support an intra-device protocol (e.g., IDI), a storage connection protocol (e.g., SMI3), and a structure-based protocol (e.g., IOSF).In some embodiments, accelerator 410 may include bus logic 435 including a device TLB 437. In some embodiments, bus logic 435 may be or include PCIe logic. In various embodiments, bus logic 435 may communicate over link 480 using a fabric-based protocol (e.g., IOSF) and / or a peripheral component interconnect express (PCIe or PCI-E) protocol. In various embodiments, communication over link 480 may be used for various functions including, without limitation, discovery, register access (e.g., registers of accelerator 410 (not shown)), configuration, initialization, interrupts, random access memory, and / or address translation services (ATS).The accelerator 410 may include a core 420 having a host memory cache 422 and an accelerator memory cache 424. The core 420 may communicate using the link 481, for example, via an in-device protocol (e.g., IDE), for various functions such as coherent requests and memory flows. In various embodiments, accelerator 410 may include coherency logic 425 that includes or accesses bias mode information 427. Coherency logic 425 may communicate using link 482, for example, via a memory link protocol (e.g., SMI3). In some embodiments, communication over link 482 may be used for memory flows. The accelerator 410 may be operatively connected to the accelerator memory 430 (e.g., as memory connected to the accelerator) that may store bias information 432.In various embodiments, host processor 445 may be operatively coupled to host memory 440, and may include coherency logic (or coherency and cache logic) 455 having a last level cache (LLC) 457. Coherence logic 455 may communicate using various connections such as connections 484 and 485. In some embodiments, connections 484 and 485 may include a storage connection protocol (e.g., SMI3) and / or an in-device protocol (e.g., IDI). In some embodiments, the LLC 457 may include a combination of at least a portion of the host memory 440 and the accelerator memory 430.The host processor 445 may include bus logic 460 including an input-output memory management unit (IOMMU) 462. In some embodiments, bus logic 460 may be or include PCIe logic. In various embodiments, bus logic 460 may communicate over links 486 and 488 using a fabric-based protocol (e.g., IOSF) and / or a peripheral component interconnect express (PCIe or PCI-E) protocol. In various embodiments, host processor 445 may include a plurality of cores 465 a- n, each including a cache 467 a- n. In some embodiments, cores 465 a- nmay include Intel® Architecture (IA) cores. Each of the kernels 465 a- nmay communicate with coherency logic 455 via connections 487 a- n. In some embodiments, connections 487 a- nmay support an in-device protocol (e.g., IDI). In various embodiments, the host processor may include means 470 operable to communicate with bus logic 460 over a link 488. In some embodiments, device 470 may include an I-O device such as a PCIe-IO device.In some embodiments, the apparatus 405 is operable to perform a coherency bias process applicable to various configurations such as a system including an accelerator 410 and a host processor 445 (e.g., a computing processing complex including one or more computing processor chips), wherein the accelerator 410 is communicatively coupled to the host processor 445 via a multi-protocol link 489, and wherein memory is directly coupled to the accelerator 410 and the host processor 445 (e.g., accelerator memory 430 and host memory 440, respectively). The coherency bias process provided by device 405 may provide several technological advantages over conventional systems, such as for both accelerator 410 and "host" software executing on processing cores 465 a- n, accessing accelerator memory 430. The coherence bias process provided by the apparatus may include a host bias process and a device bias process (collectively: bias protocol flows), and a plurality of options for modulating and / or selecting bias protocol flows for specific memory accesses.In some embodiments, the bias protocol flows may be implemented at least in part using protocol layers (e.g., "bias protocol layers") at the multi-protocol link 489. In some embodiments, the bias protocol layers may include an in-device protocol (e.g., IDI) and / or a storage connection protocol (e.g., SMI3). In some embodiments, the bias protocol flows may be enabled using various information of the bias protocol layers, adding new information to the bias protocol layers, and / or adding support for protocols. For example, the bias protocol flows may be implemented using existing opcodes for an intra-device protocol (e.g., IDI), adding opcodes to a storage link protocol standard (e.g., SMI3), and / or adding support for a storage link protocol (e.g., SMI3) at the multi-protocol link 489 (e.g., conventional multi-protocol links may include only an intra-device protocol (e.g., IDI) and a fabric-based protocol (e.g., IOSF)).In some embodiments, device 405 may be associated with at least one operating system (OS). The OS may be configured not to use the accelerator memory 430 or not to use certain portions thereof. Such an OS may have support for "memory-only NUMA modules" (e.g., no CPU). The device 405 may execute a driver (e.g., including the driver 118) to perform various accelerator storage services. Illustrative and non-limiting driver implemented accelerator storage services may include driver discovery and / or recording / allocating accelerator storage 430, providing allocation APIs and mapping pages via the OS pages mapping service, providing processes for managing multi-process storage congestion and scheduling, providing APIs to enable, as software applications, the bias mode of storage regions of the accelerator storage 430 to be established and changed, and / or enabling APIs that return pages to the driver's free page list and / or return pages to a default bias mode.FIG. 5 a illustrates an example of an operating environment 500, which may be representative of various embodiments. The operating environment 500 illustrated in FIG. 5 amay provide a host bias process flow, in accordance with some embodiments. As shown in FIG. 5 a, a device 505 may include a CPU 510 communicatively coupled to an accelerator 520 via a link 540. In some embodiments, link 540 may include a multi-protocol link. The CPU 510 may include coherency controllers 530, and may be communicatively coupled to host memory 512. In various embodiments, coherency controllers 530 may be operable to provide one or more standard cache coherency protocols. In some embodiments, coherence controllers 530 may include and / or be associated with various agents, such as a home agent. In some embodiments, the CPU 510 may include and / or be communicatively coupled to one or more IO devices. The accelerator 520 may be communicatively coupled to accelerator memory 522.Host bias process flows 550 and 560 may include a set of data flows that lock all requests to accelerator memory 522 through coherency controllers 530 in CPU 510, including requests from accelerator 520. In this manner, accelerator 522 takes a cumbersome way to access accelerator memory 522, but allows accesses from both accelerator 522 and CPU 510 (including requests from IO devices via CPU 510) to be maintained as coherent using standard cache coherency protocols of coherency controllers 530. In some embodiments, host bias process flows 550 and 560 may use an in-device protocol (e.g., IDI). In some embodiments, host bias process flows 550 and 560 may use default opcodes of an intra-device protocol (e.g., IDI) to output requests to coherency controllers 530, for example, via multi-protocol link 540. In various embodiments, coherency controllers 530 may issue various coherency messages (e.g., snoops) resulting from requests from accelerator 520 to all peer processor chips and internal processor agents for accelerator 520. In some embodiments, the various coherency messages may include point-to-point protocol coherency messages (e.g., UPI) and / or intra-device protocol messages (e.g., IDI).In some embodiments, coherence controllers 530 may conditionally issue memory access messages to an accelerator memory controller (not shown) of accelerator 520 via multi-protocol link 540. Such memory access messages may be the same as or substantially similar to memory access messages that coherency controllers 530 may send to CPU memory controllers (not shown), and may have new opcodes that allow data to be passed directly back to an internal agent of accelerator 520 instead of forcing data to be passed back to coherency controllers, and then passed back again to accelerator 520 via multi-protocol link 540 in response to an in-device protocol (e.g., IDI).The host bias process flow 550 may include a flow resulting from a request or store operation for the accelerator memory 522 coming from the accelerator. The host bias process path 560 may include a flow resulting from a request or memory operation for the accelerator memory 522 coming from the CPU 510 (or from an IO device or software application associated with the CPU 510). When device 505 is active in a host bias mode, host bias process flows 550 and 560 may be used to access accelerator memory 522 as shown in FIG. 5 a. In various embodiments, in host bias mode, all requests from CPU 510 that target accelerator memory 522 may be sent directly to coherency controllers 530. Coherency controllers 530 may apply standard cache coherency protocols and send standard cache coherency messages. In some embodiments, coherence controllers 530 may send memory link protocol (e.g., SMI3) commands over multi-protocol link 540 for such requests, where memory link protocol (e.g., SMI3) flows return data over multi-protocol link 540.FIG. 5 b illustrates an example of an operating environment 500, which may be representative of various embodiments. The operating environment 500 illustrated in FIG. 5 amay provide a host bias process flow, in accordance with some embodiments. When device 505 is active in a device bias mode, a device bias path 570 may be used to access accelerator memory 522 as shown in FIG. 5. For example, device bias flow or path 570 may enable accelerator 520 to directly access accelerator memory 522 without retrieval from coherency controllers 530. In particular, a device bias path 570 may allow the accelerator 520 to directly access accelerator memory 522 without having to send a request over the multi-protocol link 540.In device bias mode, requests from CPU 510 for accelerator memory may be issued in the same or substantially similar manner as described for host bias mode, in accordance with some embodiments, but differ in the portion of the memory interconnect protocol (e.g., SMI3) of path 580. In some embodiments, requests from CPU 510 in device bias mode may be completed as if issued as "uncachted" requests. Generally, uncachted request data is not cached in the cache hierarchy of the CPU during device bias mode. In this manner, for example, accelerator 520 is allowed to access data in accelerator memory 522 during device bias mode without retrieval from coherency controllers 530 of CPU 510. In some embodiments, unclocked requests may be implemented on the device internal protocol bus (e.g., IDI) of the CPU 510. In various embodiments, uncachted requests may be implemented using a globally maintained, once used (GO-UO) protocol on the device internal protocol bus (e.g., IDI) of the CPU 510. For example, a response to an uncachted request may return a piece of data to CPU 510 and instruct CPU 510 to use that piece of data only once, for example, to prevent caching of the piece of data and to support the use of an uncachted data flow.In some embodiments, device 505 and / or CPU 510 may not support GO-UO. In such embodiments, unclocked flows (e.g., path 580) may be implemented using multiple message response sequences on a storage link protocol (e.g., SMI3) of the multiple protocol link 540 and the device internal protocol bus (e.g., IDI) of the CPU 510. For example, if the CPU 510 targets a "device bias" side of the accelerator 520, the accelerator 520 may establish one or more states to block future requests to the target memory region (e.g., a cache line) from the accelerator 520 and send a "device bias hit" response to the device link protocol line (e.g., SMI3) of the multi-protocol link 540. In response to the "device bias hit" message, coherency controller 530 (or agents thereof) may return data to a requesting processor core immediately followed by a snoop invalidation message. When the associated processor core acknowledges that snoop invalidation is complete, coherence controller 530 (or agents thereof) may send a device bias block completed message to accelerator 520 on the memory link protocol line (e.g., SMI3) of multi-protocol link 540. In response to receiving the device bias block completed message, the accelerator may clear the associated blocking state.Referring to FIG. 4, the bias mode information 427 may include a bias indicator configured to indicate an active bias mode (e.g., the device bias mode or the host bias mode). The selection of the active bias mode may be determined by the bias information 432. In some embodiments, bias information 432 may include a bias table. In various embodiments, the bias table may include bias information 432 for certain regions of accelerator memory, such as pages, rows, and / or the like. In some embodiments, the bias table may include bits per memory page of the accelerator memory 430 (e.g., 1 or 3 bits). In some embodiments, the bias table may be implemented using RAM, such as SRAM at accelerator 410 and / or an opposing portion of accelerator memory 430, with or without caching within accelerator 410.In some embodiments, bias information 432 may include bias table entries in the bias table. In various embodiments, the bias table entry associated with each access to accelerator memory 430 may be accessed prior to actual access to accelerator memory 430. In some embodiments, local requests from accelerator 410 finding their page in device bias may be forwarded directly to accelerator memory 430. In various embodiments, local requests from accelerator 410 that find their page in host bias may be forwarded to host processor 445, for example, as an in-device protocol request (e.g., IDI) at multi-protocol link 489. In some embodiments, host processor requests 445, for example using the memory interconnect protocol (e.g., SMI3), finding their page in device bias may complete the request using an uncachted flow (e.g., path 580 of FIG. 5 b). In some embodiments, host processor requests 445 may complete the request as a default memory read from accelerator memory (e.g., via path 560 of FIG. 5 a), for example using the memory link protocol (e.g., SMI3) that finds its page in host bias.The bias mode of a bias indicator of the bias mode information 427 of a region of the accelerator memory 430 (e.g., a memory page) may be changed via a software-based system, a hardware-assisted system, a hardware-based system, or a combination thereof. In some embodiments, the bias indicator may be changed via a programming interface (API) call (e.g., OpenCL), which in turn may call the device driver (e.g., driver 118) of accelerator 410. The accelerator 410 device driver may send a message to the accelerator 410 (or set a command descriptor) to instruct the accelerator 410 to change the bias indicator. In some embodiments, a change in the bias indicator may be accompanied by a cache flush operation in the host processor 445. In various embodiments, a cache flush operation may be required for a transition from host bias mode to device bias mode, but may not be required for a transition from device bias mode to host bias mode. In various embodiments, software may change a bias mode of one or more memory regions of accelerator memory 430 via a work request sent to accelerator 430.In certain cases, the software may not be able or readily able to determine when to make a bias transition API call and identify memory regions that require a bias transition. In such cases, the accelerator 410 may provide a bias transition notification process, where the accelerator 410 determines a need for bias transition and sends a message to the accelerator driver (e.g., driver 118) indicating the need for bias transition. In various embodiments, the bias transition hint process may be activated in response to a bias table search that triggers accesses of the accelerator 410 to host bias memory regions or accesses of the host processor 445 to device bias mode memory regions. In some embodiments, the bias transition notification process may signal the need for a bias transition to the accelerator driver via an interrupt. In various embodiments, the bias table may include a bias state bit for enabling the bias transition state values. The bias state bit may be used to enable access to memory regions during the process of bias change (e.g., when caches are partially flushed and incremental cache loading must be suppressed by subsequent requests).Included herein are one or more logic flows representative of example methodologies in the practice of novel aspects of the disclosed architecture. While the one or more methodologies shown herein are shown and described as a sequence of acts for purposes of simplicity of explanation, those skilled in the art will recognize and understand that the methodologies are not limited by the order of acts. Some operations may accordingly occur in a different order and / or simultaneously with operations other than those shown and written herein. For example, those skilled in the art understand and recognize that methodology may alternatively be represented as a series of contiguous states or events as in a state diagram. Moreover, not all of the operations illustrated in one methodology may be required for a novel implementation.Logic flow may be implemented in hardware, firmware, software, or a combination thereof. In software and firmware embodiments, logic flow may be implemented by computer-executable instructions stored on a non-transitory computer-readable medium or machine-readable medium such as optical, magnetic, or semiconductor memory. The embodiments are not limited in this respect.FIG. 6 illustrates one embodiment of a logic flow 600. Logic flow 600 may be representative of some or all of the operations performed by one or more embodiments described herein as in devices 105, 305, 405, and 505. In some embodiments, logic flow 600 may be representative of some or all of the operations for a coherency bias process, in accordance with some embodiments.As shown in FIG. 6, logic flow 600 may set a bias mode for accelerator memory pages to the host bias mode at block 602. For example, a host software application (e.g., software application 140) may set the bias mode of accelerator device memory 430 to the host bias mode via a driver and / or an API call. The host software application may use an API call (e.g., an OpenCL API) to transition associated (or target) pages of accelerator memory 430 that stores operands to host bias. Because the allocated pages transition from device bias mode to host bias mode, no cache flush is initiated. The device bias mode may be indicated in a bias table of bias information 432.At block 604, the logic may move from 600 operands and / or data to accelerator memory pages. For example, accelerator 420 may perform a function for the CPU that requires certain operands. The host software application may move operands from a peer CPU core (e.g., core 465a) to associated pages of accelerator memory 430. Host processor 445 may generate operand data in associated pages in accelerator memory 430 (and in arbitrary positions in host memory 440).Logic flow 600 may move accelerator memory pages to device bias mode at block 606. For example, the host software application may use an API call to place operand memory pages of the accelerator memory 430 in device bias mode. When the device bias transition is complete, the host software application may assign operations to the accelerator 430. Accelerator 430 may perform the function associated with the allocated work without host-related coherence overhead.Logic flow 600 may generate results using operands via the accelerator and store the results in accelerator memory pages at block 608. For example, the accelerator 420 may perform a function (e.g., a floating point operation, a graphics calculation, an FFT operation, and / or the like) using operands to generate results. The results may be stored in the accelerator memory 430. In addition, the software application may use an API call to cause a work descriptor to be input to flush operand pages from the host cache. In some embodiments, cache flush may be performed using a cache (or cache line) flush routine (such as CLFLUSH) on an in-device protocol (e.g., the IDI protocol). The results generated by the function may be stored in associated pages of the accelerator memory 430.The logic flow may set the bias mode for accelerator memory pages to the host bias mode at block 610. For example, the host software application may use an API call to place operand memory pages of the accelerator memory 430 in the host bias mode without causing coherency processes and / or flush operations. The host CPU 445 may access results, cache, and share them. At block 612, the logic flow 600 may provide results from accelerator memory pages to host software. For example, the host software application may access the results directly from accelerator memory pages 430. In some embodiments, associated accelerator memory pages may be enabled by the logic flow. For example, the host software application may use a driver and / or API call to release the associated memory pages of the accelerator memory 430.FIG. 7 is a block diagram illustrating a structure according to one or more examples of the present specification. In this case, a coherent accelerator structure 700 is provided. The coherent accelerator structure 700 connects to an IAL endpoint 728 that communicatively couples the coherent accelerator structure 700 to a host device such as the host devices disclosed in the above FIGURES.The coherent accelerator structure 700 is provided to communicatively couple the accelerator 740 and its associated memory 722 to the host device. The memory 722 includes a plurality of memory controllers 720- 1 to 720- n. In one example, 8 memory controllers 720 may service 8 separate memory banks.Fabric controller 736 includes a set of controllers and connections for providing coherent memory fabric 700. In this example, texture controller 736 is divided into n separate slices to service the n memory banks of memory 722. Each slice may be substantially independent of any other slice. As discussed above, fabric controller 736 includes both "vertical" links 706 and "horizontal" links 708. Vertical connections are generally to be understood as connecting devices connected upstream or downstream from one another. For example, a last level cache (LLC) 734 vertically connects to the LLC controller 738 and a fabric thereon to the die internal connection block (F2IDI), thereby communicatively coupling the fabric controller 736 to the accelerator 740. F2IDI 730 provides a downstream link to structure stop 712 and may also provide a bypass link 715. Bypass link 715 connects an LLC controller 738 directly to a fabric to memory link 716 where signals are collectively output to a memory controller 720. On the unbridged path, requests pass from the F2IDI 730 along the horizontal link to the host and then back to the structure stop 712, then to a structure coherency engine 704, and then to F2MEM 716.Horizontal buses include buses that connect fabric stops 712 together and connect the LLC controllers together.In an example, the IAL endpoint 728 may receive from the host device a packet having an instruction to perform an accelerated function along with payload data including snoops on which the accelerator is to operate. The IAL endpoint 728 passes it to L2FAB 718, which acts as a host device link for the fabric controller 736. L2FAB 718 may act as a link controller of the structure, including providing the IAL interface controller (although in some embodiments additional IAL control elements may also be provided and generally any combination of elements providing IAL interface control may be referred to as an "IAL interface controller"). L2NB 718 controls requests from the accelerator to the host and vice versa. L2FAB 718 may also become an IDI agent and may be forced to act as a sort agent between the IDI requests from accelerators and snoops from the host.L2FAB 718 may then operate structure stop 712- 0 to fill memory 722 with the values. The structure stop L2FAB 718 may apply a load balancing algorithm, such as a simple address-based hash, to label payload data for particular target memory banks. When memory banks in memory 722 have been filled with the corresponding data, accelerator 740 operates texture controller 736 to fetch values from memory into LLC 734 via LLC controller 738. The accelerator 740 performs its accelerated computation, then writes outputs to the LLC 734, where they are then forwarded in the downstream direction and output to the memory 722.Fabric stops 712, F2MEM controllers 716, multiplexers 710, and F2IDIs 730 may all be standard buses and connections that provide connectivity according to already known principles, in some examples. The foregoing connections may provide virtual and physical channels, connections, buses, switching elements, and scheduling mechanisms. They may also provide a conflict resolution mechanism with respect to interactions between requests issued by the accelerator or device agent and requests issued by the host. The structure may have physical buses in the horizontal direction, with ring-like switching operations of the server as the buses traverse the different slices. The structure may also include a special optimized horizontal connection 739 between LLC controllers 738.Requests from the F2IDI 730 may go through the hardware to share and bundle traffic to the host between the horizontal fabric link and the slice optimized paths between the LLC controller 738 and the memory 722. This involves multiplexing traffic and directing it either to an IDI block where it crosses the traditional path over the structure stop 712 and FC 704, or the preliminary use of IDI to route traffic to transition links 715. F2IDI 730- 1 may also include hardware to manage input and output to and from the horizontal fabric connection, such as by providing appropriate signals to fabric stops 712.The IAL interface controller 718 may be a PCIe controller if appropriate. The IAL interface controller provides the interface between the packetized IAL bus and the fabric interconnect. It is responsible for setting and providing flow control for IAL messages and controlling IAL messages to the respective physical and virtual fabric channels. L2NB 718 may also provide arbitration between multiple classes of IAL messages. It may also enforce IAL sorting rules.At least three control structures in the structure controller 736 provide novel and advantageous features of the structure controller 736 of the present specification. These include LLC controllers 738, FCEs 704, and power management module 750.Advantageously, the LLC controllers 738 may also provide bias control functions according to the IAL bias protocol. Thus, the LLC controllers 738 may include hardware for performing cache searches, hardware for checking the IAL base for a cache miss request, hardware for controlling requests on the designated interconnect path, and logic for responding to snoops issued by the host processor or by an FCE 704.In controlling requests from the fabric stop 712 to the host via L2FAB 718, the LLC controller 738 determines where traffic is to be routed via the fabric stop 712, via the bypass link 715 directly to the F2MEM 716, or via the horizontal bus 739 to a different memory controller.Note that in some embodiments, the LLC controller 738 is a device or block physically separate from the FCE 704. It is possible to provide a single blog that provides functions to both the LLC controller 738 and the FCE 704. However, slicing the two blocks and providing the IAL bias logic in the LLC controller 738 is possible to provide a bypass link 715 and thus speed up certain memory operations. Advantageously, in some embodiments, disconnecting the LLC controller 738 and the FCE 704 may also support selective power routing in portions of the fabric for more efficient use of resources.The FCE 704 may include hardware for setting, processing (e.g., issuing snoops to the LLC), and tracking SMI requests from the host. This provides coherence with the host device. The FCE 704 may also include hardware for setting requests on the per slice optimized path to a memory bank in memory 722. Embodiments of an FCE may also include hardware for agitating and multiplexing the two above-mentioned request classes onto a CMI memory subsystem interface, and may include hardware or logic for resolving conflicts between the two above-mentioned request classes. Other embodiments of an FCE may provide support for sorting requests from a direct vertical link and requests from the FCE 704.The power management module (PMM) 750 also provides advantages for embodiments of the present specification. For example, consider the case where each independent slice in the texture controller 736 vertically supports 1 GB per second bandwidth. 1 GB per second is provided merely as an illustrative example, and examples of texture controller 736 under real conditions may be either substantially faster or substantially slower than 1 GB per second.The LLC 734 may have substantially higher bandwidths, for example 10 times the bandwidth of the vertical bandwidth of a slice of the structure controller 736. Thus, the LLC 734 may have a bandwidth of 10 GB per second, which may be bi-directional, through which LLC 734 is 20 GB per second. Thus, for 8 slices of fabric controller 736 that each support 20 GB per second bi-directional, accelerator 740 may recognize a total bandwidth of 160 GB per second over horizontal bus 739. Operation of the LLC controllers 738 and the horizontal bus 739 at full speed thus consumes large amounts of current.However, as mentioned above, the vertical bandwidth may be 1 GB per slice and the total IAL bandwidth may be about 10 GB per second. Thus, horizontal bus 739 provides a bandwidth that is approximately an order of magnitude over the bandwidth of overall fabric controller 736. For example, horizontal bus 739 may comprise thousands of physical cables, while vertical connections may comprise hundreds of physical cables. The horizontal structure 708 may support the entire bandwidth of the IAL, i.e., 10 GB second in each direction at a total of 20 GB per second.The accelerator 740 may perform computations and operate LLCs 734 at substantially higher speeds than the host device may be able to consume data. Thus, data may be clustered into the accelerator 740 and may be subsequently consumed by the host processor as needed. After the accelerator 740 completes its computations and enters the corresponding values into LLCs 734, maintaining full bandwidth between LLC controllers 738 consumes a large amount of power that is substantially wasted because the LLC controllers 738 no longer need to communicate with each other while the accelerator 740 is inactive. Thus, while the accelerator 740 is idle, the LLC controllers 738 may be disabled such that the horizontal bus 739 is disabled, while required vertical buses, such as from the fabric stop 712 to the FCE 704 to the F2MEM 716 to the memory controller 720, remain enabled and the horizontal bus 708 is also maintained. Because the horizontal bus 739 operates at about an order of magnitude or more over the rest of the structure 700, this may save about an order of magnitude of power while the accelerator 740 is inactive.Note that some embodiments of the coherent accelerator structure 700 may also provide isochronous controllers that may be used to provide isochronous traffic to elements that are delay prone or time sensitive. For example, if the accelerator 740 is a display accelerator, then an isochronous display path may be provided to a display generator (DG) such that a connected display receives isochronous data.The entire combination of agents and connections in coherent accelerator structure 700 implements IAL functions such that high performance is present without stalling and voids. This occurs while maintaining energy and providing increased efficiency by bypassing links 715.FIG. 8 is a flow diagram of a method 800 according to one or more examples of the present specification. Method 800 illustrates a method for saving energy, as may be provided by PMM 750 of FIG. 7.Input from the host device 804 may reach the coherent accelerator structure, including an instruction to perform a calculation and payload for the calculation. If the horizontal connections between LLC controllers are disconnected, the PMM turned on the connection to its full bandwidth in block 808.In block 812, the accelerator calculates results corresponding to the normal function. During the calculation of these results, it can operate the coherent accelerator structure with its full available bandwidth, including the full bandwidth of the horizontal connections between LLC controllers.If the results are present, the accelerator structure may offload the results to the local memory 820 in block 816.In decision block 824, the PMM determines whether there is new available data from the host that can be operated upon. If new data is available, control then returns to block 812 and the accelerator continues to perform its accelerated function. Meanwhile, the host device may consume data directly from local memory 820, which may be coherently mapped to the host memory address space.Returning to block 824, if no new data is available from the host, then in block 828, the PMM reduces power so that the LLC controllers are shut down and thus the high bandwidth horizontal link between the LLC controllers is disabled. As described above, since local memory 820 maps to the host address memory space, the host may continue to consume data from local memory 820 at full IAL bandwidth, which in some embodiments is substantially less than the full bandwidth between the LLC controllers.In block 832, control waits for a new input from the host device, and if new data is received, the connection may then be re-enabled.FIGS. 9-11 illustrate an example of IAL.mem tunneling over PCIe. The described packet formats include standard PCIe packet fields except gray highlighted fields. Gray fields are those that provide the new tunnelling fields.FIG. 9 is a block diagram of an IAL.mem read operation under PCIe operation, in accordance with one or more examples of the present specification. New fields include:• MemOpcode (4 bits) - Memory opcode. It contains information as to which memory transaction is to be processed. For example, reads, writes, zeroes, etc.• MetaField and MetaValue (2 bits) - Metadata field and Metadata value. Together, these indicate which metadata field in the memory is to be modified and to which value. A metadata field in memory typically contains information associated with the actual data. For example, QPI stores directory states in metadata.• TC (2 bits) - Traffic class. Used for differentiating traffic belonging to different quality of service classes.• Snp type (3 bits) - snoop type. Used to maintain coherency between the host and device caches.• R (5 bits) - ReservedFIG. 10 is a block diagram of a write operation of an IAL.mem under PCIe operation, in accordance with one or more examples of the present specification. New fields include:• MemOpcode (4 bits) - Memory opcode. It contains information as to which memory transaction is to be processed. For example, reads, writes, zeroes, etc.• MetaField and MetaValue (2 bits) - Metadata field and Metadata value. Together, these indicate which metadata field in the memory is to be modified and to which value. A metadata field in memory typically contains information associated with the actual data. For example, QPI stores directory states in metadata.• TC (2 bits) - Traffic class. Used for differentiating traffic belonging to different quality of service classes.• Snp type (3 bits) - snoop type. Used to maintain coherency between the host and device caches.• R (5 bits) - ReservedFIG. 11 is a block diagram of an IAL.mem termination with data under PCIe operation, in accordance with one or more examples of the present specification. New fields include:• R (1 bit) - Reserved• opcode (3 bits) - IAL.io opcode• MetaField and MetaValue (2 bits) - Metadata field and Metadata value. Together, these indicate which metadata field in the memory is to be modified and to which value. A metadata field in memory typically contains information associated with the actual data. For example, QPI stores directory states in metadata.• PCLS (4 bits) - Previous cache line state. Used for critical coherence transitions.• PRE (7 bits) - performance coding. Used by performance monitor meters in the host.FIG. 12 illustrates an embodiment of a structure composed of point-to-point connections connecting a group of components, according to one or more examples of the present specification. System 1200 includes processor 1205 and system memory 1210 coupled to controller hub 1215. The processor 1205 includes any processing element, such as a microprocessor, a host processor, an embedded processor, a co-processor, or other processor. Processor 1205 is coupled to controller hub 1215 through a front side bus (FSB) 1206. In one embodiment, an FSB 1206 is a serial point-to-point connection, as described below. In another embodiment, link 1206 comprises a serial, differential interconnect architecture that conforms to differential interconnect standards.System memory 1210 includes any storage device, such as random access memory (RAM), non-volatile memory (NV), or other memory that can be accessed by devices in system 1200. System memory 1210 is coupled to a controller hub 1215 through memory interface 1216. Examples of a memory interface include a double data rate (DDR) memory interface, a dual channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.In one embodiment, the controller hub 1215 is a root hub, root complex, or root controller in a peripheral component interconnect express (PCIe or PCIE) connection hierarchy. Examples of the control hub 1215 include a chipset, a memory control hub (MCH), a north bridge, an interconnect controller hub (ICH), a south bridge, and a root controller / hub. Often, the term chipset refers to two physically separate control hubs, i.e., a memory controller hub (MCH) coupled to an interconnect controller hub (ICH).It should be appreciated that current systems often include the MCH integrated with the processor 1205, while the controller 1215 is to communicate with I / O devices in a similar manner as described below. In some embodiments, peer-to-peer routing is optionally supported by root complex 1215.Here, the control hub 1215 is coupled to a switch / bridge 1220 through the serial link 1219. Input / output modules 1217 and 1221, which may also be referred to as interfaces / ports 1217 and 1221, include / implement a layered protocol stack to provide communication between the controller hub 1215 and the switch 1220. In one embodiment, multiple devices may be coupled to switch 1220.Switch / bridge 1220 routes packets / messages from device 1225 upstream, i.e., up a hierarchy towards a root complex, to controller hub 1215 and downstream, i.e., down a hierarchy away from a root controller, from processor 1205 or system memory 1210 to device 1225. Switch 1220 is referred to in one embodiment as a logical array of multiple PCI to PCI virtual bridge devices.The device 1225 includes any internal or external device or component to be coupled to an electronic system, such as an I / O device, a network interface (NIC), an expansion card, an audio processor, a network processor, a hard disk, a storage device, a CD / DVD-ROM, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a Firewire device, a Universal Serial Bus (USB) device, a scanner, and other input / output devices. Often, such devices are referred to as an endpoint in PCIe usage. Although not specifically shown, device 1225 may include a PCIe-to-PCI / PCI-X bridge to support legacy PCI devices or other versions of PCI devices. Endpoint devices in PCIe are often classified as adopted endpoints, PCIe, or integrated root complex endpoints.Accelerator 1230 is also coupled to control hub 1215 via serial link 1232. In one embodiment, graphics accelerator 1230 is coupled to an MCH, which is coupled to an ICH. Switch 1220 and accordingly I / O device 1225 are then coupled to the ICH. I / O modules 1231 and 1218 are also intended to implement a layered protocol stack to communicate between graphics accelerator 1230 and control hub 1215. Similar to the MCH discussion above, a graphics controller or graphics accelerator 1230 itself may be integrated into processor 1205.In some embodiments, accelerator 1230 may be an accelerator, such as accelerator 740 of FIG. 7, that provides coherent memory with processor 1205.To support IAL over PCIe, the controller hub 1215 (or other PCIe controller) may have extensions to the PCIe protocol, including, by way of non-limiting example, a mapping engine 1240, a tunneling engine 1242, a host bias-to-device bias flip engine 1244, and a QoS engine 1246.The mapping engine 1240 may be configured to provide opcode mapping between PCIe instructions and IAL.io (IOSF) opcodes. IOSF provides a non-coherent ordered semantic protocol and may provide services such as device discovery, device configuration, error message, interrupt provision, interrupt handling, and DMA-like data transfers, by way of non-limiting example. Native PCIe may provide corresponding instructions, such that the mapping may be a one-to-one mapping in some cases.Tunneling engine 1242 provides IAL.mem (SMI) tunneling over PCIe. This tunneling allows the host (e.g., the processor) to map accelerator memory into the host memory address space and read and read in a coherent manner into and out of the accelerator memory. SMI is a pipelined memory interface that can be used by a coherent engine on the host to tunnel IAL transactions over PCIe in a coherent manner. Examples of modified packet structures for such tunnelling are shown in Figures 9-11. In some cases, special fields for this tunneling may be allocated within one or more DVSEC fields of a PCIe packet.The host bias-to-device bias flip engine 1244 provides the accelerator device with the ability to flush host cache lines (required for host-to-device bias flip). This may be done using a zero length non-mapping write (i.e., a non-set active byte write) from the accelerator device on the PCIe at cache line granularity. Non-mapping semantics can be described using transaction and processing hints at the transaction layer packets (TLPs). For example:• TH=1, PH=01This allows the device to invalidate a given cache line so that it is allowed to access its own memory space without loss of coherency. The device may issue a read operation following the bias reversal of a page to ensure that all rows have been emptied. The device may also implement a CAM to ensure that no new requests to the line are received from the host during the execution of a flip-over.The QoS engine 1246 may subdivide IAL traffic into two or more virtual channels to optimize the connection. For example, this could include a first virtual channel (VC0) for MMIO and configuration operations, a second virtual channel (VC1) for host-to-device writes, and a third virtual channel (VC2) for host-to-device reads.FIG. 13 illustrates an example of an embodiment of a layered protocol stack according to one or more examples of the present specification. A layered protocol stack 1300 includes any form of layered communication stack, such as a quick path interconnect (QPI) stack, a PCIe stack, a next generation high performance computing interconnect stack, or another layer stack. Although the discussion is presented immediately below with reference to FIGS. 12-15 with respect to a PCIe stack, the same concepts are applicable to other interconnect stacks. In one embodiment, protocol stack 1300 is a PCIe protocol stack that includes a transaction layer 1305, a link layer 1310, and a physical layer 1320.An interface, such as interfaces 1217, 1218, 1221, 1222, 1226, and 1231 in FIG. 12, may be represented as communication protocol stack 1300. The representation as a communication protocol stack may also be referred to as a module or an interface that implements / comprises a protocol stack.PCIe uses packets to communicate information between components. Packets are formed in the transaction layer 1305 and the data link layer 1310 to carry the information from the transmitting component to the receiving component.As the transmitted packets flow through the other layers, they are enhanced with additional information necessary to handle packets at these layers. On the receiving side, the reverse process occurs and packets are converted from their physical layer 1320 representation to the data link layer 1310 representation and finally (for transaction layer packets) to the form that can be processed by the receiving device transaction layer 1305.Transaction LayerIn one embodiment, transaction layer 1305 is to provide an interface between a processing core of a device and the interconnect architecture, such as data link layer 1310 and physical layer 1320. In this regard, a primary admission to the transaction layer 1305 is the assembly and disassembly of packets (i.e., transaction layer packets (TLPs)). The transaction layer 1305 typically manages credit-based flow control for TLPs. A PCIe implements split transactions, i.e., time-separated request-response transactions, which allows a link to carry other traffic while the destination device collects data for the response.In addition, PCIe uses credit-based flow control. In this scheme, a device advertises an initial credit set for each of the receive buffers in the transaction layer 1305. An external device at the opposite end of the link, such as the control hub 115 in Figure 1, counts the number of credits consumed by each TLP. A transaction may be transmitted if the transaction does not exceed a credit limit. Upon receiving a response, a credit scope is recovered. An advantage of the credit system is that the latency of credit return does not affect performance provided that the credit limit is not reached.In one embodiment, four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more read requests and write requests for transferring data to / from a location mapped in memory. In one embodiment, memory space transactions are suitable for using two different address formats, for example, a short address format such as a 32-bit address or a long address format such as a 64-bit address. Configuration space transactions are used to access configuration space of the PCIe devices. Transactions to the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.Therefore, in one embodiment, transaction layer 1305 assembles packet header / payload 1306. The format for current packet headers / payload is found in the PCIe specification on the PCIe specification website.FIG. 14 illustrates an embodiment of a PCIe transaction descriptor according to one or more examples of the present specification. In one embodiment, transaction descriptor 1400 is a mechanism for carrying transaction information. In this regard, transaction descriptor 1400 supports the identification of transactions in a system. Other potential uses include monitoring changes in standard transaction sorting and associations of transactions with channels.Transaction descriptor 1400 includes global identifier field 1402, attribute field 1404, and channel identifier field 1406. In the illustrated example, a global identifier field 1402 is depicted, including a transaction local identifier field 1408 and a source identifier field 1410. In one embodiment, the global transaction identifier 1402 is the same for all outstanding requests.According to one implementation, the local transaction identifier field 1408 is a field generated by a requesting agent and stands for all outstanding requests that require completion for that requesting agent. Furthermore, in this example, source identifier 1410 uniquely identifies the requesting agent within a PCIe hierarchy. Thus, along with source ID 1410, local transaction identifier field 1408 provides global identification of a transaction within a hierarchy domain.The attribute field 1404 specifies characteristics and relationships of the transaction. In this regard, the attribute field 1404 is potentially used to provide additional information that allows a change in the standard processing of transactions. In one embodiment, the attribute field 1404 includes a priority field 1412, a reserved field 1414, a sort field 1416, and a no snoop field 1418. Here, priority sub-field 1412 can be changed by an initiator to assign a priority to the transaction. Reserved attribute field 1414 is reserved for future or proprietary use. Possible usage models using priority or security attributes may be implemented using the reserved attribute field.In this example, sort attribute field 1416 is used to provide optional information that mediates the sort type that can modify default sort rules. According to an exemplary implementation, a sort attribute "0" means that default sort rules are to be applied, a sort attribute "1" denotes relaxed sort, wherein writes may pass writes in the same direction and read concludes may pass writes in the same direction. Snoop attribute field 1418 is used to determine whether transactions are snooped. As depicted, the channel ID field 1406 identifies a channel with which a transaction is associated.Transmission LayerA transmission layer 1310, also referred to as a data transmission layer 1310, acts as an intermediate between the transaction layer 1305 and the physical layer 1320. In one embodiment, a responsibility of the data transfer layer 1310 is to provide a reliable mechanism for exchanging TLPs between two associated components. One side of the data transmission layer 1310 accepts TLPs assembled by the transaction layer 1305, applies a packet sequence identifier 1311, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., CRC 1312, and presents the modified TLPs to the physical layer 1320 for transmission over a physical to an external device.Physical LayerIn one embodiment, physical layer 1320 includes a logical sub-block 1321 and an electrical sub-block 1322 to physically transmit a packet to an external device. Here, the logical sub-block 1321 is responsible for the "digital" functions of the physical layer 1321. In this regard, a logical sub-block includes a transmit portion to prepare outgoing information for transmission by physical sub-block 1322, and a receiver portion to identify and prepare received information before it is passed to transmission layer 1310.The physical block 1322 includes a transmitter and a receiver. The transmitter is supplied with symbols by the logical sub-block 1321, which the transmitter serializes and transmits to an external device. The receiver is supplied with serialised symbols from an external device and converts the received signals into a bit stream. The bitstream is de-serialised and provided to logical sub-block 1321. In one embodiment, an 8b / 10b transmission code is employed in which ten-bit symbols are transmitted / received. Here, special symbols are used for framing a packet with frames 1323. Additionally, in one example, the receiver also provides a symbol clock that is recovered from the incoming serial stream.As noted above, a layered protocol stack is not limited in this respect, although transaction layer 1305, link layer 1310, and physical layer 1320 are discussed with respect to a specific embodiment of a PCIe protocol stack. Indeed, any layered protocol may be included / implemented. As an example, a port / interface represented as a layered protocol includes: (1) a first layer to assemble packets, i.e., a transaction layer; a second layer to sequence packets, i.e., a link layer; and a third layer to transmit packets, i.e., a physical layer. As a specific example, a layered common standard interface (CSI) protocol is used.FIG. 15 illustrates an embodiment of a PCIe serial point-to-point structure according to one or more examples of the present specification. Although an embodiment of a PCIe serial point-to-point link is illustrated, a serial point-to-point link is not limited in this respect because it includes any transmission path for transmitting serial data. In one embodiment, a basic PCIe link includes two differentially driven low voltage signal pairs: a transmit pair 1506 / 1511 and a receive pair 1512 / 1507. Device 1505 thus includes transmit logic 1506 to transmit data to device 1510 and receive logic 1507 to receive data from device 1510. In other words, two transmission paths, i.e., paths 1516 and 1517 and two reception paths, i.e., paths 1518 and 1519, are included in one PCIe link.A transmission path refers to any path for transmitting data, such as a transmission line, a copper line, an optical line, a wireless communication channel, an infrared communication link, or another communication path. A connection between two components, such as component 1505 and component 1510, is referred to as a link, such as link 1515. A link may support a path, each path representing a set of differential signal pairs (a pair for transmission, a pair for reception). To scale bandwidth, a link may aggregate multiple lanes, referred to as xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.A differential pair refers to two transmission paths, such as lines 1516 and 1517, to transmit differential signals. For example, when line 1516 transitions from a low voltage level to a high voltage level, i.e., at a rising edge, line 1517 controls from a high logic level to a low logic level, i.e., at a falling edge. Differential signals potentially have better electrical properties, such as better signal integrity, i.e. cross coupling, voltage overshoot / undershoot, sound, etc. This allows a better time window that allows faster transmission frequencies.This specification may provide representations in block diagram format, with certain features disclosed in separate blocks. These are to be broadly understood to disclose how various features cooperate, but are not intended to imply that the features in question necessarily need to be embodied in separate hardware or software. Further, if a single block discloses more than one feature in the same block, the respective features need not necessarily be embodied in the same hardware and / or software. For example, under certain circumstances, a computer "memory" could be distributed or mapped between multiple levels of cache or local memory, main memory, battery-backed volatile memory, and various forms of persistent storage such as a hard disk, a storage server, an optical disk, a tape drive, or the like. In certain embodiments, some of the components may be omitted or consolidated. In a general context, the arrangements depicted in the figures may be more logical in their representations, while a physical architecture may have various permutations, combinations, and / or hybrids of these elements. Numerous possible design configurations may be used to achieve the operational objectives outlined herein. Accordingly, the associated infrastructure includes a variety of replacement arrangements, concept choices, device capabilities, hardware configurations, software implementations, and equipment options.Reference may be made herein to a computer readable medium which is a tangible and non-transitory computer readable medium. As used in this specification and in the claims, a "computer readable medium" is intended to include one or more computer readable media of the same type or different types. A computer readable medium may include, by way of non-limiting example, an optical drive (e.g., CD / DVD / Blu-ray), a hard disk drive, a solid state drive, flash memory, or other non-transitory medium. A computer readable medium could also include a medium such as a read only memory (ROM), an FPGA, or an ASIC configured to execute the desired instructions, stored instructions for programming an FPGA or an ASIC to execute the desired instructions, an IP block for intellectual property that can be incorporated into hardware in other circuits, or instructions that can be directly encoded in hardware or microcode on a processor such as a microprocessor, a digital signal processor (DSP), a microcontroller, or in any other suitable component, device, element, or object, as appropriate and based on particular requirements. A non-transitory storage medium is expressly intended to include any non-transitory, special purpose, or programmable hardware configured to provide the disclosed operations or to cause a processor to perform the disclosed operations.Various elements may be "communicatively", "electrically", "mechanically", or otherwise "coupled" to one another throughout this specification and claims. Such coupling may be a direct point-to-point coupling or may include intermediate means. For example, two devices may be communicatively coupled to each other via a controller that enables communication. Devices may be electrically coupled to one another via intermediate devices such as signal amplifiers, voltage dividers or buffers. Mechanically coupled devices may be indirectly mechanically coupled.Any "module" or "engines" disclosed herein may refer to software, a software stack, a combination of hardware, firmware, and / or software, circuitry configured to perform the function of the engine or module, or any computer readable media disclosed herein. Such modules or engines may be provided in any suitable circumstance on or in connection with a hardware platform having hardware computing resources, such as a processor, memory, storage media, connections, networks and network interfaces, accelerators, or other suitable hardware. Such a hardware platform may be provided as a single monolithic device (e.g., in a PC form factor) or with a portion of the function being distributed (e.g., a "composite node" in a high performance data center where compute, storage, storage media, and other resources may be dynamically allocated and need not be local to each other).Flow diagrams, signal flow diagrams, or other diagrams may be disclosed herein that show operations performed in a particular order. Unless expressly stated otherwise or required in a particular context, the order is to be understood as a non-limiting example only. Further, in cases where one operation is shown to follow another, other intervening operations may also occur that may or may not be related thereto. Some operations may also be performed simultaneously or in parallel. In cases where an operation is indicated to "be based on" or "correspond to" another detail or operation, it is to be understood that it implies that the operation is based at least in part on or at least partially corresponds to the other detail or operation. This is not to be construed as implying that the operation is based on or only corresponds to the detail or operation, only or exclusively.All or part of any hardware elements disclosed herein may be readily provided in a system-on-a-chip (SoC), including also a central processing unit (CPU) package. An SoC represents an integrated circuit (IC) that integrates components of a computer or other electronic system into a single chip. Thus, for example, client devices or server devices may be provided wholly or partly in a SoC. The SoC may include digital, analog, mixed signal, and radio frequency functions, all of which may be provided on a single chip substrate. Other embodiments may include a multi-chip module (MCM), wherein a plurality of chips are positioned in a single electronic package and are configured to closely cooperate with each other through the electronic package.In a general context, any suitable configured circuit or processor may execute any type of instructions associated with the data to achieve the operations listed herein. Any processor disclosed herein could convert an element or object (e.g., data) from one state or ping to another state or ping. Further, the tracked, transmitted, received, or processor stored information could be provided in any databases, registers, tables, caches, queues, control lists, or memory structures based on particular requirements and implementations, all of which can be referenced in any suitable timeframe. Any of the storage or storage media elements disclosed herein are intended to be included in the broad terms "memory" and "storage media" where applicable.Computer program logic implementing all or part of the functionality described herein is embodied in various forms including, but not limited to, source code form, computer executable form, machine instructions or microcode, programmable hardware, and various intermediate forms (e.g., forms generated by an SM la, compiler, linker, or locator). In one example, source code comprises a series of computer program instructions implemented in various programming languages, such as object code, assembler language, or higher-level language such as OpenCL, FORTRAN, C, C++, JAVA, or HTML for use with various operating systems or operating environments, or in hardware description languages such as Pace, Verilog, and VHDL. The source code may define and use various data structures and communication messages. The source code may be in a computer-executable form (e.g., via an interpreter), or the source code may be converted (e.g., via a translator, assembler, or compiler) to a computer-executable form or converted to an intermediate form such as byte code. Where appropriate, any of the foregoing may be used to construct or describe suitable discrete or integrated circuits, whether sequential, combinatorial, state machines, or otherwise.In an exemplary embodiment, any number of electrical circuits of the FIGURES may be implemented on a board of an associated electronic device. The board may be a general circuit board that houses various components of the internal electronic system of the electronic device and further provides terminals to other peripherals. Any suitable processors and memories may be coupled to the board as appropriate based on particular configuration requirements, processing requirements, and computational concepts. It should be appreciated that in the numerous examples provided herein, interaction may be described using two, three, four, or more electrical components. However, this is done for purposes of clarity and example only. It will be appreciated that the system may be consolidated or re-laid in any suitable manner. Among similar concept alternatives, any of the illustrated components, modules, and elements of the FIGURES may be combined in various possible configurations, all of which are within the broad scope of this specification.Example ImplementationsIn one example, a peripheral component interconnect express (PCIe) controller is disclosed for providing coherent memory mapping between accelerator memory and host memory address space, comprising: a PCIe control hub comprising extensions for providing a coherent accelerator link (CAI) for providing bias-based coherence tracking between the accelerator memory and the host address memory space; wherein the extensions comprise: a mapping engine for providing opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions for the CAI; and a tunneling engine for providing scalable memory interconnect (SMI) tunneling of host memory operations to accelerator memory via the CAI.Further disclosed is an example wherein the opcode map is a one-to-one map.Further disclosed is an example wherein the CAI is a Intel® Accelerator Link (IAL) compliant link.Further disclosed is an example wherein the OSF instructions comprise instructions to perform an operation selected from the group consisting of device recognition, device configuration, error message, interrupt provision, interrupt handling, and direct memory access (DMA) type data transfers.Further disclosed is an example further comprising a host bias-to-device bias (HBDB) flip engine to enable the accelerator to flush a host cache line.Further disclosed is an example wherein the HBDB flip engine is to provide a transaction layer packet (TLP) indication including TH=1, PH=01.Further disclosed is an example further comprising a QoS engine comprising a plurality of virtual channels.Further disclosed is an example where the virtual channels include a VC0 for memory mapped input / output (MMIO) and configuration traffic, a VC1 for host-to-accelerator writes, and a VC2 for host-to-accelerator reads.Further disclosed is an example wherein the extensions are configured to provide an opcode for non-posted writes (NP Wr).Further disclosed is an example wherein the "NP Wr" opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.Further disclosed is an example wherein the extensions are configured to provide a packet format for reads under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein the extensions are configured to provide a packet format for writes under PCIe including a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein the extensions are configured to provide a packet format for read completion with data under PCIe, comprising a 3-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 4-bit PCLS, and a 7-bit PRE.Werner is an example of a link including the PCIe controller.Further disclosed is an example of a system comprising the compound.Further disclosed is an example of the system comprising a system-on-a-chip.Further disclosed is an example of the system comprising a multi-chip module.Further disclosed is an example of the system wherein the accelerator is a software-assisted device memory.Further disclosed is an example of the system wherein the accelerator is autonomous device storage.Further disclosed is an example of one or more tangible, non-transitory computer readable media having instructions stored thereon for providing peripheral component interconnect express (PCIe) control on a host platform for providing coherent memory mapping between accelerator memory and a host memory address space, comprising instructions for providing a PCIe control hub comprising extensions for providing a coherent accelerator link (CAI) for providing bias-based coherence tracking between the accelerator memory and the host address memory space; wherein the extensions include: a mapping engine to provide opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions to the CAI; and a tunneling engine to provide scalable memory interconnect (SMI) tunneling from host memory operations to accelerator memory via the CAI.Further disclosed is an example wherein the opcode map is a one-to-one map.Further disclosed is an example wherein the CAI is a Intel® Accelerator Link (IAL) compliant link.Further disclosed is an example wherein the OSF instructions comprise instructions to perform an operation selected from the group consisting of device recognition, device configuration, error message, interrupt provision, interrupt handling, and direct memory access (DMA) type data transfers.Further disclosed is an example wherein the instructions are further to provide a host bias-to-device bias (HBDB) flip engine to enable the accelerator to flush a host cache line.Further disclosed is an example wherein the HBDB flip engine is to provide a transaction layer packet (TLP) indication including TH=1, PH=01.Further disclosed is an example further comprising a QoS engine comprising a plurality of virtual channels.Further disclosed is an example where the virtual channels include a VC0 for memory mapped input / output (MMIO) and configuration traffic, a VC1 for host-to-accelerator writes, and a VC2 for host-to-accelerator reads.Further disclosed is an example wherein the extensions are configured to provide an opcode for non-posted writes (NP Wr).Further disclosed is an example wherein the "NP Wr" opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.Further disclosed is an example wherein the extensions are configured to provide a packet format for reads under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein the extensions are configured to provide a packet format for writes under PCIe including a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein the extensions are configured to provide a packet format for read completion with data under PCIe, comprising a 3-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 4-bit PCLS, and a 7-bit PRE.Further disclosed is an example of a computer-implemented method for providing peripheral component interconnect express (PCIe) control on a host platform for providing coherent memory mapping between accelerator memory and host memory address space, comprising: providing PCIe control hub services; providing extensions of the PCIe control hub for providing a coherent accelerator link (CAI) for providing bias-based coherency tracking between the accelerator memory and the host address memory space; wherein providing extensions comprises: providing opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions to the CAI; and providing scalable memory interconnect (SMI) tunnels of host memory operations to accelerator memory via the CAI.Further disclosed is an example wherein the opcode map is a one-to-one map.Further disclosed is an example wherein the CAI is a Intel® Accelerator Link (IAL) compliant link.Further disclosed is an example wherein the OSF instructions comprise instructions to perform an operation selected from the group consisting of device recognition, device configuration, error message, interrupt provision, interrupt handling, and direct memory access (DMA) type data transfers.Further disclosed is an example further comprising providing host bias to device bias (HBDB) flip services to enable the accelerator to flush a host cache line.Further disclosed is an example wherein providing HBDB flip services comprises providing a Transaction Layer Packet (TLP) indication comprising TH=1, PH=01.Further disclosed is an example further comprising providing QoS services comprising a plurality of virtual channels.Further disclosed is an example where the virtual channels include a VC0 for memory mapped input / output (MMIO) and configuration traffic, a VC1 for host-to-accelerator writes, and a VC2 for host-to-accelerator reads.Further disclosed is an example wherein the provision of extensions comprises providing an opcode for non-posted writes (NP Wr).Further disclosed is an example wherein the "NP Wr" opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.Further disclosed is an example wherein providing the extensions comprises providing a packet format for reads under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein providing the extensions comprises providing a packet format for writes under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.Further disclosed is an example wherein providing the extensions comprises providing a packet format for read completion with data under PCIe comprising a 3-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 4-bit PCLS, and a 7-bit PRE.Further disclosed is an example of an apparatus comprising means for performing the method.Further disclosed is an example wherein the means for performing the method comprises a processor and a memory.Further disclosed is an example, wherein the memory comprises machine readable instructions that, when executed, cause the apparatus to perform the method.Further disclosed is an example wherein the device is a computing system.Also disclosed is an example of at least one computer readable medium comprising instructions that, when executed as described in the preceding examples, implement a method or create a device.
Claims
"Peripheral Component Interconnect Express" (PCIe) control for providing coherent memory mapping between accelerator memory (522) and host memory address space (512), comprising: a PCIe control hub (1215) comprising extensions for providing a coherent accelerator link (CAI) for providing bias-based coherency tracking between the accelerator memory and the host memory address space; wherein the extensions comprise: a mapping engine (1240) for providing opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions for the CAI; and a tunneling engine (1242) to provide scalable memory interconnect (SMI) tunneling of host memory operations to accelerator memory via the CAI.The PCIe controller of claim 1, wherein the opcode map is a one-to-one map.The PCIe controller of claim 1, wherein the CAI is a Intel® Accelerator Link (IAL) compliant link.The PCIe controller of claim 1, wherein the OSF instructions comprise instructions to perform an operation selected from the group consisting of device detection, device configuration, error message, interrupt provision, interrupt handling, and direct memory access (DMA) type data transfers.The PCIe controller of claim 1, further comprising a host bias-to-device bias (HBDB) flip engine (1244) to enable the accelerator to flush a host cache line.The PCIe controller of claim 5, wherein the HBDB flip engine is to provide a transaction layer packet (TLP) indication including TH=1, PH=01.The PCIe controller of claim 1, further comprising a QoS engine (1246) comprising a plurality of virtual channels.The PCIe controller of claim 7, wherein the virtual channels comprise a memory mapped input / output (MMIO) and configuration traffic VC0, a host-to-accelerator write VC1, and a host-to-accelerator read VC2.The PCIe controller of claim 1, wherein the extensions are configured to provide an opcode for non-posted writes (NP Wr).The PCIe controller of claim 9, wherein the "NP Wr" opcode includes reserved PCIe fields Fmt[2:0]=011b and Type[4:0]=1101b.The PCIe controller of any of claims 1-10, wherein the extensions are configured to provide a packet format for reads under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.The PCIe controller of any of claims 1-10, wherein the extensions are configured to provide a packet format for writes under PCIe comprising a four-bit opcode, a 2-bit meta-field, a 2-bit meta-value, a 2-bit time code, and a 3-bit snp type.The PCIe controller of any of claims 1-10, wherein the extensions are configured to provide a packet format for read completion with data under PCIe comprising a 3-bit opcode, a 2-bit meta field, a 2-bit meta value, a 4-bit PCIS, and a 7-bit PRE.A link comprising the PCIe controller of any of claims 1-13.A system comprising the compound of claim 14.The system of claim 15, comprising a system on a chip.The system of claim 15, comprising a multi-chip module.The system of any of claims 15-17, wherein the accelerator is a software-assisted device memory.The system of any of claims 15-17, wherein the accelerator is an autonomous device memory.One or more tangible, non-transitory computer readable media having instructions stored thereon for providing peripheral component interconnect express (PCIe) control on a host platform for providing coherent memory mapping between accelerator memory (255) and host memory address space (512) comprising instructions for: providing a PCIe control hub (1215) comprising extensions for providing a coherent accelerator link (CAI) for providing bias-based coherence tracking between the accelerator memory and the host memory address space; wherein the extensions include: a mapping engine (1240) to provide opcode mapping between PCIe instructions and on-chip system fabric (OSF) instructions to the CAI; and a tunneling engine (1242) to provide scalable memory interconnect (SMI) tunneling of host memory operations to the accelerator memory via the CAI.The one or more tangible, non-transitory media of claim 20, wherein the opcode map is a one-to-one map.One or more tangible, non-transitory media according to claim 20, wherein the CAI is a Intel® Accelerator Link (IAL) compliant connection.The one or more tangible, non-transitory media of claim 20, wherein the OSF instructions comprise instructions for performing an operation selected from the group consisting of device recognition, device configuration, error message, interrupt provision, interrupt handling, and direct memory access (DMA) type data transfers.The one or more tangible, non-transitory media of claim 20, wherein the instructions are further to provide a host bias-to-device bias (HBDB) flip engine (1244) to enable the accelerator to flush a host cache line.The one or more tangible, non-transitory media of claim 24, wherein the HBDB flip engine is to provide a transaction layer packet (TLP) indication comprising TH=1, PH=01.
Citation Information
Patent Citations
Data processor
US20100153656A1
Providing Hardware Support For Shared Virtual Memory Between Local And Remote Physical Memory
US20110072234A1