Unified address space for multiple connections

DE102018127751B4Active Publication Date: 2025-10-16INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102018127751
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-12-09
Filing Date
2018-11-07
Publication Date
2025-10-16
Estimated Expiration
2038-11-07

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Device comprising: a plurality of physical connections for communicatively coupling an accelerator device (508) to a host device (504), the host device (504) executing software; and an address translation module, ATM, for providing address mapping between host physical address spaces, HPA, and guest physical address spaces, GPA, to the accelerator device (508), wherein the plurality of physical connections share a common GPA domain and wherein address mapping is associated with only one of the plurality of physical connections such that the plurality of physical connections present themselves as a single connection to the software.
Need to check novelty before this filing date? Find Prior Art

Description

Field of SpecificationThis disclosure generally addresses the field of network computing and, more particularly, though not exclusively, relates to a system and method for providing a uniform address space for multiple connections.Prior ArtIn some modern data centers, the function of a device or device may not be tied to a particular, fixed hardware configuration. Rather, processing, memory, storage, and accelerator functions may in some cases be merged from different locations into a virtual "federation node.". A current network may include a data center that houses a large number of generic hardware server devices, for example, contained in a server rack and controlled by a hypervisor. A virtual device may be executed on one or more instances of a virtual device, such as a payload server or a virtual desktop.US 2006 / 069899, A1 relates to input / output (I / O) virtualization in the field of microprocessors. To improve address translation performance, a register stores capability indicators to indicate the capability supported by a circuit in a chipset to address translate a guest physical address to a host physical address. A plurality of multi-level page tables are used for page walking in address translation. Each of the page tables has page table entries. Each of the page table entries has at least one entry identifier corresponding to the capability indicated by the capability indicators.US 2013 / 0 007 408 A1 relates to a method and an apparatus for managing a software-controlled cache for translating the physical memory access of a virtual machine between different levels of translation units. An example method allows a guest operating system to change an entry in a TLB directly without involvement of a hypervisor. Upon receipt of a guest TLB miss exception, a guest operating system issues a TLBWE (TLB Write Entry) instruction to logic. The logic executes the TLBWE command in supervisor mode without invoking a hypervisor. The TLB may include entries in a guest page table and entries in a host page table.SUMMARY OF THE INVENTIONThe solution according to the invention is defined in the apparatus according to the main claim 1 as well as the block of mental property according to the subordinate claim 14, the accelerator device according to the subordinate claim 15, the computing system according to the subordinate claim 19, the one or more tangible, non-transitory storage media according to the subordinate claim 22 and the computer-implemented method according to the subordinate claim 25.Brief Description of the DrawingsThe present disclosure will be best understood from the following detailed description taken in conjunction with the accompanying drawings. It is underlined that, according to standard practice in the industry, various features are not necessarily drawn to scale and are used for illustrative purposes only. Unless expressly or implicitly indicated to a scale, this provides only an illustrative example. In other embodiments, the dimensions of the various features may be arbitrarily increased or decreased in the interest of discussion. FIG. 1 is a block diagram of selected components of a network connectivity data center according to one or more examples of the present application. FIG. 2 is a block diagram of selected components of an end user computing device according to one or more examples of the present specification. FIG. 3 is a block diagram of a system that may reside, for example, in a data center or other computing resource, in accordance with one or more examples of the present specification. FIG. 4 is a block diagram of a second system that may also achieve advantages in accordance with the teachings of the present specification. FIG. 5 a illustrates an embodiment wherein an address translation module (ATM) is provided as an address translation and security check (ATU) unit in an accelerator according to one or more examples of the present specification. FIG. 5 b illustrates an embodiment wherein a central processing unit (CPU) operates with three accelerator devices, according to one or more examples of the present specification. FIG. 6 is a flow diagram of a method for providing address translation via an ATU in an accelerator or other suitable system configuration according to one or more examples of the present specification. FIG. 7 is a block diagram for using an address translation cache as the ATM according to one or more examples of the present specification. FIG. 8 is a flow diagram of a method that may be performed by a system where the address translation module is executed in an ATC on the accelerator device, according to one or more examples of the present specification. FIG. 9 is a block diagram of components of a computing platform in accordance with one or more examples of the present specification. FIG. 10 illustrates an embodiment of a structure composed of point-to-point connections that connect a group of components, according to one or more examples of the present specification. FIG. 11 illustrates an embodiment of a layered protocol stack according to one or more embodiments of the present specification. FIG. 12 illustrates an embodiment of a Peripheral Component Interconnect Express (PCIe) transaction descriptor according to one or more examples of the present specification. FIG. 13 illustrates an embodiment of a serial point-to-point PCIe structure according to one or more examples of the present specification. FIG. 14 illustrates an embodiment of multiple potential multi-manifold configurations according to one or more examples of the present specification. FIG. 15 illustrates an embodiment of a layered stack for a high-performance interconnect architecture according to one or more examples of the present specification.Embodiments of the DisclosureThe following disclosure provides several different embodiments or examples for implementing different features of the present disclosure. Specific examples of components and arrangements are described below for the purpose of simplifying the present disclosure. These are, of course, merely examples which are not intended to be limiting. Further, in the present disclosure, reference numerals and / or letters may be repeated in the various examples. This repetition is for purposes of simplicity and clarity and does not in itself dictate a relationship between the various embodiments and / or configurations discussed. Different embodiments may have different advantages, and in any embodiment, no particular benefit of one is necessarily required.A current computing platform, such as a hardware platform provided by Intel® or similar providers, may have a capability to monitor device performance and make decisions to provide resources. For example, in a large data center, such as may be provided by a cloud service provider (CSP), the hardware platform may include rack-located servers having computing resources such as processors, memories, storage pools, accelerators, and similar resources. As used herein, cloud computing includes network-related computing resources and technology that enables ubiquitous access to data, resources, and / or technology (often worldwide). Cloud resources are generally characterized by great flexibility for dynamically allocating resources according to current payload data and requests. This may be achieved, for example, by virtualization, where resources such as hardware, memory, and networks are provided to a virtual machine (VM) via a software abstraction layer and / or containerization, where instances of network functions are provided in "containers" that are separate from each other but share the underlying operating system, memory, and driver resources.In a modern data center, such as may be provided by a cloud service provider (CSP) or high-performance computing (HPC) cluster, computing resources including cycles of the central processing unit (CPU) may be one of the primary monetary resources. It is therefore advantageous to provide other resources that can provide offloading of special functions, thus enabling CPU cycles for utilization and increasing overall efficiency of the data center.To this end, a modern data center may provide facilities such as accelerators for offloading various functions. For example, in an HPC cluster executing an artificial intelligence task, such as a convolutional neural network (CNN), training tasks may be offloaded to an accelerator so that more cycles are available to execute the actual folds on the CPU. In a CSP, accelerators may be provided, e.g., via Intel® Accelerator Link (IAL). In other embodiments, accelerators may also be provided as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), coprocessors, digital signal processors (DSPs), graphics processing units (GPUs), or other processing entities that may optionally be tuned to provide the accelerator function. Accelerators may perform offloaded tasks that increase the efficiency of network operations or that enable other special functions. As a non-limiting example, accelerators may provide compression, decompression, encryption, decryption, deep packet inspection, security services, or any other removable functions.Because of the high data volume incurred in a data center, it is often desirable to minimize latency and maximize the bandwidth available to the CPU for communication with the accelerator. Data centers may provide certain connections known for high bandwidth and / or low latency, such as Intel® Ultra Path Interconnect (UPI), QuickPath Interconnect (QPI), Gigabit Ethernet, Infiniband, or Peripheral Component Interconnect Express (PCIe). Although many different types of connections are possible, the present specification uses UPI as a representative cache coherent connection and PCIe as a representative non-cache coherent connection. While PCIe and UPI are provided herein as specific and illustrative examples, it should be understood that they are non-limiting examples and that they may be optionally replaced with any other suitable types of connections.As accelerators and other off-site devices become more prevalent, the demand for bandwidth between the CPU and these devices increases. Further, interactions between the CPU and accelerators are also becoming increasingly mature, with a great need for low latency, high bandwidth, coherence in some cases, and other semantics.It becomes increasingly difficult to meet the overall connection demand using a single interface. For example, PCIe provides very high bandwidth, but current embodiments of PCIe do not have cache coherency. On the other hand, UPI provides cache coherency and very low latency, but less bandwidth than PCIe. It is possible to use multiple connections with different characteristics between the CPU and the accelerator to meet the need for the connection. However, this can complicate the software running on the CPU and the firmware or instructions executed on the device, as each connection requires its own drivers and supporting communication using different address spaces for each connection.For example, consider a case where a connection between a CPU and an FPGA requires:• Very high bandwidth• low latency, and• Coherence.In this case, PCIe alone may not be sufficient because PCIe is not a cache coherent link and thus does not provide coherent access to system memory for the accelerator device. Coherence may be provided by using a coherent port-to-port connection such as QPI or UPI. However, QPI and UPI may not provide the high bandwidth of PCIe, and are less efficient for volume transfer of data in many cases. To address this problem, some system designers may provide both a UPI and a PCIe link to the accelerator from the CPU.In many processors, such as the current Intel® Xeon® server CPUs, the CPU downstream port and the coherent interface (e.g., UPI) are considered part of the secure domain of the CPU complex. External devices under the CPU data stream port are protected at the root of the hierarchy in the CPU root complex. To provide isolation between memory accesses from each device, the CPU may implement address translation and security checks on the root complex. Thus, a dual-link device (e.g., with a PCIe and UPI link) connected to the CPU may be mapped to two separate routes. The device therefore appears to the operating system as two separate devices, namely a PCIe device and a UPI device. Further, memory accesses are mapped to two separate secure address domains. This means that the software developer must manage two separate devices to address a single accelerator, thereby adding software complexity for a multi-link device. Further, multiple PCIe interfaces may be provided to provide even greater bandwidth.As a non-limiting and illustrative example, this specification provides two cases, where multiple PCIe connections are provided to a single logical accelerator device.In the first example, shown in FIG. 3 below, three PCIe interfaces to a single accelerator device are provided in addition to a UPI interface. The three PCIe interfaces provide the accelerator with the ability to handle extremely high bandwidth transactions across the plurality of PCIe devices. The addition of a UPI device also provides the ability to handle cache coherent transactions with extremely low latency. However, in this example, the single accelerator illustrated in FIG. 3 is addressable as four separate devices because four discrete interfaces are provided.In a second example, shown below in FIG. 4, a plurality of devices may operate cooperatively and have device-to-device connections to communicate with one another. These devices may, for example, perform different phases of the same operation, or may perform the same operation on three separate data in parallel. The CPU may have a PCIe connection to each, but there may be logical reasons for the three devices being addressable as a single logical device.In both instances presented above, and in many other instances, it is advantageous to provide multiple physical connections to the single physical or logical device, besides connection aggregation schemes that allow multiple connections to be considered logically as a single connection in software. This provides the ability to support the benefit of a plurality of links without having to add the software complexity of separately addressing the links.Embodiments of the system and method of the present specification use an address translation and security check unit (ATU) and / or an address translation cache (ATC) to provide a mapping between a guest physical address (GPA) and a host physical address (HPA). Embodiments of the present specification separate the device's logical view from the physical connection topology. In other words, the multi-connector can be logically treated as a single connector from the physical point of view of the system and its software.One operating principle of the present specification is a "one software view" of a multi-link device such that the software on link 0 sees a single device link. When performing an allocation operation (such as "malloc"), it allocates a buffer to a particular host physical address for link 0 only. In other words, the CPU allocates a buffer at HPA-X for link 0 over malloc. The device may then send GPA-X Prime (GPA-X') over any connection to achieve the same HPA-X.In a multi-connection device as shown in Fig. 3, the CPU and device may be designed so that the software can only recognize the device on the physical hierarchy of connection 0. Note that physical connection 0 is illustrated here as a logical choice for the sole physical connection at which the device is recognizable, but this is a non-limiting example. In other embodiments, any of the plurality of connections may be selected as the discoverable connection. The other unselected physical connections are hidden from view of the software.In some embodiments of the present specification, guest-to-host physical address translation may be implemented on the device itself. This can be done by implementing the ATU in the facility upon disabling the CPU side ATU on all connections. Note that this requires the extension of the CPU root to the device. In another embodiment, guest-to-host address translation may be implemented by an ATC in the device with all ATUs on the CPU side disabled except for the selected link ATU (e.g., link 0). This link is always used to carry the host physical address.In each embodiment, the request identifier (RID) may be used by the system to identify which address domain the request is to be mapped to. In other words, each physical interface may have its own address domain, and the CPU may map, in hardware or microcode, a request to the address domain of a physical connection that the RID uses. All requests from the device may use the same RID for address translation. By using the same RID, all requests are mapped to the same address domain.This solution provides advantages with respect to other possible solutions. For example, in the case of multiple connections, I / O virtualization may be disabled so that the accelerator device cannot be virtualized. This may not be a widely acceptable solution in many data centers that rely heavily on virtualization.In other cases, virtualization may be enabled, but the device is considered to be a device with N partitions, where N is equal to the number of physical interfaces. These N partitions may need to be managed by independent software drivers or instances of software drivers. This makes it difficult for the device to perform dynamic load balancing of bandwidth across connections and further increases the software complexity for managing the end devices, rather than treating the device as a single device having a single connection. For example, a 100 GB network interface card (NIC) connected over two PCIe G3x8 connections may require execution of two independent network stacks.The system and method of the present specification can be advantageously managed as a single device, from a single software driver, but still use two or more physical connections to realize the additional bandwidth of these connections. This helps scale the performance of the device without adding additional complexity to the device and driver designers. Further, this enables portability for the application. An application written for use with a device having a single connection can be easily ported to a system that provides multiple connections, in some cases without any reprogramming. Rather, the software can dynamically adjust the number of available connections since the CPU itself manages the connections as a single logical connection.A system and method for providing a uniform address space for multiple connections is described below with particular reference to the accompanying FIGURES. Note that certain labels may be repeated throughout the FIGURES to indicate that a particular device or block is fully or partially consistent throughout the FIGURES. However, this is not intended to imply a particular relationship between the various embodiments disclosed. In certain examples, a class of elements may be denoted by a particular label ("widget 10"), while individual types or examples of the class may be denoted by a bar label label ("first specific widget 10- 1" and "second specific widget 10- 2").FIG. 1 is a block diagram of selected components of a data center having connectivity to the network 100 of a cloud service provider (CSP) 102, according to one or more examples of the present specification. It should be appreciated that embodiments of the data center disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. The CSP 102 may be, as a non-limiting example, a traditional enterprise data center or may provide a "private cloud" or a "public cloud" and services such as "infrastructure as a service" (IaaS), "platform as a service" (PaaS), or "software as a service" (SaaS). In some cases, the CSP 102 may provide high-performance computing (HPC) platforms or services instead of or in addition to the cloud services. Thus, although not expressly identical, HPC clusters ("supercomputers") may structurally resemble cloud data centers, and unless expressly stated otherwise, the teachings of this specification may be applied to both types.The CSP 102 may provide a number of payload clusters 118, which may be clusters of individual servers, blade servers, rackmount servers, or any other suitable server topology. In this illustrative example, two payload clusters 118- 1 and 118- 2 are shown, each providing rackmount servers 146 in a chassis 148.In this illustration, payload clusters 118 are shown as modular payload clusters complying with the rack unit ("U") standard, wherein a standard 19 inch wide rack is constructed to accommodate 42 units (42U) each 1.75 inch high, at about 36 inches depth. In this case, computing resources such as processors, memory, memory, accelerators, and switches may fit into several multiples of the rack units from one to 42.Each server 146 may hold a stand-alone operating system and provide a server function, or servers may be virtualized, in which case they may be controlled by a virtual machine manager (VMM), hypervisor, and / or orchestrator, and may hold one or more virtual machines, virtual servers, or virtual applications. These server racks may be located in a single data center, or they may be located in different geographic data centers. Depending on the contractual agreements, some servers 146 may be dedicated specifically to particular customers or tenants of an enterprise while others may be shared.The various devices in a data center may be interconnected via a switching fabric 170, which may include one or more high speed routing and / or switching devices. The switching fabric 170 may provide both "north-south" traffic (e.g., traffic to and from the wide area network (WAN), such as the Internet), and "east-west" traffic (e.g., traffic over the data center). Historically, north-south traffic has made the bulk of network traffic, but with more complex and distributed web services, the volume of east-west traffic has increased. In many data centers, east-west traffic is now the major part of the traffic.Further, as the capacity of each server 146 increases, the traffic volume may increase further. For example, each server 146 may provide multiple processor slots, each slot having a processor having four to eight cores, as well as sufficient memory for the cores. Thus, each server can hold a number of VMs that each generate their own traffic.To handle the large volume of traffic in a data center, a high-performance switching fabric 170 may be provided. The switching fabric 170 is shown in this example as a "flat" network, where each server 146 may have a direct connection to a top-of-rack (ToR) switch 120 (e.g., a star configuration), and each ToR switch 120 may be coupled to a core switch 130. This two-stage flat network architecture is shown only as an illustrative example. In other examples, other architectures may be employed, such as three-membered star or leafback (also referred to as "fat tree" topologies) based on the "clos" architecture, "hub and spoke" topology, mesh topologies, ring topologies, or 3-D mesh topologies as non-limiting examples.The structure itself may be provided by any suitable connection. For example, each server 146 may include a Intel® host fabric interface (HFI), a NIC, a host channel adapter (HCA) or other host interface. For convenience and convenience, throughout this specification, these may be referred to as a "host fabric interface" (HFI), which is broadly intended to be an interface for communicatively coupling the host to the data center fabric. The HFI may couple one or more host processors via a link or bus such as PCI, PCIe, or the like. In some cases, this interconnect bus may be considered part of the fabric 170, among other "local" connections (e.g., core-to-core "ultra path interconnect"). In other embodiments, the UPI (or other local coherent link) may be treated as part of the secure domain of the processor complex and thus not as part of the structure.The interconnect technology may be provided by a single interconnect or a hybrid interconnect, e.g., where PCIe provides on-chip communication, 1 GB or 10 GB copper Ethernet provides relatively short connections to a ToR switch 120, and optical cabling provides relatively longer connections to the core switch 130. Connection technologies that may be present in the data center include, by way of non-limiting examples, Intel® Omni-PathTM Architecture (OPA), TrueScaleTM, Ultra Path Interconnect (UPI) (legacy QPI or KTI), FibreChannel, Ethernet, FibreChannel over Ethernet (FCoE), InfiniBand, PCI, PCIe, or fiber optic cables, to name a few. The structure may be cache and memory coherent, non-cache and non-memory coherent, or a hybrid form of coherent and non-coherent interconnects. Some connections for particular purposes or functions are more popular than others, and selection of an appropriate structure for the particular application is made with the usual knowledge. For example, OPA and Infiniband are commonly used in high performance computing (HPC) applications, while Ethernet and FiberChannel are more popular in cloud data centers. However, these examples are expressly not limiting and as data center fabric technologies continue to develop, they will similarly progress.It should be appreciated that while sophisticated structures such as OPA are provided herein for illustrative purposes, the structure 170 may more generally be any suitable connection or bus for the particular application. This could in some cases include legacy connections such as local area networks (LAN), token ring networks, synchronous optical networks (SONET), asynchronous transfer mode (ATM) networks, wireless networks such as WiFi and Bluetooth, plain old telephone system (POTS) connections, or the like. It is also expressly anticipated that new networking technologies will come up in the future to supplement or replace some of those listed herein, and any such future networking topologies and technologies may form part of structure 170.In certain embodiments, the fabric 170 may provide communication services on various "layers" as originally outlined in the seven-layer network model open system interconnection (OSI). In actual practice, the OSI model is not strictly adhered to. In common usage, layers 1 and 2 are often referred to as an "Ethernet" layer (although Ethernet may be displaced or supplemented by newer technologies in some data centers or supercomputers). Layers 3 and 4 are often referred to as the Transmission Control Protocol / Internet Protocol (TCP / IP) layer (which can be further divided into TCP and IP layers). Layers 5- 7 may be referred to as the "application layer.". These layer definitions are disclosed as a useful framework, but are not intended to be limiting.FIG. 2 is a block diagram of selected components of an end user computing device 200 in accordance with one or more examples of the present specification. Note that embodiments of computing device 200 disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. As above, computing device 200 may provide cloud service, high-performance computing, telecommunication services, enterprise data center services, or any other computing services that utilize computing device 200, as appropriate.In this example, structure 270 is provided to connect various aspects of computing device 200. The structure 270 may be the same as the structure 170 of FIG. 1, or it may be a different structure. As above, the structure 270 may be provided by any suitable connection technology. In this example, Intel® Omni-PathTM is used as an illustrative and non-limiting example.As shown, the computing device 200 includes a number of logic elements that form a plurality of nodes. It should be appreciated that each node may be provided by a physical server, a group of servers, or other hardware. Each server may execute one or more virtual machines according to the need of its application.Node 0 208 is a processing node, includes a processor port 0 and a processor port 1. Data node 0 208 may be configured to provide network or payload functions, such as by incorporating a plurality of virtual machines or virtual applications.Onboard communication between processor port 0 and processor port 1 may be provided by an onboard uplink 278. This may provide a very high speed and short length connection between the two processor ports so that virtual machines executing on node 0 208 may communicate with each other at very high speeds. To facilitate this communication, a virtual switch (vSwitch) may be provided at node 0 208, which may be considered part of fabric 270.Node 0 208 connects to fabric 270 via an HFI 272. The HFI 272 may be connected to a Intel® Omni-PathTM structure. In some examples, communication with fabric 270 may be tunneled, such as in providing UPI tunneling via Omni-PathTM.Because the computing device 200 may provide many functions in a distributed manner that were provided on the board at previous generations, a powerful HFI 272 may be provided. The HFI 272 may operate at speeds of several gigabits per second, and in some cases may be closely coupled to the node 0 208. For example, in some embodiments, the logic for HFI 272 is directly integrated with the processors on a system-on-a-chip. This provides very high speed communication between HFI 272 and the processor ports without the need for intermediary bus devices that can introduce additional latency into the fabric. However, this is not intended to imply that embodiments are to be excluded in which HFI 272 is provided over a traditional bus. Rather, it is expressly stated that HFI 272 may be provided on a bus, such as a PCIe bus, which is a serialized version of PCI, that provides higher speeds than traditional PCI, in some examples. Throughout computing device 200, different nodes may provide different types of HFIs 272, such as onboard HFIs and plug-in HFIs. It should also be appreciated that certain blocks may be provided in a system-on-a-chip as Intellectual Property (IP) blocks, which may be incorporated into an integrated circuit as a modular unit. Thus, in some cases, the HFI 272 may be derived from such an IP block.Note that in the "Network is the device" architecture, node 0 208 may provide limited or no onboard memory or storage. Rather, node 0 208 may be primarily based on distributed services, such as a memory server and a networked storage server. On-board, node 0 208 may provide only sufficient memory and storage for boot-loading the device and establishing communication with fabric 270. This type of distributed architecture is possible due to the very high speeds of current data centers, and may be advantageous because it is not necessary to provide spare resources for each node. Instead, a large inventory of high speed or special storage may be dynamically allocated between a number of nodes so that each node has access to a large inventory of resources, but these resources do not remain inactive when the node in question does not require them.In this example, a node 1 storage server 204 and a node 2 storage server 210 provide the working memory and storage capacities of node 0 208. For example, storage server node 1 204 may provide network direct memory access (RDMA), whereby node 0 208 may access storage resources at node 1 204 via structure 270 in a direct memory access manner, similar to accessing its own onboard memory. The memory provided by the memory server 204 may be traditional memory, such as double data rate type 3 dynamic random access memory (DRAM) (DDR3), which is volatile, or it may be a more exotic memory, such as persistent fast memory (PFM) such as Intel® 3D XPoint (3DXPTM), which operates at DRAM-like speeds but is non-volatile.Similarly, instead of providing an onboard hard disk to node 0 208, storage server node 2 210 may be provided. The storage server 210 may provide a networked bundle of hard disks (NMOD), PFM, a redundant array of independent hard disks (RAID), a redundant array of independent nodes (RAIN), network-bound storage (NAS), optical storage, tape drives, or other non-volatile storage solutions.Thus, when executing its intended function, node 0 208 may access memory from storage server 204 and store results to the storage provided by storage server 210. Each of these devices is coupled to the fabric 270 via an HFI 272, thereby providing fast communication that makes these technologies possible.For further illustration, node 3 206 is also depicted. Node 3 206 also includes an HFI 272 in addition to two processor ports internally connected by an uplink. Unlike node 0 208, however, node 3 206 has its own onboard memory 222 and memory 250. Thus, the note 3 206 may be configured to perform its functions primarily on the board and it may not be necessary to build on the memory server 204 and the storage server 210. In appropriate circumstances, node 3 206 may supplement its own onboard memory 222 and distributed resource memory 250 similar to node 0 208.The computing device 200 may also include accelerators 230. These may provide various accelerated functions, including hardware or coprocessor acceleration for functions such as packet processing, encryption, decryption, compression, decompression, network security, or other accelerated functions in the data center. In some examples, accelerators 230 may include deep learning accelerators directly connected to one or more cores in nodes such as node 0 208 or node 3 206. Examples of such accelerators may include, by way of non-limiting example, Intel® QuickData Technology (QDT), Intel® QuickAs Technology (QAT), Intel® Direct Cache Access (DCA), Intel® Extended Message Signaled Interrupt (MSI-X), Intel® Receive Side Coalescing (RSC), and other acceleration technologies.The basic building block of the various components disclosed herein may be referred to as "logic elements.". Logic elements may include hardware (including, for example, a software programmable processor, an ASIC, or an FPGA), external hardware (digital, analog, or mixed signal), software, mutual software, services, drivers, interfaces, components, modules, algorithms, sensors, components, firmware, microcode, programmable logic, or objects that may be coordinated to achieve a logical operation. Further, some logic elements are provided by a tangible, non-transitory computer readable medium having executable instructions stored thereon for instructing a processor to perform a particular task. Such a non-transitory medium could include, for example, a hard disk, solid state memory or a solid state disk, read only memory (ROM), persistent fast storage (PFM) (e.g., Intel® 3DXTM), external storage, a redundant array of independent hard disks (RAID), a redundant array of independent nodes (RAIN), networked storage (NAS), optical storage, a tape drive, a backup system, cloud storage, or any combination of the foregoing, as non-limiting examples. Such a medium could also have instructions programmed into an FPGA or encoded in hardware on an ASIC or processor.FIG. 3 is a block diagram of a system 300 that may reside, for example, in a data center or other computing resource, in accordance with one or more examples of the present specification. The system 300 includes a processor 304 and an accelerator 308. Processor 304 and accelerator 308 are coupled together via a plurality of connections. For example, processor 304 in this case has a UPI connection to accelerator 308 that can provide low latency and cache coherency, such that accelerator 308 becomes possible to coherently access the memory space of processor 304, and if accelerator 308 includes its own onboard memory, that onboard memory can coherently map to the native address space of processor 304. The system 300 also provides three separate PCIe interfaces between the processor 304 and the accelerator 308, namely PCIe 0, PCIe 1 and PCIe 2. Thus, the combination of one UPI link and three PCIe links, provided merely as an illustrative and non-limiting example, advantageously provides cache coherency and low latency over the UPI link and extremely high bandwidth over the three PCIe links.However, as discussed above, in traditional addressing modes, software executing on processor 304 would need to address accelerator 308 as four separate devices to implement the advantages of the various connections. This significantly complicates the design and implementation of software on processor 304.Embodiments of I / O virtualization technology map each device to a unique address domain by creating a page table per device. For example, the page table's approach uses the {requester ID, guest physical address] to determine the host physical address. For example, the RID is on a PCIe link {Bus, Device, Function].Some existing CPUs may assume that a device on each link has a unique RID. Thus, if a request with the same GPA goes through two different connections, they make a page call with the same GPA but with a different RID. This invokes two different page tables which result in different host addresses. As part of the secure design principles of virtualization, certain existing VMMs do not allow two different RIDs to share the same page tables or have identical copies of the two page tables.Thus, while the system 300 of FIG. 3 implements increased bandwidth across the plurality of connections provided, the CPU I / O virtualization technology maps a device behind each connection as a separate device with an independent address domain. In this case, the device will eventually have four address domains. This can make it difficult to aggregate bandwidth over the three PCIe connections, in addition to the additional complexity incurred by the UPI connection.This type of configuration is usually present in existing systems, such as SKX+FPGA MCP.As discussed above, these types of implementations may use the teachings of the present specification by providing a single address domain (or optionally a single address domain per link type) such that software complexity is greatly reduced and use of the available bandwidth may be optimized.Note that in certain embodiments of the present specification, the single UPI link shown in this figure may be treated as a separate address domain from the three PCIe links provided. This may have security impact because UPI is a coherent link, and the accelerator 308 may be treated as part of the security domain or root complex of the processor 304 as long as communication is over a UPI link. On the other hand, PCIe is not a coherent link, and thus communication over the PCIe links cannot be treated as part of the root domain. Thus, in some embodiments, accelerator 308 may appear as two separate devices, a first device connected via the UPI link that is treated as part of the secure domain of the root complex and a second device connected via a single logical PCIe link (with three aggregated physical connections) that is not treated as part of the secure domain of processor 304.In other embodiments, all four connections may be treated as a single logical connection, with communication managed at a low level to implement optimized transactions. For example, low latency transactions do not require high bandwidth to be sent over the UPI link, while high bandwidth transactions may be sent over one or more of the PCIe links. In this case, the accelerator 308 may appear as a single logical device that is not in the secure domain of the processor 304 and that has a single logical connection. Those skilled in the art will appreciate that many other types of combinations are possible.FIG. 4 is a block diagram of a second system 400 that may also implement advantages in accordance with the teachings of the present specification.In this case, the CPU 404 is communicatively coupled to three devices 408- 1, 408- 2, and 4083 via three independent PCIe connections. This provides a distributed computational model in which N devices, more generally speaking, are connected to the processor 404 via N connections. The devices 408 may also communicate with each other via one or more device-to-device connections provided for the purpose of dynamically moving tasks across devices.Existing CPU I / O virtualization models require each device to operate in a different address domain. However, when each device operates in a different address domain, it is relatively difficult to migrate tasks from one device to another. Thus, the system 400 illustrated in FIG. 4 may also benefit from the ability to address devices 408- 1, 408- 2, and 408- 3 as a single logical device connected to the CPU 404 via a single logical connection including connections PCIe 0, PCIE 1, and PCIE 2.Addressing a single device having a plurality of connections or a plurality of devices each having one or more connections as a single logical device having a single logical connection may be realized via an address translation module (ATM) provided on the device itself.Figures 5a, 5b and 7 below illustrate three different embodiments of an address translation module that may be provided in accordance with the present specification. Note that the ATM of FIG. 5 and the address translation cache of FIG. 7 are provided as non-limiting examples of ATMs.An ATM as described in this specification may include hardware, microcode instructions, firmware, software, a coprocessor, intellectual property block, or other hardware platform, or part of a hardware platform configured to provide the operations of the ATM. In some embodiments, an ATM may include one or more tangible, non-transitory computer readable media having instructions stored thereon to direct a processor to provide the ATM or having instructions stored thereon to provide an ATM (e.g., register transfer language (RTL) instructions or policies or instructions or policies of another hardware description language to provide the logic of an ATM into an FPGA, an ASIC, an IP block, or other hardware or module), as well as instructions that could be executed on a programmable processor.FIG. 5 a illustrates an embodiment wherein the ATM is provided as an ATU in accelerator 508, according to one or more examples of the present specification.As shown in FIG. 5a, CPU 504 may have its own address translation unit 512 which provides three separate PCIe connections to accelerator 508. These connections may be marked according to RIDs, for example, RID 0 for PCIe 0, RID 1 for PCIe 1, and RID 2 for PCIe 2.The CPU 504, as well as the other CPUs and processors illustrated herein (e.g., in FIGS. 5 band 7 ), are provided by way of non-limiting example only. In other embodiments, any suitable host device may be employed for the illustrated CPUs.In existing systems, each RID may be associated with a separate logical device, even though the three connections are provided to the same physical accelerator 508.To enable CPU 504 to address accelerator 508 as a single logical device with a single logical PCIe connection, ATU 512 may be disabled at CPU 504 and ATU 516 may be provided at accelerator 508. The RID0 may be used for all address translations while the actual transactions are mapped to one of the physical buses via the RIDx (where "x" is the physical bus identifier for handling the transaction).As used in this specification, an ATU may be an address translation and security verifier, which in some embodiments may be compatible with the specification of Intel® directed I / O (VT-d) virtualization technology. One purpose of the ATU is to provide address space isolation and force approval checks for I / O devices. This may involve a state machine that traverses OS-managed page tables and caches to store intermediate translations. The ATU 516 may provide HPA-to-GPA and / or "guest virtual address" (GVA) mapping. A guest physical address may be a guest operating system or virtual machine physical address. Note that a GPA may not be a real physical address as recognized by the machine and actual DRAMs. The host physical address is a physical address from the perspective of the VMM or machine. This may be the address used to access the physical DRAMs.By placing the ATU 516 at the accelerator 508 and disabling the ATU 512 at the CPU 504, the accelerator device logical view may be separated from the physical topology. The multiple connections to accelerator 508 may be logically treated as a single logical connection from the perspective of software executing on CPU 504. Software executing on CPU 504 views a single device at PCIe 0 and reserves a buffer over malloc at HPA-X only for PCIe 0. The accelerator 508 may send a GPA-X' over any of the PCIe connections to achieve the same HPA-X.In this embodiment, each link (e.g., RID0, RID1, RID2) may traverse its own memory pages.Figure 5b illustrates a similar embodiment. However, in the example of FIG. 5 b, the CPU 554 operates with three accelerator devices 558- 1, 558- 2, 558- 3. While it is possible that all address translations could be performed on a single ATU (e.g., positioned on device 1 558-1), availability of a plurality of devices makes it possible to distribute the address translation across multiple devices using a distributed ATU (DATU). For example, as shown herein, device 1 558- 1 includes DATU 560- 1, device 2 558- 2 includes DATU 560- 2, and device 3 558- 3 includes DATU 560- 3. Load balancing address translation between the DATUs may be accomplished by any suitable method, including a method in which each device 558 handles the address translation for its own transactions and may perform its own page pass. Note that the CPU 554 does not need to know the distributed nature of the DATU 560. Rather, the three devices of the CPU appear as a single logical device with a single logical ATU. Note that in FIG. 7, an address translation cache (ATC) is used instead of an ATC. The dio Di DiFIG. 6 is a flow diagram of a method 600 for providing address translation via an ATU 516 of an accelerator 508 or in any other suitable system configuration according to one or more examples of the present specification. In method 600, at block 604, the system disables the ATU, such as the ATU 512 of CPU 504 implemented for the plurality of connections to the accelerator device. This substantially disables the virtualization capabilities provided by the CPU for these connections, such as the three PCIe connections shown in FIG. 5 a. Note, however, that virtualization technology is a system-wide capability, which means that it is not possible in some cases to have a system in which some devices are virtualized and other devices are not. This is provided merely as a non-limiting example.In block 608, the system provides and / or activates an ATU at the accelerator device, such as the ATU 516 of the accelerator 508. This ATU covers one of the connections, such as the PCIe 0 connection. However, any of the connections could be selected, such as PCIe 1 or PCIe 2.Note that in some cases, the system may assume that the ATU is integrated into the CPU. Moving the ATU to the accelerator device, such as device 508 of FIG. 5, may require some changes to the device enumeration procedure, enabling access to additional memory mapped I / O base address registers (MMIO BARs) for the VT-d unit. These reminders may depend on the CPU architecture for some platforms, and the PCIe endpoint may need to be declared as a root complex integrated endpoint (RCiEP). It is anticipated that making such changes to the CPU architecture would be within the skill of those skilled in the art.In block 612, the system exposes the device ATU to the VMM or OS on a single line such as PCIe 0.At block 616, a request 620 is received from the device such as the accelerator device 508. In some embodiments, all requests from the accelerator device are examined at the accelerator device ATU using the RID of link 0 to obtain the host physical address. Such aggregated queries are shown at device 508 of FIG. 5.In block 624, the ATU at the accelerator receives the host physical address.At block 628, the request for address translation is sent over the link with the host physical address and one RID per link. For example, a request over PCIe could have 0 {RID0, HPA}. A request over PCIe 1 could have {RID1,HPA}. A request over PCIe 2 could include {RID2,HPA}.At block 632, error detection or routing assumptions may be provided and incorporated into the RID in some embodiments on the CPU side. This replaces the single RID with connection specific unique RIDs before issuing the request to the CPU. This is acceptable because connection virtualization capabilities on the CPU side may be disabled.FIG. 7 is a block diagram illustrating use of an address translation cache as the ATM according to one or more examples of the present specification.As used in this specification, an address translation cache (ATC) is a translation cache that can cache addresses translated by the address translation and security checking unit (ATU). The ATC may not be able to perform page passes. If a request fails the ATC, the ATC is coordinated with the ATU to perform a page pass. This may include, for example, address translation services as defined in the PCIe specification.In the embodiment of FIG. 7, the CPU 704 is provided with the ATU 712, while the accelerator 708 is provided with the ATC 716. Note that unlike the ATU 512 of the CPU 504 of FIG. 5 a, the ATU 712 of the CPU 704 is not disabled. Rather, ATU 712 may enable address translation services only on a single link such as PCIe 0. Thus, as in FIG. 5a, GPA-to-HPA translation is always via PCIe 0 (with RID0). Accelerator 708 then provides the HPA to CPU 704 via PCIe 0 with the correct RIDx, and the actual transaction is performed on PCIe x.As before, the ATC 716 aggregates requests into {RID0, HPA}. However, the actual communication may be performed with a specific RID. For example, for PCIe 0, {RID0, HPA}. For PCIe 1, {RID1, HPA}. For PCIe 2, {RID2, HPA}. In this embodiment, only the link associated with address translation (e.g., RID0) may traverse the address page.In the above FIGURES, the address translation modules (ATMs), i.e., ATU 516 (FIG. 5 a), DATU 560 (FIG. 5 b), and ATC 716 (FIG. 7 ), are shown as part of the accelerator 508. However, other embodiments include an ATM that is not integrated into the accelerator device. In these embodiments, the ATM may include separate elements such as ASIC, FPGA, IP block, circuitry, programmable logic, ROM, or other devices that provide the address translation services between the host device and the accelerator device.In some embodiments, as shown in any of the preceding figures, an ATM may also provide nested translation, GPA-to-GVA (guest virtual address).FIG. 8 is a flow diagram of a method 800 that may be performed, for example, by a system such as that disclosed in FIG. 7, wherein the address translation module is executed in an ATC at the accelerator device, according to one or more examples of the present specification.At block 801, the CPU may implement address translation services to support an ATC at the accelerator device.In block 808, the system may deactivate the CPU ATU anywhere except for a connection. For example, in the embodiment of FIG. 7, the ATU may be disabled anywhere except for the PCIe 0 link. PCIe 0 is chosen here as an illustrative example, but it should be understood that the selection of an appropriate link for activation is within the ordinary skill in the art.In block 820, the system implements an ATC on the accelerator device that can be used to translate requests over all three connections. Each miss of the ATC may be handled via an address translation service request, for example, to PCIe 0.In block 824, a request 816 is received from the device. All requests from the device are looked up at the ATC using the RID of link 0 to obtain the host physical address.At block 828, the request for address translation may be sent over the connection to the HPA and one RID per connection. For example, requests over PCIe use 0 {RID0, HPA}. Requests over PCIe 1 use {RID1, HPA}. Requests over PCIe 2 use {RID2, HPA}.On the CPU side of the link, error detection or routing assumptions established around the RID may be present. This replaces the single RID with connection specific unique RIDs before issuing the request to the CPU. This is allowed because the virtualization capabilities on connections 1 or 2 on the CPU side may be disabled.Note that in certain embodiments, the same concepts for implementing device-side ATU or ATC may be extended to shared virtual memory flows as may be used in a distributed system as shown in FIG. 4.FIG. 9 is a block diagram of components of a computing platform 902A, in accordance with one or more examples of the present specification. Note that embodiments of computing device 902A disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. In the illustrated embodiment, platforms 902A, 902B, and 902C are coupled to data center management platform 906 and data analytics engine 904 via network 908. In other embodiments, a computing system may include any suitable number of platforms (i.e., one or more). In some embodiments (e.g., when a computing system has only a single platform), all or a portion of system management platform 906 may be included on platform 902. Platform 902 may include platform logic 910 including one or more central processing units (CPUs) 912, memories 914 (which may include any number of different modules), chipsets 916, communication interfaces 918, and any other suitable hardware and / or software for executing hypervisor 920 or other operating system capable of executing operations associated with applications executing on platform 902. In some embodiments, a platform 902 may function as a host platform for one or more guest systems 922 that invoke these applications. Platform 902A may represent any suitable computing environment, such as a high-performance computing environment, a data center, a communication service provider infrastructure (e.g., one or more portions of an evolved packet core), an in-memory computing environment, a computing system of a vehicle (e.g., an automobile or aircraft), an Internet of Things environment, an industrial control system, another computing environment, or a combination thereof.In various embodiments of the present disclosure, cumulative loading and / or cumulative loading rates of a plurality of hardware resources (e.g., cores and uncores) may be monitored and entities (e.g., system management platform 906, hypervisor 920, or other operating system) of computing platform 902A may allocate hardware resources to platform logic 910 to perform operations corresponding to the loading information. In some embodiments, even diagnostic capabilities may be combined with stress monitoring to more accurately determine the state of the hardware resources. Each platform 902 may include platform logic 910. Platform logic 910 includes, among other logic that enables the functionality of platform 902, one or more CPUs 912, memory 914, one or more chipsets 916, and communication interfaces 928. Although three platforms are shown, the computing platform 902A may be connected to any suitable number of platforms. In various embodiments, platform 902 may be on a board installed in a chassis, rack, or other suitable structure that includes multiple platforms coupled together by network 908 (which may include, e.g., a rack or backplane switch).The CPUs 912 may each include any suitable number of processor cores and supporting logic (e.g., uncores). The cores may be coupled to each other, to memory 914, to at least one chipset 916, and / or to a communication interface 918 through one or more controllers resident on the CPU 912 and / or chipset 916. In particular embodiments, a CPU 912 is embodied within a port that is permanently or removably coupled to platform 902A. Although four CPUs are shown, a platform 902 may include any suitable number of CPUs.The memory 914 may include any form of volatile or non-volatile memory including, but not limited to, magnetic media (e.g., one or more tape drives), optical media, random access memory (RAM), read only memory (ROM), flash memory, removable media, or any other suitable local or remote storage components or components. The memory 914 may be used for short, medium, and / or long term storage by the platform 902A. The memory 914 may store any suitable data or information used by the platform logic 910, including software embedded in a computer readable medium and / or encoded logic (e.g., firmware) integrated in hardware or otherwise stored. The memory 914 may store data used by cores of the CPUs 912. In some embodiments, memory 914 may also include memory space for instructions executed by cores of CPUs 912 or other processing elements (e.g., logic present on chipset 916) to provide functionality associated with manageability engine 926 or other components of platform logic 910. A platform 902 may also include one or more chipsets 916 including any suitable logic to support the operation of the CPUs 912. In various embodiments, chipset 916 may be on the same die or package as a CPU 912 or on one or more different dies or packages. Each chipset may support any suitable number of CPUs 912. A chipset 916 may also include one or more controllers to couple other components of platform logic 910 (e.g., communication interface 918 or memory 914) to one or more CPUs. In the illustrated embodiment, each chipset 916 also includes a manageability engine 926. The manageability engine 926 may include any suitable logic to support the operation of the chipset 916. In a particular embodiment, a manageability engine 926 (which may also be referred to as an innovation engine) is capable of capturing real-time telemetry data from chipset 916, CPU(s) 912, and / or memory 914 managed by chipset 916, other components of platform logic 910, and / or various connections between components of platform logic 910. In various embodiments, the detected telemetry data includes the loading information described herein.In various embodiments, a manageability engine 926 operates as an out-of-band asynchronous compute agent capable of interfacing with the various elements of platform logic 910 to acquire telemetry data on CPUs 912 without or with minimal disruption to running processes. For example, the manageability engine 926 may include a dedicated processing element (e.g., a processor, controller, or other logic) on the chipset 916 that provides functionality to the manageability engine 926 (e.g., by executing software instructions) such that processing cycles of the CPUs 912 for operations performed by the platform logic 910 associated with the operations are maintained. In addition, dedicated logic for the manageability engine 926 may operate asynchronously with respect to the CPUs 912, and may capture at least some of the telemetry data without increasing the load on the CPUs.A manageability engine 926 may process telemetry data it captures (specific examples of processing stress information are provided herein). In various embodiments, the manageability engine 926 reports the data it captures and / or the results of its processing to other elements in the computing system, such as one or more hypervisors 920 or other operating systems and / or system management software (which may be executed on any suitable logic such as the system management platform 906). In particular embodiments, a critical event such as a core that has accumulated an excessive amount of load may be reported prior to the normal interval for telemetry data reporting (e.g., a notification may be sent immediately upon detection).Additionally, the manageability engine 926 may include programmable code configurable to specify which CPU(s) 912 manages a particular chipset 916 and / or which telemetry data is captured.The chipsets 916 further each include a communication interface 928. Communication interface 928 may be used to communicate signaling and / or data between chipset 916 and one or more I / O devices, one or more networks 908, and / or one or more devices coupled to network 908 (e.g., system management platform 906). For example, communication interface 928 may be used to send and receive network traffic such as data packets. In a particular embodiment, a communication interface 928 includes one or more physical network interface controllers (NICs), also known as network interface cards or network adapters. A NIC may include electronic circuitry to communicate using any suitable physical layers and data link layer standards such as Ethernet (e.g., defined by an IEEE 802.3 standard), Fiber Channel, InfiniBand, Wi-Fi, or other suitable standard. A NIC may include one or more physical ports that may be coupled to a cable (e.g., an Ethernet cable). A NIC may enable communication between any suitable elements of chipset 916 (e.g., manageability engine 926 or switch 930) and another device coupled to network 908. In various embodiments, a NIC may be integrated with the chipset (i.e., it may be on the same integrated circuit or board as the rest of chipset logic), or it may be on a different integrated circuit or board electromechanically coupled to the chipset.In particular embodiments, communication interface 928 may enable data communication (e.g., between manageability engine 926 and data center management platform 906) in conjunction with management and monitoring functions performed by manageability engine 926. In various embodiments, the manageability engine 926 may use elements of the communication interfaces 928 (e.g., one or more NICs) to report the telemetry data (e.g., to the system management platform 906) to reserve the use of NICs of the communication interface 918 for operations associated with operations performed by the platform logic 910.Switches 930 may be coupled to various ports (e.g., provided by NICs) of communication interface 928, and may switch data between these ports and various components of chipset 916 (e.g., one or more Peripheral Component Interconnect Express (PCIe) lanes coupled to CPUs 912). Switches 930 may be a physical or virtual (i.e., software) switch fabric.Platform logic 910 may include an additional communication interface 918. Similar to communication interfaces 928, communication interfaces 918 may be used to communicate signaling and / or data between platform logic 910 and one or more networks 908 and one or more devices coupled to network 908. For example, communication interface 918 may be used to send and receive network traffic such as data packets. In a particular embodiment, communication interfaces 918 include one or more physical NICs. These NICs may enable communication between any suitable elements of platform logic 910 (e.g., CPUs 912 or memory 914) and another device coupled to network 908 (e.g., elements of other platforms or remote computing devices coupled to network 908 through one or more networks).Platform logic 910 may receive and perform any suitable types of operations. An operation may include any request to use one or more resources of platform logic 910, such as one or more cores or associated logic. For example, an operation may include a request to instantiate a software component such as an I / O device driver 924 or guest system 922; a request to process a network packet received from a virtual machine 932 or device external to platform 902A (such as a network node coupled to network 908); a request to execute a process or thread associated with a guest system 922, an application executing on platform 902A, a hypervisor 920, or other operating system executing on platform 902A; or other suitable processing request.A virtual machine 932 may emulate a computer system with its own dedicated hardware. A virtual machine 932 may execute a guest operating system over the hypervisor 920. The components of platform logic 910 (e.g., CPUs 912, memory 914, chipset 916, and communication interfaces 918) may be virtualized to appear to the guest operating system that virtual machine 932 has its own dedicated components.A virtual machine 932 may include a virtualized NIC (vNIC) used by the virtual machine as its network interface. A vNIC may be associated with a media access control (MAC) address or other identifier, so that it becomes possible to individually address multiple virtual machines 932 in a network.The VNF 934 may include a software implementation of a functional building block with defined interfaces and behaviors that may be implemented in a virtualized infrastructure. In particular embodiments, a VNF 934 may include one or more virtual machines 932 that collectively provide specific scopes of functions (e.g., wide area network (WAN) optimization, virtual private network (VPN) termination, firewall operations, load balancing operations, security functions, etc.). A VNF 934 executing on platform logic 910 may provide the same functionality as traditional network components implemented by dedicated hardware. For example, a VNF 934 may include components for performing any suitable NFV operations, such as virtualized evolved packet core (vEP) components, mobility management entities, 3rd generation partnership project (3GPP) control and data plane components, etc.The SFC 936 is a group of VNFs 934 organized as a chain for performing a series of operations such as network packet processing operations. Service function chaining may provide the capability to provide an ordered list of network services (e.g., firewalls, load balancers) that are merged in the network to generate a service chain.A hypervisor 920 (also known as a virtual machine monitor) may include logic to generate and execute guest systems 922. Hypervisor 920 may include guest operating systems executing by virtual machines with a virtual operating platform (i.e., from the virtual machine perspective, they are executing on separate physical nodes while actually consolidated on a single hardware platform) and manage execution of the guest operating systems by platform logic 910. The services of hypervisor 920 may be provided by virtualization in software, or by hardware-assisted resources that require minimal software intervention, or by both. Multiple instances of a plurality of guest operating systems may be managed by hypervisor 920. Each platform 902 may have a separate instantiation of a hypervisor 920.The hypervisor 920 may be a native or bare metal hypervisor running directly on the platform logic 910 to control the platform logic and manage the guest operating systems. Alternatively, hypervisor 920 may be a legacy hypervisor executing on a host operating system and abstracts the guest operating systems from the host operating system. Supervisor 920 may include a virtual switch 938 that may provide virtual switching and / or routing functions to virtual machines of guest systems 922. The virtual switch 938 may include a logical switching fabric that couples the vNICs of the virtual machines 932 together and thus creates a virtual network over which virtual machines may communicate with each other.The virtual switch 938 may include a software element that is executed using components of the platform logic 910. In various embodiments, hypervisor 920 may be in communication with any suitable entity (e.g., an SDN controller) that causes hypervisor 920 to re-establish the parameters of virtual switch 938 in response to changing conditions in platform 902 (e.g., adding or deleting virtual machines 932 or specifying optimizations that may be made to improve the performance of the platform).Hypervisor 920 may also include resource allocation logic 944, which may include logic to determine the allocation of platform resources based on telemetry data (which may include load information). Resource allocation logic 944 may also include logic for communicating with various components of platform logic 910 entities of platform 902A to implement such optimization as platform logic 910 components.Any suitable logic may make one or more of these optimization decisions. For example, system management platform 906, resource allocation logic 944 of hypervisor 920, or other operating system or logic of computing platform 902A may be capable of making such decisions. In various embodiments, system management platform 906 may receive telemetry data via multiple platforms 902 and manage the placement of operations thereon. The system management platform 906 may communicate with hypervisors 920 (e.g., out of band) or other operating systems of the various platforms 902 to implement placements of operations targeted to the system management platform.The elements of platform logic 910 may be coupled together in any suitable manner. For example, a bus may couple any components together. A bus may include any known connection, such as a multi-drop bus, a mesh connection, a ring connection, a point-to-point connection, a serial connection, a parallel bus, a coherent (i.e., cache coherent) bus, a layered protocol architecture, a differential bus, or a Gunning Transceiver Logic (GTL) bus.Elements of computing platform 902A may be coupled to each other in any suitable manner, such as via one or more networks 908. A network 908 may be any suitable network or combination of one or more networks operating using one or more suitable network protocols. A network may represent a number of nodes, points, and connected communication paths for receiving and transmitting information packets propagating over a communication system. For example, a network may include one or more firewalls, routers, switches, security applications, anti-virus servers, or other suitable network devices.FIG. 10 illustrates an embodiment of a structure composed of point-to-point connections that connect a group of components, according to one or more examples of the present specification. It should be appreciated that embodiments of the structure having a uniform address space for multiple connections disclosed in this figure may be provided and utilize as described in the present specification. System 1000 includes processor 1005 and system memory 1010 controlled by controller hub 1015. Processor 1005 includes any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or other processor. Processor 1005 is coupled to controller hub 1015 via front side bus (FSB) 1006. In one embodiment, the FSB 1006 is a serial point-to-point-to-point connection as described below. In another embodiment, link 1006 comprises a serial, differential link architecture that complies with differential link standards.System memory 1010 includes any storage device such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible by devices in system 1000. System memory 1010 is coupled to controller hub 1015 through memory interface 1016. Examples of a memory interface include a double data rate (DDR) memory interface, a dual channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.In one embodiment, the controller hub 1015 is a root hub, a root complex, or a root controller in a peripheral component interconnect express (PCIe) connection hierarchy. Examples of the control hub 1015 include a chipset, a memory control hub (MCH), a north bridge, an interconnect control hub (ICH), a south bridge, and a root controller / hub. The term chipset often denotes two physically separate control hubs, i.e. a storage control hub (MCH) coupled to an interconnect control hub (ICH). Note that current systems often have the MCH integrated with processor 1005, while controller 1015 is to communicate with I / O devices in a similar manner as described below. In some embodiments, peer-to-peer routing is optionally supported by root complex 1015.Here, the control hub 1015 is coupled to the switch / bridge 1020 via the serial link 1019. Input / output modules 1017 and 1021, which may also be referred to as interfaces / ports 1017 and 1021, include / implement a layered protocol stack for providing communication between the control hub 1015 and the switch 1020. In an embodiment, the plurality of devices may be coupled to the switch 1020.Switch / bridge 1020 routes packets / messages from device 1025 up, i.e., up in a hierarchy to a root complex to controller hub 1015, and down, i.e., down in a hierarchy from a root controller, processor 1005, or system memory 1010 to device 1025. The switch 1020, in one embodiment, is referred to as a logical array of multiple PCI-to-PCI virtual bridge devices. The device 1025 includes internal or external devices or components for coupling to an electronic system such as an I / O device, a network interface controller (NIC), an add-on card, an audio processor, a network processor, a hard disk, a storage device, a CD / DVD-ROM, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a firewire device, a universal serial bus (USB) device, a scanner, and other input / output devices. In PCIe usage, such a device is often referred to as an endpoint. Although not specifically shown, device 1025 may include a PCIe-to-PCI / PCI-X bridge to support legacy devices or other versions of PCI devices. Endpoint devices in PCIe are often classified as legacy, PCIe, or root complex integrated endpoints.Graphics accelerator 1030 is also coupled to controller hub 1015 through serial link 1032. In one embodiment, graphics accelerator 1030 is coupled to an MCH, which is coupled to an ICH. The switch 1020, and accordingly the I / O device 1025, is then coupled to the ICH. I / O modules 1031 and 1018 are also intended to implement a layered protocol stack to communicate between graphics accelerator 1030 and control hub 1015. Similar to the MCH discussion above, a graphics controller or graphics accelerator 1030 itself may be integrated into processor 1005.FIG. 11 illustrates an embodiment of a layered protocol stack according to one or more embodiments of the present specification. Note that embodiments of the layered protocol stack disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. Layered protocol stack 1100 includes any form of layered communication stack, such as a quick path interconnect (QPI) stack, a PCIe stack, a next generation high performance computing interconnect stack, or another layered stack. Although the discussion immediately below with reference to FIGS. 10-13 is presented with reference to a PCIe stack, the same concepts may be applied to other interconnect stacks. In one embodiment, protocol stack 1100 is a PCIe protocol stack including a transaction layer 1105, a link layer 1110, and a physical layer 1120. An interface such as interfaces 1017, 1018, 1021, 1022, 1026, and 1031 in FIG. 1 may be represented as communication protocol stack 1100. The representation as a communication protocol stack can also relate to a module or an interface which implement / have a protocol stack.PCIe uses packets to communicate information between components. Packets are formed in the transaction layer 1105 and data link layer 1110 to bring the information from the sending component to the receiving component. As the transmitted packets flow through the other layers, they are enhanced with additional information required to handle the packets in these layers. At the receiving side, the reverse process occurs and packets are converted from the playback of their physical layer 1120 to the playback of data link layer 1110 and finally (for transaction layer packets) to the form that can be processed by transaction layer 1105 at the receiving side.Transaction LayerIn one embodiment, transaction layer 1105 is to provide an interface between a processing core of a device and the connection architecture such as data link layer 1110 and physical layer 1120. In this regard, a first responsibility of the transaction layer 1105 is the assembly and disassembly of packets, i.e., transaction layer packets (TLPs). The transaction layer 1105 typically manages credit-based flow control for TLPs. PCIe implements shared transactions, i.e., transactions in which the request and response are separated in time, so that a link can carry more traffic while the target device acquires data for the response.In addition, PCIe uses credit-based flow control. In this scheme, a device reports an initial credit amount for each of the receive buffers in the transaction layer 1105. An external device at the opposite end of the link, such as controller hub 115 in Figure 1, counts the number of credits consumed by each TLP. A transaction may be transmitted if the transaction does not exceed a credit limit. Upon receipt of a response, a credit amount is recovered. An advantage of a credit scheme is that latency of credit feedback does not affect performance unless the credit limit is reached.In one embodiment, four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more read requests and write requests to transfer data to / from a location mapped to memory. In one embodiment, memory space transactions may use two different address formats, e.g., a short address format such as a 32-bit address or a long address format such as a 64-bit address. Configuration space transactions are used to access configuration space of the PCIe devices. Transactions to the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.Thus, in one embodiment, transaction layer 1105 assembles packet header / payload 1106. The format for current packet header / payload is found in the PCIe specification at the PCIe specification website.FIG. 12 illustrates an embodiment of a PCIe transaction descriptor according to one or more examples of the present specification. Note that embodiments of the PCIe transaction descriptor disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. In one embodiment, transaction descriptor 1200 is a mechanism for carrying transaction information. In this regard, transaction descriptor 1200 supports the identification of transactions in a system. Other possible applications include tracking modifications to the default transaction order and assigning transactions to channels.Transaction descriptor 1200 includes global identifier field 1202, attribute field 1204, and channel identifier field 1206. In the illustrated example, global identifier field 1202 is shown to include local transaction identifier field 1208 and source identifier field 1210. In one embodiment, the global transaction identifier 1202 is unique to all outstanding requests.According to one implementation, the local transaction identifier field 1208 is a field generated by a requesting agent and is unique to all outstanding requests that require completion for the requesting agent in question. Further, in this example, the source identifier 1210 uniquely identifies the requesting agent in a PCIe hierarchy. Accordingly, the local transaction identifier 1208, along with the source ID 1210, provides global identification of a transaction in a hierarchy domain.The attribute field 1204 indicates properties and relationships of the transaction. In this regard, the attribute field 1204 is potentially used to provide additional information that enables modifying the standard handling of transactions. In one embodiment, the attribute field 1204 includes a priority field 1212, a reserved field 1214, an ordering field 1216, and a no snoop field 1218. Here, priority sub-field 1212 may be modified by an initiator to assign a priority to the transaction. Reserved attribute field 1214 is left reserved for future or provider-defined use. Possible usage models using priority or security attributes may be implemented using the reserved attribute field.In this example, the ordering attribute field 1216 is used to provide optional information indicating the type of ordering that the default ordering rules can modify. According to an example implementation, an order attribute "0" denotes default order rules to apply, while an order attribute "1" denotes relaxed order, wherein writes may forward writes in the same direction and read concludes may forward writes in the same direction. Snoop attribute field 1218 is used to determine whether transactions are undergoing snoop checking. As depicted, the channel ID field 1206 identifies a channel with which a transaction is associated.Bonding LayerThe link layer 1110, also referred to as data link layer 1110, acts as an intermediate between the transaction layer 1105 and the physical layer 1120. In one embodiment, a responsibility of data link layer 1110 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two connected components. A data link layer 1110 side accepts TLPs composed by the transaction layer 1105, applies the packet sequence identifier 1111, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., CRC 1112, and directs the modified TLPs to the physical layer 1120 for physical transmission to an external device.Physical LayerIn one embodiment, physical layer 1120 includes a logical sub-block 1121 and an electrical sub-block 1122 to physically transmit a packet to an external device. Here, the logical sub-block 1121 is responsible for the "digital" functions of the physical layer 1121. In this regard, the logical sub-block includes a transmitting portion to prepare outgoing information for transmission through the physical sub-block 1122, and a receiving portion to identify and prepare received information before it is forwarded to the link layer 1110.The physical block 1122 includes a transmitter and a receiver. The transmitter receives symbols from logical sub-block 1121 that the transmitter serializes and transmits to an external device. The receiver receives serialised symbols from an external device and converts the received signals into a bit stream. The bit stream is deserialized and passed to logical sub-block 1121. In one embodiment, an 8b / 10b transmission code is employed, wherein 10-bit symbols are transmitted / received. Here, special symbols are used to frame a packet with frames 1123. Additionally, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.As noted above, although transaction layer 1105, link layer 1110, and physical layer 1120 are discussed with respect to a specific embodiment of a PCIe protocol stack, a layered protocol stack is not so limited. Indeed, any layered protocol may be included / implemented. For example, a port / interface, represented as a layered protocol, includes: (1) a first layer for assembling packets, i.e., a transaction layer; a second layer for tracking packets, i.e., a link layer; and a third layer for sending the packets, i.e., a physical layer. As a specific example, a common standard interface (CSI) layered protocol is used.FIG. 13 illustrates an embodiment of a serial point-to-point PCIe structure according to one or more examples of the present specification. Note that embodiments of the serial point-to-point PCIe structure disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. Although an embodiment of a PCIe serial point-to-point link is illustrated, a serial point-to-point link is not so limited as it has any transmission path for transmitting serial data. In the illustrated embodiment, a base PCIe link includes two differentially driven low voltage signal pairs: a transmit pair 1306 / 1311 and a receive pair 1312 / 1307. Accordingly, device 1305 includes transmit logic 1306 to transmit data to device 1310 and receive logic 1307 to receive data from device 1310. In other words, two transmission paths, i.e., paths 1316 and 1317 and two reception paths, i.e., paths 1318 and 1319, are included in a PCIe link.A transmission path denotes any path for transmitting data, such as a transmission line, a copper line, an optical line, a wireless communication channel, an infrared communication channel, or another communication path. A connection between two devices such as device 1305 and device 1310 is referred to as a link such as link 1315. A link may support a path - each path representing a group of differential signal pairs (a pair for transmission, a pair for reception). To scale bandwidth, a link may aggregate multiple lanes denoted by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.A differentiated pair refers to two transmission paths, such as lines 1316 and 1317, for transmitting differential signals. For example, when line 1316 switches from a low voltage level to a high voltage level, i.e., a rising edge, line 1317 controls from a high voltage level to a low voltage level, i.e., a falling edge. Differential signals potentially have better electrical properties, such as better signal integrity, i.e. cross coupling, voltage overshoot / undershoot, ringing, etc. This allows for a better time window that allows for faster transmission frequencies.In one embodiment, a new high power connection (HPI) is provided. HPI is a next generation cache coherent link-based link. As an example, HPI may be used on high performance computing platforms such as workstations or servers, where PCIe is commonly used to connect accelerators or I / O devices. However, HPI is not limited thereto. Rather, HPI may be deployed in any of the systems or platforms described herein. Further, the individual ideas developed can be applied to other connections such as PCIe. In addition, HPI can be extended to compete on the same market as other connections, e.g., PCIe. To support multiple devices, in one implementation, HPI has an agnostic instruction set architecture (ISA) (i.e., HPI may be implemented in multiple different devices). In another scenario, HPI may also be used to connect high performance I / O devices and not just processors or accelerators. For example, a high-performance PCIe device may be coupled to HPI via a corresponding translation bridge (i.e., HPI to PCIe). In addition, the HPI connections may be used in various ways (e.g., stars, rings, meshes) in many HPI-based devices.FIG. 14 illustrates one embodiment of multiple potential multi-port configurations. Note that embodiments of the multi-port configurations disclosed in this figure may be provided with and utilize a uniform address space for multiple connections as described in the present specification. A two port configuration 1405 as shown has two HPI links; however, in other implementations, one HPI link may be used. For larger topologies, any configurations may be used as long as an ID is allocable and a form of virtual path exists. As shown, 4-port configuration 1410 includes an HPI link among all processors. However, in the 8-port implementation shown in configuration 1415, not all ports are directly interconnected by an HPI link. However, if there is a virtual path between the processors, the configuration is supported. A range of supported processors includes 2-32 in a native domain. Higher numbers of processors may be achieved through the use of multiple domains or other connections between node controllers.The HPI architecture has a definition of a layered protocol architecture that is similar to PCIe because it also has a layered protocol architecture. In one embodiment, HPI defines protocol layers (coherent, non-coherent, and optionally other memory-based protocols), a routing layer, a link layer, and a physical layer. Further, as with many other connection architectures, HPI has improvements in performance manager, design for test and debugging (DFT), fault handling, registry, security, etc.FIG. 15 A illustrates an embodiment of potential layers in the layered HPI protocol stack; however, these layers are not required and may be optional in some implementations. Note that embodiments of potential layers disclosed in this figure may be provided in and utilize layered HPI protocol stack with a uniform address space for multiple connections as described in the present specification. Each layer handles its own granularity level and information scope (the protocol layer 1505a,b with packets 1530, the link layer 1510a,b with flits 1535, and the physical layer 1505a,b with phits 1540). Note that in some embodiments, a packet may include partial flits, a single flit, or multiple flits based on the implementation.As a first example, a width of a phit 1540 includes a 1:1 mapping of the link width to bits (e.g., a 20-bit link width includes a phit of 20 bits, etc.). Flits may be larger in size, such as 184, 192, or 200 bits. Note that if the phit 1540 is 20 bits wide and the size of flit 1535 is 184 bits, then a fraction of phits 1540 is needed to transmit a flit 1535 (e.g., 9.2 phits with 20 bits to transmit a 184-bit flit 1535 or 9.6 at 20 bits to transmit a 192-bit flit). Note that the widths of the basic link on the physical layer may vary. For example, the number of lanes per direction may comprise 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, link layer 1510 a,bmay be capable of embedding multiple portions of different transactions into a single flit, and multiple headers (e.g., 1, 2, 3, 4) may be embedded in the flit. Here, HPI divides the headers into corresponding slots to allow multiple messages in the flit that are destined for different nodes.The physical layer 1505 a,bis responsible for quickly transferring information to the physical medium (electrical or optical, etc.), in one embodiment. The physical link is point-to-point between two link layer entities, such as layer 1505a and 1505b. Link layer 1510 a,b abstracts physical layer 1505 a,bfrom the upper layers and provides the capability for reliable data transfer (as well as for requests) and manages flow control between two directly connected entities. It is also responsible for virtualizing the physical channel into multiple virtual channels and message classes. Protocol layer 1520 a,b builds on link layer 1510 a,bto map protocol messages into the respective message classes and virtual channels before passing to physical layer 1505 a,bfor transmission over the physical links. The link layer 1510 a,bmay support multiple messages, such as request, snoop, response, writeback, non-coherent data, etc.In one embodiment, to provide reliable transmission through link layer 1510 a,b, cyclic redundancy check (CRC) error checking techniques are provided to isolate the effects of routine bit errors that may occur on the physical link. Link layer 1510a generates the CRC at the transmitter and checks at receiver link layer 1510b.In some embodiments, link layer 1510 a,buses a credit scheme for flow control. During initialization, a transmitter receives a given number of credits for transmitting packets or flits to a receiver. When a packet or a receipt is sent to the receiver, the transmitter down-converts its credit counters by a credit representing either a packet or a flit, depending on the type of virtual network used. If a buffer is enabled at the receiver, credit for the relevant type of buffer is fed back to the transmitter. In one embodiment, when the credits of the transmitter for a given channel are exhausted, it stops transmitting flits in that channel. Essentially, credits are returned after the receiver consumes the information and releases the associated buffers.In one embodiment, routing layer 1515a,b provides a flexible and distributed path for routing packets from a source to a destination. In some platform types (e.g., single processor and dual processor systems), this layer may not be explicit, but could be part of link layer 1510 a,b; in such a case, this layer is optional. It is based on the virtual network and the message class abstraction provided by the link layer 1510 a,bas part of the function of determining how to route the packets. In one implementation, the routing function is defined by the implementation of specific routing tables. Such a definition enables a variety of use models.In one embodiment, protocol layer 1520 a,b implements the communication protocols, ordering rules and coherency maintenance, I / O, interrupts, and other higher level communication. Note that in one implementation, protocol layer 1520 a,b provides messages for negotiated the performance states for components and the system. As a potential addition, the physical layer 1505a,b may also specify power states of the individual links independently or in cooperation.Multiple agents may be interconnected with an HPI architecture, such as a home agent (issues requests to memory), caching (issues requests to coherent memory and responds to snoops), configuration (handles configuration transactions), interrupts (operates interrupts), legacy (handles legacy transactions), non-coherent (handles non-coherent transactions), and others. A more detailed discussion of the layers for HPI is provided below.An overview of some potential features of HPI: does not use pre-allocation at home nodes; does not use ordering requirements for a series of message classes; packs multiple messages into a single flit (protocol header) (i.e., a packed flit that can accommodate multiple messages in defined slots); wide link with scaling capability of 4, 8, 16, 20, and more lanes; wide error checking scheme that can employ 8, 16, 32, or up to 64 bits for error protection; and using an embedded timing scheme.HPI Physical LayerThe physical layer 1505a,b (or PHY) of HPI is above the electrical layer (i.e., electrical conductors connecting two components) and below the link layer 1510a,b as shown in Figure 15. The local and remote electrical layers are connected by physical media (e.g., cable, conductor, optical, etc.). The physical layer 1505 a,bhas two main phases, initialization and operation, in one embodiment. During initialization, the link is opaque to the link layer and the signaling may involve a combination of clocked states and handshake events. During operation, the link is transparent to the link layer and signalling occurs at a rate with all lanes cooperating as a single link. During the operational phase, the physical layer transports flits from agent A to agent B and from agent B to agent A. The connection is also referred to as a link and abstracts some physical aspects, including media, width, and speed, from the connection layers while passing and control / status of the current configuration (e.g., being enabled) are exchanged with the connection layer. The initialization phase includes subphases, e.g., polling and configuration. The operating phase also includes slave phases (e.g., link power management states).In one embodiment, physical layer 1505 a,b also includes: compliance with a reliability / fault standard, tolerates a fault on a path on a link and jumps to a fraction of the desired width, tolerates single faults in the opposite direction of a link, support addition / removal during ongoing operation, enable / disable PHY ports, timeout in initialization attempts when the number of attempts has exceeded a predetermined threshold, etc.In one embodiment, HPI uses a rotating bit pattern. For example, if a flit size is not aligned with a multiple of the lanes in an HPI link, the flit may not be able to be sent in an integer multiple of transmissions across the lanes (e.g., a 192-bit flit is not an accurate multiple of an example 20-lane link. Thus, at x20 flits may be offset to avoid waste of bandwidth (e.g., sending a partial flit at a location without using the rest of the lanes). The offset is determined in one embodiment to optimize latency of key fields and multiplexers in the transmitter (Tx) and receiver (Rx). The fixed pattern also potentially provides clean and fast transition to / from a narrower width (e.g., x8) and seamless operation at the new width.In one embodiment, HPI uses an embedded clock, such as a 20-bit embedded clock or another number of bits embedded clock. Other high performance interfaces may use a forwarded clock or another clock for in-band reset. Clock embedding in HPI potentially reduces pin occupancy. However, using an embedded clock may result in different devices and methods for handling in-band reset in some implementations. As a first example, after initialization, a blocked link status is employed to retain the link's flit transfer and to enable PHY usage (described in detail in Appendix A). As a second example, electrically ordered groups such as an electrically idle ordered set (EIOS) may be used during initialization.In one embodiment, HPI may use a first bit width direction without a forwarded clock and a second, smaller bit width link for power management. As an example, HPI has a partial link width transmission state using a partial width (e.g., a full width (x20) and a partial width (x8)); however, the widths are merely illustrative and may vary. Here, the PHY can handle partial width power management without support or intervention of the link layer. In one embodiment, a blocking link state (BLS) protocol is used to enter partial width transmit status (PWTS). Leaving PWTS may use the BLS protocol or detection of the noise trap, in one or more implementations. Because of the lack of a forwarded clock, exiting PWTLS may include re-equalizing that maintains the determination of the link.In one embodiment, HPI uses Tx adaptation. As an example, loopback status and hardware are used for Tx adaptation. As an example, HPI may count actual bit errors; this may be enabled by performing injection of special patterns. As a result, HPI should be able to achieve better electrical margins at lower power. When using loopback status, a direction can be used as a hardware return channel, with metrics being sent as part of training sequence (TS) payload.In one embodiment, HPI may provide latency determination without exchanging synchronization counter values in a TS. Other connections may perform the determination of latency based on such an exchange of a synchronization counter value in each TS. Here, HPI may use periodically recurrent electronically idle exit ordered sets (EIEOS) as a proxy for the synchronization counter value, aligning the EIEOS with the synchronization counter. This potentially saves TS payload space, eliminates aliasing and DC balance problems, and simplifies the calculation of the latency to add.In one embodiment, HPI provides software and timer control for transitions of a link state machine. Other connections may support a semaphore (hold bit) that is set by hardware upon entering an initialization state. The exit from state occurs when the hold bit is cleared by software. In one implementation, HPI allows the software to control this type of mechanism for entering a sending link state or a status of a loopback pattern. In one embodiment, HPI allows handshake status to be exited based on a software programmable time-out after handshake, potentially facilitating software testing.In one embodiment, HPI uses Pseudo Random Bit Sequence (PRBS) scrambling of the TS. As an example, a 23-bit PRBS (PRBS 23) is used. In one embodiment, the PRBS is implemented by a self-starting storage element of similar bit size as a linear feedback shift register. As an example, a fixed UI pattern with bypass to an adaptation state may be used for scrambling. However, by scrambling the TS with PRBS23, the Rx adaptation can be performed without the bypass. In addition, offset and other errors during clock recovery and sampling can be reduced. The HPI approach relies on the use of LFSRs, which may be self-starting during specific portions of the TS.In one embodiment, HPI supports simulated slow mode without changing the PLL clock frequency. Some concepts may use separate PLLs for slow and fast speeds. Nevertheless, in one implementation, HPI uses emulated slow mode (i.e., PLL clock is running at fast speed; TX repeats bits multiple times; RX performs oversampling to locate edges and identify the bit). This means that ports sharing a PLL can coexist at low and fast speeds. In an example where the multiple is an integer ratio of fast speed to slow speed, different fast speeds may operate at the same slow speed, which may be used during the detection phase of a port while operating.In one embodiment, HPI supports a slow common mode frequency for connection during ongoing operation. The emulated slow mode As described above, HPI ports allow sharing of a PLL for coexistence at slow and fast speeds. If a designer sets the emulation multiple as an integer ratio of fast speed to slow speed, different fast speeds may operate at the same slow speed. Thus, two agents supporting at least one common frequency can be connected in running operation regardless of the speed at which the host port is operated. Software detection may then use the slow mode link to identify and establish the optimal link speeds.In one embodiment, HPI supports reinitializing the link without termination changes. One could provide re-initialization from in-band reset, including clock path terminations changed for the recognition process used for reliability, availability, and maintainability (RAS). In one embodiment, re-initialization for HPI may occur without changing termination values when HPI has an RX screening of incoming signaling to identify good lanes.In one embodiment, HPI supports robust low power link status (LPLS) input. As an example, HPI may have a minimum dwell in LPLS (i.e., a minimum amount of time, UI, counter value, etc. for dwell of a link in LPLS before leaving). Alternatively, the LPLS input may be negotiated and an intra-band reset then used to enter LPLS. However, this may mask an actual intra-band reset, which in some cases returns to the second agent. In some implementations, HPI allows a first agent to enter LPLS and a second agent to enter reset. The first agent is not responsible for a period of time (i.e., the minimum dwell time), which allows the second agent to complete the reset and subsequently wake up the first agent, thus allowing a much more efficient, robust entry into LPLS.In one embodiment, HPI supports features such as debouncing detection, wake-up, and continuous screening for web failures. HPI may search for a specified signaling pattern for an extended period of time to detect a valid wake-up from an LPLS, thus reducing the possibility of spurious wake-up. The same hardware can also be used in the background to continually screen for bad lanes during the initialization process and create a more robust RAS feature.In one embodiment, HPI supports deterministic exit with lockstep and restart / playback. In HPI, when operating at full width, some TS boundaries may coincide with flit boundaries. Here, HPI may identify and specify the exit boundaries so that lockstep behavior may be maintained with another link. In addition, HP may indicate timers that may be used to maintain lockstep with a link pair. After initialization, HPI may also support intra-band reset disabled operation to support some variants of lockstep operation.In one embodiment, HPI supports the use of the TS header instead of payload for initialization key parameters. Alternatively, TS payload data may be used to exchange unit parameters such as ACKs and path numbers. And DC levels for communicating the lane polarity may also be employed. HPI can still use matched DC codes in the TS header for key parameters. This potentially reduces the number of bytes required for operation and potentially allows for the use of an entire PRBS23 pattern for TS scrambling, thus reducing the need for DC matching of the TS.In one embodiment, HPI supports measures to increase the noise immunity of active lanes during the ingress / egress of inactive lanes in Partial Width Transmitting Link State (PWTLS). In one embodiment, null flits (or other non-retestable flits) around the width change point may be used to increase the noise immunity of active lanes. In addition, HPI may employ null flits around the beginning of the PWTLS exit (i.e., the null flits may be broken by data flits). HPI may also use specialized signaling, whose format may vary to reduce opportunities for detection of spurious wake-up.In one embodiment, HPI supports the use of specialized patterns during PWTLS egress to enable non-blocking equalization. Alternatively, inactive lanes cannot be removed at PWTLS exit, as they retain distortion by a propagated clock. By using an embedded clock, HPI may still use specialized signaling, whose format may vary to reduce spurious wake-up detection opportunities and also to enable equalization without blocking the flit stream. This also enables more robust RAS by seamless disabling of failed lanes, re-establishing them, and re-online locations without blocking the flow of flits.In one embodiment, HPI supports entry into low power link status (LPLS) Without link layer support and more robust LPLS exit. Alternatively, the link layer negotiation may depend on between predetermined master and slave for entering LPLS from Transmitting Link State (TLS). In HPI, the PHY can handle negotiation using blocking link state (BLS) codes and support both agents as master or initiator as well as entry into LPLS directly from PWTLS. The exit from LPLS may be based on debouncing a noise trap using a specific pattern, followed by handshake between the two sides and a time-out induced in-band reset if one fails.In one embodiment, HPI supports control of nonproductive loopings during initialization. Alternatively, an initialization error (e.g., lack of good lanes) may result in re- and frequent attempted initialization, potentially wasting power and difficult to analyze. In HPI, the link pair may attempt initialization at the predetermined frequency before abort and power down to a reset state, and the software may make adjustments before re-attempting initialization. This potentially improves the RAS of the system.In one embodiment, HPI IBIST (Interconnect Built-in Self-test) options support. In one embodiment, a pattern generator may be used that enables two non-correlated maximum length PRBS23 patterns for any pins. In one embodiment, HPI may be able to support four such patterns and provide the capability to control the length of these patterns (i.e., dynamically vary test patterns and PRBS23 length).In one embodiment, HPI provides enhanced logic for equalizing lanes. As an example, the TS boundary after TS lock can be used to equalize the webs. In addition, HPI can equalize by comparing trajectory PRBS patterns in the LFSR during specific points in the operation. Such warping may be useful in test chips that may lack the ability to detect TS or state machines to manage the warping.In one embodiment, the exit from the call transfer initialization occurs at a planetary alignment TS boundary. In addition, from this point on, HPI may support negotiated delay. In addition, the order of exit between the two directions can be controlled by using master-slave determinants, allowing one instead of two planetary alignment controls for the link pair.Some implementations use a fixed 128UI pattern for TS scrambling. Others use a fixed 4k PRBS23 for TS scrambling. HPI, in one embodiment, allows the use of PRBS of any length, having an entire (8M-1) PRBS23 sequence.In some architectures, adaptation is fixed duration. In one embodiment, the escape of adaptation is effected by means of handshake instead of time lapse. This means that adaptation times between the two directions can be asymmetric and need not be as long as one of the sides requires.In one embodiment, a state machine may bypass states if such state operations need not be re-executed. However, this may result in more complex designs and validation exceptions. HPI does not use a bypass - instead, operations are distributed so that short timers can be used in any state to perform the operations and bypasses are avoided. This provides potentially more uniform and synchronized state machine transitions.In some architectures, passed clock is used for in-band reset and the link layer is used for performing partial width transfer and low power link input. HPI uses functions similar to state codes with block association. These codes could potentially have bit errors that result in 'misses' in Rx. HPI has a protocol for handling misses and means for handling requests for asynchronous reset, low power link state and partial width link status.In one embodiment, a 128 UI scrambler is used for loopback TS. However, this can result in aliasing for TS blocking when loopback begins; thus, certain architectures change the payload to only 0. In another embodiment, HPI uses a uniform operation and periodically occurring unencrypted EIEOS for TS locking.Certain architectures use encrypted TS during initialization. In one embodiment, HPI defines supersequences that are combinations of scrambled TSs of various lengths and unencrypted EIEOSs. This allows for more random transitions during initialization and also simplifies the TS lock, latency determination, and other operations.HPI Link LayerReferring again to Figure 15, an embodiment of a logical block for link layer 1510a,b is illustrated. In one embodiment, link layer 1510 a,bsecure reliable data transfer between two protocol or routing entities. It abstracts the physical layer 1505a,b from the protocol layer 1520a,b; is responsible for flow control between two protocol agents (A,B) and provides virtual channel services to the protocol layer (message classes) and the routing layer (virtual networks). The interface between protocol layer 1520a,b and link layer 1510a,b is typically at the packet level. In one embodiment, the smallest transmission unit at the link layer is referred to as a flit with a predetermined number of bits, such as 192. Link layer 1510 a,b builds on physical layer 1505 a,bto frame physical layer 1505 a,b(Phit) transmission unit in link layer 1510 a,b(Flit). In addition, link layer 1510 a,bmay be logically broken into two parts, a transmitter and a receiver. A transmitter / receiver pair at one entity may be connected to a receiver / transmitter pair at another entity. Flow control is often performed on both a flit and a packet base. Error detection and correction are also potentially performed based on a flit level.In one embodiment, flits are extended to 192 bits. However, any range of bits, such as 81-256 (or more), may be used in various variations. Here, the CRC field is also increased (e.g., 16 bits) to handle more bulky payload data.In one embodiment, TIDs (transaction IDs) are 11 bits long. As a result, pre-assignment and activation of distributed home agents can be dispensed with. Further, in some implementations, the use of 11 bits allows the use of TID without having to employ an extended TID mode.In one embodiment, header flits are divided into 3 slots, of which 2 is the same size (slots 0 and 1) and another smaller slot (slot 2). A flow field may be available for use by either slot 0 or slot 1. The messages that can use slots 1 and 2 are optimized, thereby reducing the number of bits required to encode the opcodes of these slots. When a header that needs more bits than slot 0 enters the link layer, slot algorithms are available that allow it to take payload bits from slot 1 for additional space. Special control flits (e.g., LLCTRL) may occupy all 3 slots with bits for their need. Slot algorithms may also be present for partially used connection cases to allow the use of individual slots while other slots carry no information. Other connections may allow a single message instead of multiple messages per flit. The size of the slots in the flit and the types of messages that can be placed in each slot potentially provide the increased bandwidth of HPI even at a reduced flit rate. A more detailed description of flits and the multi-slot header can be found in the Flit Definition section in Appendix B.In HPI, a large CRC baseline may improve error detection. For example, 16-bit CRC is employed. As a result of the larger CRC, more comprehensive payload data may also be used. The 16 CRC bits in combination with a polynomial used with these bits enhance error detection. As a result, there are a minimum number of gates for providing: 1) 1-4 bit errors are detected, 2) errors of burst length 16 or less are detected.In one embodiment, rolling CRC is used based on two CRC-16 equations. Two 16-bit polynomials can be used - the polynomial of HPI CRC-16 and a second polynomial. The second polynomial has the smallest number of gates for implementing, while the characteristics of 1) all 1-7 detected bit errors, 2) burst protection per lane in x8 link widths, 3) all burst length 16 errors or less continue to be detected.In one embodiment, a reduced maximum flit rate (9.6 rather than 4 UI) is used, but increased throughput of the connection is achieved. As a result of the increased flit rate, the introduction of multiple slots per flit, the optimized use of payload bits (altered algorithms to remove or re-place rarely used fields), more connection efficiency is achieved.In one embodiment, a portion of the 3 slot bearer comprises 192-bit flits. The flow field asserts 11 extra payload bits for either slot 0 or slot 1. note that if a larger flit is used, more flowing bits can be used. And as a result, if a smaller flit is employed, fewer bits are provided flowing. By allowing a field to flow between the two slots, we can provide the extra bits needed for particular messages while maintaining the extent of the 192 bits and maximizing bandwidth usage. Alternatively, providing an 11 bit HTID field for each slot may use an additional 11 bits in the flit that could not be used as efficiently.Some connections may transmit a viral state in protocol-level messages and a poisson state in data flits. In one embodiment, HPI protocol level and Poisson status messages are moved to control flits. Since these bits are used rarely (only in the case of errors), their removal from the protocol level messages potentially increases flit usage. Their injection using control flits allows the isolation of the errors.In one embodiment, CRD and ACK bits in a flit allow a number of credits, e.g., eight, or the number of ACKs, e.g., 8, to return. As part of the fully encoded credit fields, these bits are used as credit[n] and acknowledge[n] when slot 2 is encoded as LLCRD. This potentially improves efficiency by allowing each flit to return the number of VNA credits and the number of acknowledge using a total of only 2 bits, but its definitions are also possible to remain consistent when using fully encoded LLCRD feedback.In one embodiment, VNA vs. VNO / 1 coding applies (saves lightning by aligning slots with the same coding). The slots in a multi-slot header flit may be aligned with only VNA, only VN0, or only VN1. The related enforcement removes bits per slot specifying VN. This increases the efficiency of flit bit utilization and allows the potential expansion of 10-bit TIDs to 11-bit TIDs.Some fields allow feedback only in increments of 1 (for VN0 / 1), 2 / 8 / 16 (for VNA), and 8 (for acknowledge). This means that feeding back a large number of pending credits or acknowledgments can use multiple feedback messages. It also means that odd numbered feedback values for VNA and acknowledge may remain unprocessed during the cumulation of an even divisible value. HPI may have fully encoded credit and ack feedback fields, whereby an agent may return all cumulative credits or acknowledgements for a pool with a single message. This potentially improves connection efficiency and also potentially simplifies logic implementation (the feedback logic may implement a "clear" signal rather than a full down counter).Routing LayerIn one embodiment, routing layer 1515a,b provides a flexible and distributed method for routing HPI transactions from a source to a destination. The scheme is flexible because routing algorithms for multiple topologies may be specified by programmable routing tables at each router (programming is performed by firmware, software, or a combination thereof in embodiments). The routing functionality may be distributed; the routing may be done by a series of routing steps, each routing step being defined by lookup of a table in either the source, intermediate or destination router. Source lookup may be useful to inject an HPI packet into the HPI structure. Lookup at an intermediate router may be used to route an HPI packet from an input port to an output port. The target port lookup may be used to target the target HPI protocol agent. Note that in some implementations, the routing layer is thin because the routing tables, and thus the routing algorithms, are not specifically defined by the specification. This allows for a variety of usage models, including flexible platform architecture topologies for definition by the system implementation. The routing layer 1515a,b is based on the link layer 1510a,b to provide the use of up to three (or more) virtual networks (VNs) - in one example, two non-blocking VNs, VNO and VN1, are with multiple defined message classes in each virtual network. A shared adaptive virtual network (VNA) may be defined in the link layer, but this adaptive network cannot be directly exposed in routing concepts because each message class and VN may have dedicated resources and guaranteed forwarding progress.A non-exhaustive example list of routing rules includes: (1) (message class invariance): an incoming packet belonging to a particular message class may be routed to an outgoing HPI port / virtual network in the same message class; (2) (switching) HPI platforms may support the switching types "store and forward" and "virtual direct through.". In another embodiment, HPI may not support "Wormhole" or "Circuit" switching. (3) (Connection Blocking Free) HPI platforms may not be based on adaptive streams for blocking free routing. With platforms using both VN0 and VN1, the 2 VNs can be used together for non-blocking routing; and (4) (VN0 for "leaf" routers). In HPI platforms which can use both VN0 and VN1, it is allowed to use VN0 for those components whose routers are not used for routing; that is, incoming ports have HPI destinations terminating at the component in question. In such a case, packets from different VNs may be routed to VN0. Other rules (e.g., movement of packets between VN0 and VN1) may be controlled by a platform-dependent routing algorithm.Routing Step: A routing step, in one embodiment, denotes a routing function (RF) and a selection function (SF). The routing function can take as inputs an HPI port on which a packet enters and a destination NodeID; a 2-tuple - the HPI port number and the virtual network - to which the packet is to follow on its path to the destination is then output. It is permissible that the routing function is additionally dependent on the incoming virtual network. Furthermore, it is allowed that the routing step yields multiple <port#, virtual network> pairs. The resulting routing algorithms are referred to as adaptive. In such a case, a selection function SF may select a single 2-tuple based on additional state information that the router has (for example, in adaptive routing algorithms, the selection of a particular port of a virtual network may depend on the local congestion conditions). A routing step, in one embodiment, consists of applying the routing function and then the selection function to yield the 2-tuple / s.Router Table Simplifications: HPI platforms can implement allowed subsets of the virtual networks. Such subsets simplify the amount of virtual channel buffering (reduced number of columns) associated with the routing table and arbitration at the router switch. These simplifications can be at the expense of platform flexibility and features. VN0 and VN1 may be non-blocking networks that provide non-blocking freedom either together or individually depending on the usage model, usually with minimal virtual channel resources allocated to them. The flat organization of the routing table may have a size corresponding to the maximum number of NodeIDs. In such an organization, the routing table may be indexed by the destination NodeID field and possibly the virtual network ID field. The table organization can also be created hierarchically, wherein the destination NodeID field is divided into a plurality of subfields depending on the implementation. When divided into "local" and "non-local" parts, the "non-local" part of the routing is completed, for example, before the routing of the "local" part. The potential benefit of reducing the table size at each input port adds to the potential cost of forcibly assigning NodeIDs to HPI components in a hierarchical manner.Routing Algorithm: In one embodiment, a routing algorithm defines the set of allowed paths from a source module to a destination module. A particular path from the source to the destination is a subset of the allowed paths and is obtained as a series of routing steps defined by starting with the router at the source, traversing zero or more intermediary routers, and ending with the router at the destination. Note that although an HPI structure may have multiple physical paths from a source to a destination, the allowed paths are those defined by the routing algorithm.HPI Coherence ProtocolIn one embodiment, the HPI coherency protocol is included in layer 1520 a,bto assist agents in caching lines of data from memory. An agent wishing to cache memory data may use the coherency protocol to read in the line of data to be loaded into its cache. An agent wishing to modify a line of data in its cache may use the coherency protocol to acquire property on the line prior to modifying the data. After changing a line, an agent may follow the protocol requests by keeping them in its cache until it either writes the line back to memory or receives the line in response to an external request. Finally, an agent may satisfy external requests to invalidate a line in its cache. The protocol ensures coherence of the data by presetting the rules that all caching agents can follow. It also provides the means for agents without caches to coherently read and write memory data.Two conditions may be enforced to support transactions using the HPI coherency protocol. First, the protocol maintains data consistency, for example on an address basis, among the data in caches of the agents and between that data and the data in memory. Informally, data consistency may denote any valid data line in an agent's cache that represents a most recently current value of the data, and data sent in a coherency protocol packet represents the most recently current value of the data at the time of sending. If there is no valid copy of the data in the cache or in the transfer, the protocol may ensure that the last updated value of the data is in memory. Second, the protocol provides well-defined commit points for requests. Commit points for reads may indicate when the data is usable; and for writes, may indicate when the written data is globally observable and loaded by subsequent reads. The protocol may support these commement points for both cacheable and non-cacheable (UC) requests in coherent memory space.The HPI coherency protocol may also ensure that forwarding progress of coherency requests is made by an agent to an address in coherent memory space. Of course, transactions for proper system operation can eventually be executed and withdrawn. The HPI coherency protocol may also not have a retry approach to resolving conflicts in resource allocation in some embodiments. Thus, the protocol itself may be defined to contain no circular resource dependencies, and implementations may be designed in their concepts to introduce no dependencies that may result in blocking. Additionally, the protocol may indicate where concepts are capable of providing adequate access to protocol resources.Logically, in one embodiment, the HPI coherency protocol consists of three parts: coherency (or caching) protocol, home agents, and the HPI interconnect fabric that connects the agents. Coherency agents and home agents cooperate to achieve data consistency by exchanging messages over the link. Link layer 1510 a,band the associated description provide the details of the link structure and compliance with the coherency protocol requirements discussed herein. (Note that the division into coherency agents and home agents is done for clarity. A concept may include multiple agents of both types in a port or may also combine agent behavior in a single concept unit.)In one embodiment, HPI does not pre-allocate resources of a home agent. Here, a receiving agent receiving a request allocates resources for its processing. An agent sending a request allocates resources for responses. In this scenario, HPI may follow two general rules relating to resource allocation. First, an agent receiving a request may be responsible for allocating the resource to process it. Second, an agent that generates a request may be responsible for allocating resources to process responses to the request.The assignment of resources can also be extended to HTID (along with RNID / RTID) in snoop requests, the potential reduction of home agent use, and the forwarding of responses to support responses to the home agent (and data forwarding to the requesting agent).In one embodiment, home agent resources in snoop requests and forwarding responses to support responses to the home agent (and forwarding data to the requesting agent) are also not pre-allocated.In one embodiment, there is no pre-assignment of the home resource's ability to send CmpO "early" before the home agent has completed processing the requests if the requesting agent is sure to reuse its RTID resource. Generally treating snoops with similar RNID / RTID in the system is also part of the protocol.In one embodiment, conflict resolution is performed using an ordered response channel. A coherency agent uses RspCnflt as a request for a home agent to send FwdCnfltO that is ordered with CmpO (if already scheduled) for the contentioned request of the coherency agent.In one embodiment, HPI supports resolution over an ordered response channel. A coherency agent uses Snoop information to support FwdCnfltO processing, with no "type" information and no RTID to forward data to the requesting agent.In one embodiment, a coherency agent blocks forwarding for writeback requests to maintain data consistency. However, it also allows the coherency agent to use a writeback request to pass non-cacheable (UC) data before processing in the forward direction, and allows the coherency agent to write back partial text lines rather than a protocol that allows partial implicit writeback for forwarding.In one embodiment, a read invalidation request (RdInv) is supported that accepts exclusive-state data. Semantics of non-cacheable (UC) reads include flushing modified data into memory. However, some architectures allowed this to validate reads by forwarding M data, forcing the requesting agent to clean the line if M data was received. The RdInv operation simplifies the flow, but does not allow E data to be forwarded.In one embodiment, HPI supports InvItoM-to-IODOC functionality. An InvItoM requests exclusive ownership of a cache line without receiving data and with the intention of writing back soon thereafter. A required cache state may be an M state and E state, or both.In one embodiment, HPI supports WbFlu for persistent memory empty. An embodiment of a WbFlush is shown below. It may be sent as a result of a persistent handover. May flush a write to persistent memory.In one embodiment, HPI supports additional operations such as SnpF for fanout snoops generated by the routing layer. Some architectures do not have explicit support for fanout snoops. Here, an HPI home agent generates individual fanout snoop requests, and in response, the routing layer generates snoops to all peer agents in the fanout cone. The home agent can expect snoop responses from each of the agent sections.In one embodiment, HPI supports additional operations such as SnpF for fanout snoops generated by the routing layer. Some architectures do not have explicit support for fanout snoops. Here, an HPI home agent generates individual fanout snoop requests, and in response, the routing layer generates snoops to all peer agents in the fanout cone. The home agent can expect snoop responses from each of the agent sections.In one embodiment, HPI supports explicit writeback with cache push-Hint (WbPushMtoI). In one embodiment, a coherency agent writes modified data back with a hint for the home agent to be able to push the modified data into a "local" cache, storing it in the M state without writing the data to memory.In one embodiment, a coherency agent may maintain the F state when passing shared data. In one example, an F-state coherency agent receiving a sharing snoop or forwarding operation after such snoop may maintain the F-state while sending the S-state to the requesting agent.In one embodiment, protocols may be interleaved, where one table references another sub-table in the "next state" columns, and the interleaved table may include additional or finer-grained guardians to indicate which rows (behaviors) are allowed.In one embodiment, protocol tables use line bridging to indicate equally acceptable behaviors (lines), rather than adding "bias" bits, to select among the behaviors.In one embodiment, action tables are organized for use as a functionality engine for BFM (Validation Environment Tool), such that the BFM team does not need to generate a BFM engine based on its own interpretation.HPI Non-Coherent ProtocolIn one embodiment, HPI supports non-coherent transactions. As an example, a non-coherent transaction is referred to as an operation that does not participate in the HPI coherency protocol. Non-coherent transactions include requests and their corresponding terminations. In some specific transactions, a transfer mechanism.The foregoing discussion outlines features of one or more embodiments of the subject matter disclosed herein. These embodiments are provided to enable those of ordinary skill in the art (those skilled in the art) to better understand various aspects of the present disclosure. Certain already known terms, as well as underlying technologies and / or standards, may be given without detailed description. It is anticipated that those skilled in the art will have or have access to background knowledge or information about these technologies and standards sufficient to practice the teachings of the present specification.Those skilled in the art will appreciate that they may readily use the present disclosure as a basis for designing or altering other processes, structures, or variations to perform the same purposes and / or to achieve the same advantages of the embodiments introduced herein. Those skilled in the art will also appreciate that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they may make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.In the foregoing description, certain aspects of some or all of the embodiments are described in more detail than is necessarily required to implement the appended claims. These details are provided by way of non-limiting example only for the purpose of providing context and illustrating the disclosed embodiments. Such details are not to be understood as being required, and are not to be read as limitations on the claims. The phrase may refer to "one embodiment" or "embodiments.". These formulations and any further references to embodiments should be broadly understood to mean any combination of one or more embodiments. Further, the plurality of features disclosed in a particular "embodiment" may be as well spread over multiple embodiments. For example, if features 1 and 2 are disclosed in "one embodiment", embodiment A may include feature 1 lacking feature 2, while embodiment B may include feature 2 lacking feature 1.This specification may provide illustrations in block diagram format if certain features are disclosed in separate blocks. This should be broadly understood to disclose how various features interact, but is not intended to imply that the features in question must necessarily be embodied in separate hardware or software. Further, if a single block discloses more than one feature in the same block, the respective features need not necessarily be embodied in the same hardware and / or software. For example, under certain circumstances, a computer "memory" could be distributed or mapped between multiple levels of cache or local memory, main memory, backed-up volatile memory, and various forms of persistent storage such as a hard disk, a storage server, an optical disk, a tape drive, or the like. In certain embodiments, some of the components may be omitted or consolidated. In a general sense, the arrangements depicted in the figures may be more logical in their representations, while a physical architecture may have various permutations, combinations, and / or hybrids of these elements. Numerous possible design configurations may be employed to achieve the operational goals outlined herein. Accordingly, the associated infrastructure includes a variety of replacement arrangements, designs, device capabilities, hardware configurations, software implementations, and equipment options.Here, references may be made to a computer readable medium, which may be a tangible and non-transitory computer readable medium. As used in this specification and in the course of the claims, a "computer readable medium" is to be understood to include one or more computer readable media of the same or different types. A computer readable medium may include, by way of non-limiting example, an optical drive (e.g., CD / DVD / Blu-ray), a hard drive, a solid state drive, flash memory, or other non-transitory medium. A computer readable medium could also include a medium such as a read only memory (ROM), an FPGA, or an ASIC configured to perform the desired instructions; stored instructions for programming an FPGA or an ASIC to perform the desired instructions; a intellectual property (IP) block that can be integrated into hardware in other circuits or instructions directly encoded in hardware or microcode on a processor such as a microprocessor, a digital signal processor (DSP), a microcontroller, or in any other suitable component, device, element, or object where appropriate and based on particular needs. A non-transitory storage medium is expressly intended to include any non-transitory special or programmable hardware configured to provide the disclosed operations or to cause a processor to perform the disclosed operations.Various elements may be "communicatively", "electrically", "mechanically", or otherwise "coupled" to one another as during the course of this specification and the claims. Such coupling may be direct, point-to-point, or may include intervening devices. For example, two devices may be communicatively coupled to each other via a controller that enables communication. Devices may be electrically coupled to one another via intermediate devices such as signal amplifiers, voltage dividers or buffers. Mechanically coupled devices may be indirectly mechanically coupled.Any "modules" or "engines" disclosed herein may refer to or include software, a software stack, a combination of hardware, firmware, and / or software, circuitry configured to perform the function of the engine or module, or any computer readable medium in accordance with the disclosure above. Such modules or engines may be provided on or in connection with a hardware platform having hardware computing resources, such as a processor, memory, memory, connections, networks and network interfaces, accelerators, or other suitable hardware, under suitable circumstances. Such hardware platforms may be provided as a single monolithic device (e.g., in a PC form factor) or with elements or portions of the distributed function (e.g., a "federation node" in a high performance data center where computing, memory, storage, and other resources may be dynamically allocated and need not be local to each other).Flow diagrams, signal flow diagram, or other representations showing operations performed in a particular order may be disclosed herein. Unless expressly stated otherwise or unless required in a particular context, the order is to be understood as a non-limiting example only. Further, in cases where one operation is shown following another, further involved operations may also occur, which may or may not be related thereto. Some operations may also be performed simultaneously or in parallel with each other. In cases where an operation is indicated as being based on or corresponding to another element or operation, it is to be understood that it is implied that the operation is based at least in part on or at least in part executing accordingly on the other element or operation. This should not be construed to imply that the operation is solely or exclusively based on the element or operation, or is solely or exclusively performed accordingly.Any hardware elements disclosed herein may be readily provided, in whole or in part, in a system-on-a-chip (SoC) including a central processing unit (CPU) package. An SoC represents an integrated circuit (IC) that integrates components of a computer or other electronic system into a single chip. Thus, for example, client devices or server devices may be provided wholly or partly in a SoC. The SoC may include digital, analog, mixed signal, and radio frequency functions, all of which may be provided on a single chip substrate. Other embodiments may include a multichip module (MCM), having a plurality of chips positioned in a single electronic package and configured to interact tightly with each other through the electronic package.In a general sense, any suitable configured circuit or processor may execute any type of instructions associated with the data to achieve the operations recited herein. A processor disclosed herein could convert an element or article (e.g., data) from one state or ping to another state or ping. Further, the tracked, transmitted, received, or processor stored information may be provided in any database, register, table, cache, queue, control list, or storage structure based on particular needs and implementations, all of which may be referenced in any suitable timeframe. Any of the memory or storage elements disclosed herein are intended to be encompassed by the broad terms "memory" or "memory" where appropriate.Computer program logic implementing the functionality described herein or a portion thereof is embodied in various forms including, but in no way limited to, source code form, computer executable form, machine instructions or microcode, programmable hardware, and various intermediate forms (e.g., forms generated by an assembler, compiler, linker, or locator). In one example, source code includes a series of computer program instructions implemented in various programming languages, such as object code, assembler language, or higher-level language such as OpenCL, FORTRAN, C, C++, JAVA, or HTML, for use with various operating systems or operating environments, or in hardware description languages such as Pace, Verilog, and VHDL. The source code may define and use various data structures and communication messages. The source code may be in a computer-executable form (e.g., via an interpreter), or the source code may be converted (e.g., via a translator, assembler, or compiler) to a computer-executable form or converted to an intermediate form such as byte code. Where appropriate, any of the foregoing elements may be used to establish or describe suitable discrete or integrated circuits, whether sequential, combinatorial, state machines, or otherwise.Implementation ExamplesThe following examples are provided for illustrative purposes.Example 1 includes an apparatus comprising: a plurality of connections to communicatively couple an accelerator device to a host device; and an address translation module (ATM) to provide address mapping between host physical address spaces (HPA) and guest physical address spaces (GPA) to the accelerator device, wherein the plurality of devices share a common GPA domain and wherein address mapping is to be associated with only one of the plurality of connections.Example 2 includes the apparatus of example 1, wherein the ATM is an address translation unit (ATU).Example 3 includes the apparatus of example 2, wherein the ATU is a distributed ATU.Example 4 includes the apparatus of example 2, wherein each of the plurality of connections is configured to traverse a memory page.Example 5 includes the apparatus of example 1, wherein the ATM is an address translation cache (ATC).Example 6 includes the apparatus of example 5, wherein the ATC is a distributed ATC.Example 7 includes the apparatus of example 5, wherein only the connection associated with the address mapping is to traverse a memory page.Example 8 includes the apparatus of example 1, wherein the plurality of connections is of a single type.Example 9 includes the apparatus of example 8, wherein the type is a Peripheral Component Interconnect Express (PCIe) bus.Example 10 includes the apparatus of example 1, wherein the plurality of connections is of at least two types.Example 11 includes the apparatus of example 10, wherein the at least two types include a Peripheral Component Interconnect Express (PCIe) bus and an Ultra-Path Interconnect (UPI) bus.Example 12 includes the apparatus of example 1, wherein the accelerator device comprises a plurality of accelerator devices connected to a single address domain.Example 13 includes the apparatus of example 1, wherein the ATM is further to provide a translation from an nested GPA to a guest virtual address (GVA).Example 14 includes a block of intellectual property (IP) to provide the apparatus of any of Examples 1-13.Example 15 includes an accelerator device comprising the apparatus of any of Examples 1-13.Example 16 includes the accelerator device of example 15, wherein the accelerator device comprises a field programmable gate array (FPGA).Example 17 includes the accelerator device of example 15, wherein the accelerator device includes an application specific integrated circuit (ASIC).Example 18 includes the accelerator device of example 15, wherein the accelerator device comprises a coprocessor.Example 19 includes a computing system comprising the accelerator of example 15 and a host processor, wherein the host processor is to at least partially disable an address translation unit on the die.Example 20 includes the computing system of example 19, wherein the processor is to completely deactivate the address translation unit.Example 21 includes the computing system of example 19, wherein the processor is to deactivate all rows of the address translation unit on the die except one.Example 22 includes one or more tangible, non-transitory storage media having instructions stored thereon to provide a plurality of connections to communicatively couple an accelerator device to a host device; and to provide an address translation module (ATM) to provide address mapping between host physical address spaces (HPA) and guest physical address spaces (GPA) to the accelerator device, wherein the plurality of devices share a common GPA domain and wherein address mapping is to be associated with only one of the plurality of connections.Example 23 includes the one or more tangible non-transitory storage media of example 22, wherein the ATM is an address translation unit (ATU).Example 24 includes the one or more tangible non-transitory storage media of example 23, wherein the ATU is a distributed ATU.Example 25 includes the one or more tangible non-transitory storage media of example 23, wherein each of the plurality of connections is configured to traverse a memory page.Example 26 includes the one or more tangible non-transitory storage media of example 22, wherein the ATM is an address translation cache (ATC).Example 27 includes the one or more tangible, non-transitory storage media of example 26, wherein the ATC is a distributed ATC.Example 28 includes the one or more tangible non-transitory storage media of example 26, wherein only the connection associated with the address mapping is to traverse a memory page.Example 29 includes the one or more tangible, non-transitory storage media of example 22, wherein the plurality of connections is of a single type.Example 30 includes the one or more tangible, non-transitory storage media of example 29, wherein the type is a Peripheral Component Interconnect Express (PCIe) bus.Example 31 includes the one or more tangible, non-transitory storage media of example 22, wherein the plurality of connections is of at least two types.Example 32 includes the one or more tangible, non-transitory storage media of example 31, wherein the at least two types include a Peripheral Component Interconnect Express (PCIe) bus and an Ultra-Path Interconnect (UPI) bus.Example 33 includes the one or more tangible non-transitory storage media of example 22, wherein the accelerator device comprises a plurality of accelerator devices connected to a single address domain.Example 34 includes the one or more tangible non-transitory storage media of example 22, wherein the ATM is further to provide a translation from an nested GPA to a guest virtual address (GVA).Example 35 includes the one or more tangible, non-transitory storage media of any of Examples 22-34, wherein the instructions comprise instructions to provide a block of mental ownership (IP).Example 36 includes the one or more tangible, non-transitory storage media of any of Examples 22-34, wherein the instructions comprise instructions to provide a field programmable gate array (FPGA).Example 37 includes the one or more tangible, non-transitory storage media of any of Examples 22-34, wherein the instructions comprise instructions to provide an application specific integrated circuit (ASIC).Example 38 includes a computer-implemented method for providing a single address domain for a plurality of connections, comprising communicatively coupling the plurality of connections to an accelerator device and a host device; and providing an address translation module (ATM) for providing address mapping between host physical address spaces (HPA) and guest physical address spaces (GPA) to the accelerator device, wherein the plurality of devices share a common GPA domain and wherein address mapping is to be associated with only one of the plurality of connections.Example 39 includes the method of example 38, wherein the ATM is an address translation unit (ATU).Example 40 includes the method of example 39, wherein the ATU is a distributed ATU.Example 41 includes the method of example 39, wherein each of the plurality of connections is configured to traverse a memory page.Example 42 includes the method of example 38, wherein the ATM is an address translation cache (ATC).Example 43 includes the method of example 42, wherein the ATC is a distributed ATC.Example 44 includes the method of example 42, wherein only the connection associated with the address mapping is to traverse a memory page.Example 45 includes the method of example 38, wherein the plurality of compounds is of a single type.Example 46 includes the method of example 42, wherein the type is a Peripheral Component Interconnect Express (PCIe) bus.Example 47 includes the method of Example 38, wherein the plurality of compounds is of at least two types.Example 48 includes the method of example 38, wherein the at least two types include a Peripheral Component Interconnect Express (PCIe) bus and an Ultra-Path Interconnect (UPI) bus.Example 49 includes the method of example 38, wherein the accelerator device comprises a plurality of accelerator devices connected to a single address domain.Example 50 includes the apparatus of example 38, wherein the ATM is further to provide a translation from an nested GPA to a guest virtual address (GVA).Example 51 includes an apparatus comprising means for performing the method of any of Examples 38-50.Example 52 includes the apparatus of example 51, wherein the means comprises a Mental Property (IP) block.Example 53 includes an accelerator device comprising the apparatus of example 51.Example 54 includes the accelerator device of example 53, wherein the accelerator device comprises a field programmable gate array (FPGA).Example 55 includes the accelerator device of example 53, wherein the accelerator device includes an application specific integrated circuit (ASIC).Example 56 includes the accelerator device of example 53, wherein the accelerator device comprises a coprocessor.Example 57 includes a computing system comprising the accelerator of example 53 and a host processor, wherein the host processor is to at least partially disable an address translation unit on the die.Example 58 includes the computing system of example 57, wherein the processor is to completely deactivate the address translation unit.Example 59 includes the computing system of example 57, wherein the processor is to deactivate all rows of the address translation unit on the die except one.

Claims

An apparatus comprising: a plurality of physical connections for communicatively coupling an accelerator device (508) to a host device (504), wherein the host device (504) executes software; and an address translation module, ATM, for providing address mapping between host physical address spaces, HPA, and guest physical address spaces, GPA, to the accelerator device (508), wherein the plurality of physical connections share a common GPA domain and wherein address mapping is associated with only one of the plurality of physical connections such that the plurality of physical connections presents as a single connection to the software.The apparatus of claim 1, wherein the ATM is an address translation unit, ATU, (516).The apparatus of claim 2, wherein the ATU (516) is a distributed ATU.The apparatus of claim 2, wherein each of the plurality of physical connections is configured to traverse a memory page.The apparatus of claim 1, wherein the ATM is an address translation cache, ATC, (716).The apparatus of claim 5, wherein the ATC (716) is a distributed ATC.The apparatus of claim 5, wherein a corresponding one of the plurality of physical connections complies with a protocol and only the protocol of the corresponding connection associated with the address mapping is allowed to trigger a pass of a memory page.The apparatus of claim 1, wherein the plurality of physical connections belong to a single type.The apparatus of claim 8, wherein the type is a Peripheral Component Interconnect Express, PCIe, bus.The apparatus of claim 1, wherein the plurality of physical connections belong to at least two types.The apparatus of claim 10, wherein one of the at least two types comprises a bus compatible with a Peripheral Component Interconnect Express, PCIe, based protocol.The apparatus of claim 1, wherein the accelerator device (504) comprises a plurality of accelerator devices associated with a single address domain.The apparatus of claim 1, wherein the ATM is further to provide a translation from nested GPA to guest virtual address, GVA.A Intellectual Property, IP, block for providing the apparatus of any of claims 1-13.Accelerator device (508) comprising the apparatus of any of claims 1-13.The accelerator device (508) of claim 15, wherein the accelerator device (508) comprises a field programmable gate array, FPGA.The accelerator device (508) of claim 15, wherein the accelerator device (508) comprises an application specific integrated circuit, ASIC.The accelerator device (508) of claim 15, wherein the accelerator device (508) comprises a coprocessor.A computing system comprising the accelerator (508) of claim 15 and a host processor (504), wherein the host processor (504) is to at least partially disable the address translation unit (512) on the die.The computing system of claim 19, wherein the processor (504) is to completely disable the address translation unit (512).The computing system of claim 19, wherein the processor (504) is to deactivate all rows of the address translation unit (512) on the die except for a row.One or more tangible, non-transitory storage media having instructions stored thereon for providing a plurality of physical connections for communicatively coupling an accelerator device to a host device, wherein the host device executes software; and providing an address translation module, ATM, for providing address mapping between host physical address spaces, HPA, and guest physical address spaces, GPA, to the accelerator device, wherein the plurality of physical connections share a common GPA domain, and wherein address mapping is associated with only one of the plurality of physical connections such that the plurality of physical connections presents as a single connection to the software.The one or more tangible, non-transitory storage media of claim 22, wherein the ATM is an address translation unit, ATU.The one or more tangible, non-transitory storage media of claim 22, wherein the ATM is an address translation cache, ATC.A computer-implemented method for providing a single address domain for a plurality of physical connections, comprising: communicatively coupling the plurality of physical connections to an accelerator device and a host device, wherein the host device executes software; and providing an address translation module, ATM, for providing address mapping between host physical address spaces, HPA, and guest physical address spaces, GPA, to the accelerator device, wherein the plurality of physical connections shares a common GPA domain, and wherein address mapping is associated with only one of the plurality of physical connections such that the plurality of physical connections presents as a single connection to the software.The method of claim 25, wherein the ATM is an address translation unit, ATU.

Citation Information

Patent Citations

  • Performance enhancement of address translation using translation tables covering large address spaces

    US20060069899A1

  • Method and apparatus for managing software controlled cache of translating the physical memory access of a virtual machine between different levels of translation entities

    US20130007408A1