Controlling fabric-attached memory
The host bridge device translates CXL protocol requests into Ethernet packets to manage memory access, addressing the complexity of fabric-attached memory systems and reducing the host's processing burden, enabling efficient and high-performance memory operations.
Patent Information
- Application Number
- GB2023020089
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-08-06
AI Technical Summary
The increasing demand for high-capacity, low-latency memory systems in computer systems is hindered by the complexity of memory management and routing across switching fabrics, placing a processing burden on the host device.
A host bridge device that translates CXL protocol requests into Ethernet protocol packets, managing memory access and reducing the burden on the host by presenting a plurality of memory windows as if directly connected, while handling memory management tasks downstream.
The host bridge device efficiently manages memory access and reduces the processing burden on the host by isolating it from the complex fabric-attached memory system architecture, enabling seamless and high-performance memory operations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD [00011 The present application relates to fabric-attached memory in computer systems, in particular but not exclusively to a device for controlling access to fabric-attached memory, such as by a host. BACKGROUND
[0002] As the demand for artificial intelligence (Al), machine learning (ML), and other computationally expensive applications increases, so too does the need for high capacity, high bandwidth and low-latency computer memory systems.
[0003] Processors configured to implement computational tasks are becoming more powerful, and are able to execute instructions at increasing rates. However, optimal performance of these processors cannot be achieved without comparable performance of memory devices in the computer system. In this context, a processor is a hardware device which executes software for its operation. The terms 'computer' and 'processor' are substantially interchangeable. Since processors are becoming more compact, and the distance between a central processing unit (CPU) and memory impacts the speed with which memory can be accessed, significant challenges must be overcome to increase the available memory of a host without significantly compromising performance of that memory. Moreover, it is important that improvements in memory access keep up with the trend of increasing processor performance.
[0004] Recent developments in interconnect technologies are enabling low-latency and high-bandwidth memory to be attached over a switching fabric.
[0005] One recent development in the field includes Compute Express Link (CXL), an open standard for high-speed and high-bandwidth connections between devices in high performance computing systems. CXL enables lower latency memory devices, such as DDR_DRAM, to be accessible to host systems via a switching fabric.
[0006] CXL is designed to leverage the CXL protocols and the PCI Express physical layer infrastructure to provide high-bandwidth, low latency connectivity between a host processor and other CXL-enabled devices.
[0007] However, whilst the ability to attach memory over a switching fabric expands the range of system architectures that are configurable, the added complexity also brings new challenges to procedures for system setup and memory management, for example, memory device discovery and enumeration. Furthermore, there are many challenges to overcome in routing memory access requests across a switching fabric. A controlling computer (e.g., a host computer) may be responsible for managing many of these challenges, which places a processing and configuration burden on the host. By way of example, reference is made to Wang etal: 'CXL over Ethernet: A novel FPGA-based Memory Disaggregation Design in Data Centers’, which describes connecting CXL memory to a host via Ethernet. The following disclosure is directed to addressing the above problems and removing the processing and configuration burden from the host. SUMMARY
[0008] Examples of the present disclosure do not propagate these problems to the host, as dedicated software takes the burden.
[0009] According to a first aspect of the invention there is provided a host bridge device for use in accessing a fabric-attached memory system, the host bridge device comprising: an upstream port connectable to a host device and configured to operate under a CXL communication protocol to transmit and receive data between the host bridge device and the host; a downstream port connectable to a switched fabric network comprising a fabric switch and at least one fabric-attached memory device, the downstream port configured to operate under an Ethernet communication protocol to transmit and receive data between the host bridge device and the switched fabric network; at least one register storing memory address information for the at least one fabric-attached memory device of the switched fabric network, the memory address information mapping memory windows of a plurality of memory windows to Ethernet destination addresses; and packetization logic configured to: receive, via the upstream port, a first data packet comprising an indication of a memory window; construct, based on the indication of a memory window and the memory address information, a modified data packet in accordance with the Ethernet communication protocol, the modified data packet comprising an indication of an Ethernet destination address corresponding to the memory window of the first packet, and a payload that comprises a payload of the first data packet; and transmit the modified data packet via the downstream port.
[0010] In some embodiments, the host bridge, further comprises registers storing, for each memory window of the plurality of memory windows, one or more of: a dynamic bandwidth parameter indicating a bandwidth associated with a fabric attached memory device, wherein the dynamic bandwidth parameter is determined based on a measurement of bandwidth pertaining to the fabric attached memory device; and a static bandwidth parameter indicating a bandwidth associated with a fabric-attached memory device, wherein the static bandwidth parameter is determined based on an Ethernet fabric topology.
[0011] In some embodiments, the host bridge further comprises registers storing, for each memory window of the plurality of memory windows, one or more of: a dynamic latency parameter indicating a latency associated with the fabric attached memory device, wherein the dynamic latency parameter is determined based on a measurement of latency pertaining to the fabric attached memory device; and a static latency parameter indicating a latency of the fabric-attached memory device, wherein the static bandwidth and static latency parameters are determined based on an Ethernet fabric topology.
[0012] In some embodiments, the host bridge device is configured to receive, from the host via the upstream port, a request for a value of at least one of the static bandwidth parameter and the dynamic bandwidth parameter and, responsive to the request, to transmit the value of the at least one parameter via the upstream port.
[0013] In some embodiments, the host bridge device is configured to receive, from the host via the upstream port, a request for a value of at least one of the static latency parameter and the dynamic latency parameter associated with a first memory window and, responsive to the request, to transmit the value of the at least one parameter via the upstream port.
[0014] In some embodiments, the modified data packet is a layer-2 Ethernet packet, and wherein the indication of an Ethernet destination address comprises a DMAC address.
[0015] In some embodiments, the host bridge is further configured to: receive, via the downstream port, a read data packet in accordance with the Ethernet protocol, the read data packet comprising read data that is read from a fabric-attached memory device; unpack the read data from a payload of the read data packet and identify a transaction identifier in a header of the read data packet; construct a host read packet in accordance with the CXL protocol, the host read packet comprising the read data; and based on the transaction identifier, transmit the packet to the host via the upstream port of the host bridge device.
[0016] In some embodiments, the read data packet comprises error information pertaining to the fabric attached memory device.
[0017] In some embodiments, the host bridge device implements a CXL Dynamic Capacity function.
[0018] In some embodiments, the host bridge device comprises a collect buffer and is configured to: receive plural first data packets comprising respective indications of a same memory window, wherein respective ones of the plural first data packets comprise consecutive addresses within the same memory window; and construct a modified data packet comprising an indication of an Ethernet destination address corresponding to the memory window of the plural first packets, and a payload that comprises respective payloads of the plural first data packets. [00191 In some embodiments, the upstream port is connectable to the host via a CXL switch provided between the host bridge device and the host.
[0020] According to a second aspect of the invention there is provided a computer system, comprising: a host configured to operate under a CXL protocol; a host bridge device in accordance with embodiments of the first aspect, connected to the host via an upstream port of the host bridge device; an Ethernet fabric comprising at least one fabric switch, the Ethernet fabric connected to the host bridge device via a downstream port of the host bridge device; at least one fabric-attached memory device connected downstream of the Ethernet fabric; and for each fabric-attached memory device, a corresponding memory bridge configured to receive a modified data packet from the host bridge device in accordance with the Ethernet protocol, to unpack a first data packet from the modified packet, and to transmit the first data packet to the corresponding fabric-attached memory device.
[0021] In some embodiments, a first fabric attached memory device and a first corresponding memory bridge are comprised within a server appliance connected downstream of the Ethernet fabric.
[0022] In some embodiments, the host is a first host, and the host bridge device is a first host bridge, and the system further comprises: a second host configured to operate under a CXL protocol, the second host configured to operate as a fabric attached memory device comprising host memory; a second host bridge device in accordance with an embodiment of the first aspect, connected to the second host via an upstream port of the second host bridge device, wherein the second host bridge is connected to the Ethernet fabric via a first port in the Ethernet fabric; and a host memory bridge connected to the Ethernet fabric via the first port in the Ethernet fabric and connected to the second host via a downstream port of the second host.
[0023] In some embodiments, the second host bridge device and the host memory bridge are connected to respective CXL ports on the second host.
[0024] In some embodiments, the second host bridge device and the host memory bridge are implemented by a single initiator-responder device connected to the Ethernet fabric via the first port. The initiator-responder device may comprise a Type-2 CXL device.
[0025] In some embodiments the system further comprises a CXL switch provided between the host bridge device and the host, wherein the upstream port of the host bridge device is connectable to the host via the CXL switch.
[0026] In some embodiments, the host bridge is configured to receive a first data packet comprising a memory access request and an indication of a memory window, the memory window corresponding to a plurality of downstream fabric attached memory devices, wherein the Ethernet fabric comprises an Ethernet switch comprising: a first upstream port connected to the host via the host bridge device; a plurality of downstream ports, each of which is connected to a respective downstream fabric-attached memory device; and wherein the host bridge further comprises logic configured to: receive the first data packet from the host; segment a data item of the first data packet into a plurality of segments, including a parity segment; construct, based on the indication of the memory window, a plurality of modified data packets in accordance with the Ethernet communication protocol, each modified data packet comprising: an indication of an Ethernet destination address corresponding to a respective fabric attached memory device associated with the memory window of the first packet, and a payload that comprises a segment of the first data item; and transmit the modified data packets via the downstream port to an Ethernet switch of the ethernet fabric.
[0027] In some embodiments, each memory bridge is configured to unpack a first data packet from the modified packet, and to transmit the first data packet to the corresponding fabric-attached memory device based on one of a CXL protocol, or a JEDEC protocol.
[0028] According to a third aspect of the invention there is provided a method comprising: receiving, from a host device via an upstream port of a host bridge device, a first data packet comprising an indication of a memory window, the first data packet configured in accordance with a CXL protocol; accessing at least one register of the host bridge device storing memory address information for the at least one fabric-attached memory device of a switched fabric network, the memory address information mapping memory windows of a plurality of memory windows to corresponding Ethernet destination addresses of the switched fabric network; constructing, by packetization logic of the host bridge device based on the indication of a memory window and the memory address information, a modified data packet in accordance with an Ethernet communication protocol, the modified data packet comprising an indication of an Ethernet destination address of a fabric attached memory device of the switched fabric network, the Ethernet destination address corresponding to the memory window indicated in the first data packet, wherein a payload of the modified data packet comprises a payload of the first data packet; and transmitting the modified data packet via a downstream port of the host bridge device.
[0029] In some embodiments, the method further comprises receiving, at a memory bridge device via the Ethernet switched fabric network, the modified data packet from the host bridge device in accordance with the Ethernet protocol. The method may comprise unpacking, by the memory bridge device, a first data packet from the modified packet. The method may comprise transmitting the first data packet to a corresponding fabric-attached memory device indicated by the Ethernet destination address in the modified packet.
[0030] According to a fourth aspect of the invention there is provided transitory or non-transitory computer readable media embodying computer readable instructions which, when executed by one or more processor of one or more computer device cause the one or more processor to implement a method in accordance with the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] To assist understanding of the present disclosure and to show how embodiments may be put into effect, reference is made by way of example to the accompanying drawings in which:
[0032] Figure 1 shows a highly schematic block diagram illustrating an exemplary fabric-attached memory system;
[0033] Figure 2a shows a highly schematic packet representation of a write request as it travels through the exemplary system of Figure 1;
[0034] Figure 2b shows a highly schematic packet representation of a read response packet as it travels through the exemplary system of Figure 1;
[0035] Figure 3 shows a highly schematic diagram of a memory space accessible to a host device;
[0036] Figure 4 shows an exemplary computer system in which a host is connected via CXL to plurality of devices;
[0037] Figure 5 shows a second exemplary computer system in which a host is connected to memory subsystems via a CXL switch;
[0038] Figure 6 shows an exemplary fabric-attached memory subsystem in which data striping and redundancy techniques are implemented; and
[0039] Figure 7 shows a set of exemplary registers for implementing the data striping and redundancy techniques.
[0040] In the drawings, corresponding reference characters indicate corresponding components. The skilled person will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. Also, common but well-understood elements that are useful or necessary in a commercially feasible embodiment are often not depicted in order to facilitate a less obstructed view of these various example embodiments.
[0041] It will be appreciated that where reference numerals of the form xxx-a, xxx-b etc., are used to denote a particular instance of a feature of a drawing, the same reference numeral without a specified letter suffix (e.g., a, b, etc.) may denote a generic instance of the same feature. DETAILED DESCRIPTION
[0042] Recent developments in interconnect technologies are enabling low-latency and high-bandwidth memory to be attached over a switching fabric. It is primarily with respect to these fabric-attached memory (FAM) solutions that the present disclosure is directed. DDR5 (Double Data Rate), LPDDR5 (Low-Power DDR), and SCM (Storage-Class Memory) are examples of suitable fabric-attached memory devices, or FAMs. In some examples, a server appliance may be implemented as a fabric attached memory device. That is, a server comprising a low-cost CPU and DIMMs (Dual Inline Memory Modules), such as DDR DIMM modules, may be provided as a FAM. It will be appreciated that any suitable FAM devices may be utilized with the concepts herein.
[0043] In particular, the present disclosure is concerned with reducing the burden of memory management tasks on a host device, in fabric-attached memory contexts.
[0044] Examples of the present disclosure provide a host bridge device that enables loadstore requests to be transmitted across a fabric-attached memory system. In some exemplary systems, a host device is connected to a downstream bridge device (referred to herein as a host bridge). From the perspective of the host, the host bridge presents a plurality of memory windows as being accessible to the host. As described in more detail later, the host bridge presents the plurality of accessible memory windows to the host device as if there is a direct connection between the host and the memory. However, the host bridge is in fact connected to a downstream switching fabric network, comprising at least one switch and a plurality of discrete fabric-attached memory devices. It will be understood that switches may comprise a plurality of multiplexing circuits configured to route data between a first device in the network and a second device in the network. The fabric-attached memory devices may represent the true memory locations, at which the perceived memory windows are located.
[0045] For the purposes of the present description, the term 'memory window' refers to a chunk of memory, i.e., a range of memory addresses, which the host perceives is available to it, for the purpose of memory access operations. A host may perceive that it has access to a plurality of memory windows, each having a known size. The host may further have access to parameters indicative of a latency and bandwidth of each memory window. The parameters may be 'indicative' in that latency and bandwidth are not always entirely predictable, although approximate values can be estimated based on the type of memory and the route by which it is accessed. Certain approximate values may be referred to herein as 'static' bandwidth and latency values. Static latency and bandwidth parameters may indicate a latency or bandwidth based on an Ethernet fabric topology, for example, and may be calculated by a fabric controller at system boot. Telemetry data may be used to establish values of 'dynamic' latency and bandwidth parameters, as described further in the description that follows. Measurements may be conducted by the fabric controller to update the dynamic parameters, as latency and bandwidth may change as the system is running, based on factors such as operations being run on the system. The host may call the parameters when performing read / write operations. In examples where a DCD is implemented, as described below, collecting telemetry of this kind does not require a driver. The term 'memory window' does not imply any physical structure or location of the chunk of memory, but merely denotes an amount of memory that is known to be accessible, with an address useable by the host to access it.
[0046] By contrast, the term 'true memory location', refers to a physical device (i.e., a memory device) or a sub-component thereof, which provides physical memory locations which constitute the memory windows. True memory locations may have corresponding Ethernet destination addresses, to and from which data may be transmitted over an Ethernet switching fabric. Moreover, as detailed further in the description that follows, a memory window may have a true location at a FAM device attached downstream of the host bridge via a switching fabric.
[0047] The host is isolated from the system architecture downstream of the host bridge, and instead perceives the host bridge as an endpoint of the system; e.g., as if the host bridge is a directly connected memory device.
[0048] The present disclosure provides examples of how hardware and software aspects of the host bridge enable a system of the form described above. Examples also provide exemplary systems including the host bridge device, and explain how memory operations including setup, enumeration and memory accesses handling are managed in such fabric-attached memory systems. As described later, Memory accesses may be write requests (such as store operations) where data is written into a memory, or read responses (such as load operations) where data is read from a memory device.
[0049] CXL is an interconnect standard, which enables high-speed memory access and low-latency interconnection. Specification details for a first three major release families within the CXL standard (CXL 1.1, CXL 2.0, and CXL 3.1) are known.
[0050] By way of example, CXL attached memory devices may have a latency of between 170-250 ns. For context, main memory directly attached to a CPU may have a latency of, for example, 80-140 ns. It will be noted, however, that fabric-attached CXL memory may have a higher latency that the figure given above, dependent on factors such as a distance between the memory devices and the host.
[0051] CXL 1.1 further supports cache coherency, which is important in parallel and distributed computing applications. It will be understood that systems may still leverage caching techniques when the fabric-attached memory techniques described herein are implemented.
[0052] CXL 1.1 introduces three CXL protocols: CXL.io, CXL.cache, and CXL.mem. Particular reference is made herein to the CXL.mem protocol, which provides a host processor with access to device-attached memory using read / write commands. For example, the CXL.mem protocol may be used to issue load / store requests to fabric attached memory over a switch. As described above, a load request is a memory read access request to load a data item from memory. A store request is a memory write access request to store a data item into memory. Load and store requests are examples of memory access requests, specific to cache lines under the CXL.mem protocol. Whilst load / store requests may be issued in some examples, the present disclosure is not limited in this way. Examples herein will refer more broadly to read / write requests. Each memory access request has an address which identifies the memory location on which the access operates, either to write data to the memory (store) or read data from the memory (load). It is noted here that a write request also has payload in the form of data to be written to the memory location.
[0053] Memory pooling between devices may also be enabled with CXL 2.0. That is, memory resources are consolidated in a central system, and components in the system such as hosts and accelerators etc. may request allocations in the pool, as and when needed. The available memory is therefore shared among different hosts, which improves efficiency of memory usage and improves scalability.
[0054] CXL-enabled devices are grouped into three 'Types'. Type-1 devices are configured to leverage CXL.io and CXL.cache protocols, and may be configured to coherently access host memory. Type-2 devices are configured to leverage CXL.io, CXL.cache and CXL.mem protocols, and may be configured to coherently access host memory and allow the host to access memory of the Type-2 device. Type-3 devices are configured to leverage CXL.io and CXL.mem protocols, and may allow host to access and manage memory of the Type-3 device.
[0055] As discussed previously, CXL is an emerging technology, with many features being envisioned but not yet practically implemented. CXL 2.0 and 3.1 may, in future, enable configuration of systems that include accelerators and other types of device. By way of example, horizontal scaling of switches — i.e., switches linked to one or more further downstream switch — is also envisaged with CXL 3.1, enabling configuration of more complex topologies. CXL systems that provide memory sharing functionality are also envisaged with CXL 3.1, allowing a plurality of hosts to access a memory location atthe same time and still see up-to-date information at that location.
[0056] References herein to 'CXL devices' may be understood as being CXL-enabled; i.e., devices that are equipped to operate under the CXL protocols. A CXL-enabled device may, for example, comprise CXL logic implemented on a port, such that communications from other devices under CXL protocols in a CXL-based system may be interpretable by the receiving device.
[0057] Examples provided herein may use CXL devices connected to an Ethernet switching fabric. That is, the host bridge device may be connected between an upstream host and a downstream Ethernet switching fabric, to which a plurality of fabric attached memory devices are connected.
[0058] It will be understood that Ethernet is a standard networking technology that enables a wired communicative coupling of computer devices in a system. Systems configured to communicate via the Ethernet protocol may a have a 'star', or 'linear bus' topology, and communication may conform to the IEEE 802.3 group of standards. IEEE 802.3 is a collection of standards that define attributes of the physical layer and data link layers of wired Ethernet systems. In the examples described below, Ethernet switches will be understood to be operating in a lossless mode, as lossy transmission of read / write requests is not suitable. Examples of Ethernet switching provided herein may take place on the L2 (data link) layer of Ethernet. Routing of Ethernet packets across the ethernet Fabric is based on MAC addresses. Source and destination addresses are referred herein as SMAC and DMAC addresses respectively. Ethernet PEC (priority flow control) may be utilized to control flow of transmission when a transmitting device is transmitting data at a faster rate than a receiving device can receive it. The flow control mechanism may determine when more data can be transmitted. PFC may be used to isolate flow control on requests rather than responses and decouple the flow of request packets and response packets.
[0059] Examples of the present disclosure may, in future, be adapted to leverage features of Ultra Ethernet, which is an Ethernet-based communication stack architecture for high-performance networking. Ultra Ethernet is not yet widely available, though it is anticipated that better performance than present Ethernet fabrics may be realised using Ultra Ethernet.
[0060] Reference is made to Figure 1, which shows a highly schematic block diagram illustrating an exemplary fabric-attached memory system 100.
[0061] Figure 1 shows a host device 101, which may comprise any suitable computer device, such as a central processing unit (CPU), graphics processing unit (GPU), or data processing unit (DPU). Other architectures may also be implemented for carrying out data processing tasks. The host 101 comprises a software operating system (OS) that manages processes in the host and may perform memory and input / output (I / O) management tasks within the host device 101.
[0062] The system 100 of Figure 1 comprises a host bridge device 103. In some examples, the host bridge 103 device may comprise a suitably programmed processor or an application-specific integrated circuit (ASIC). However, other devices may be suitable. For example, FPGA (Field-Programmable Gate Array) devices may be suitably programmed to provide the functionalities described herein.
[0063] In some examples, the host bridge device may implement a dynamic capacity function. Dynamic Capacity is a feature of a CXL memory device that allows memory capacity to change dynamically without the need for resetting the device. A dynamic capacity device (DCD) is a CXL memory device that implements Dynamic Capacity. Unlike a traditional DPA (Device Physical Address) range that a CXL memory device may support, a Dynamic Capacity DPA range is subdivided into a plurality (e.g., 1 to 8) DC (dynamic capacity) regions, each of which is subdivided by the DCD into a number of fixed-size blocks, referred to as DC blocks.
[0064] In CXL systems where the CXL are directly accessible to the host, the host 101 executes code to implement a programming function which programs a maximum potential capacity utilizing one or more HDM (host-managed device memory) decoders to span the entire DPA range of all configured regions. The DCD controls the allocation of these DC blocks to the host and utilizes events to signal the host when changes to the allocation of these DC blocks occurs. The DCD communicates the state of these DC blocks through an Extent List that describes the starting DPA and length of all DC blocks the host can access.
[0065] DCD devices are not typically intended to be implemented upstream of a fabric. That is, DCD is typically a capability of an endpoint memory device. Examples herein leverage the functionality of the DCD to present memory windows to the host, wherein those memory windows are provided by FAMs connected via a switching fabric downstream of the DCD.
[0066] System architecture downstream of the host bridge device 103 is invisible to the host 101. Figure 1 shows an abstraction border 110, which represents a layer of isolation between the host device and the system architecture. The abstraction border 110 is not a physical attribute of the system 100 but merely represents a limit to the 'line of sight' of the host device 101. The host sees the host bridge device 103 as an endpoint of the system with a plurality of accessible memory windows. An endpoint may be understood as a device that has no further downstream connections and which provides suitable responses to the host based on its functionality. For example, an endpoint memory device would receive access requests and provide appropriate responses. CXL Class 3 devices are endpoint devices.
[0067] Processes lying below the abstraction border 110 of Figure 1 are controlled by fabric controller software 111. The fabric controller 111 may be implemented anywhere in the system. The fabric controller 111 is represented in a dashed box in Figure 1, indicating that no physical location of the controller 111 is implied by Figure 1. In some examples, however, it will be understood that the host 101 may not comprise the fabric controller 111; this is because the host is unaware that memory windows correspond to memory in FAMs 107 provided over an Ethernet fabric 105. The fabric controller 111 implements high level services such as resource management, fabric inventory management, and other services such as messaging, security, and logs. Processes below the abstraction border 110 may be said to be in a fabric controller domain of the system.
[0068] The host bridge device 103 has downstream connections to a switching fabric 105, and ultimately to one or more FAM device 107 that provide the memory windows that are accessible to the host 101. Two exemplary FAM devices 107 are shown in the example of Figure 1. The downstream connections of the host bridge device 103 to a switching fabric are invisible to the host.
[0069] The host bridge device 103 may be configured, amongst other things, to perform re-packetisation and routing operations. 'Routing operations' will be understood to comprise two paradigms: finding a correct address that defines a destination for the item being routed, and using the right address to route the item through switch or switches of a routing fabric. The host bridge device 103 performs a part of the routing operations that fall into the former category, finding the right address. The host bridge 103 is programmed, e.g., in registers, with a mapping of memory windows to Ethernet destination addresses, enabling the translation of addresses between the memory window format used by the host and the true Ethernet destination addresses (i.e., DMAC addresses) of memory devices in the Ethernet fabric. Transmission of data across the fabric, i.e., operations for using the right addresses, are performed by one or more Ethernet switch in the fabric 105 based on DMAC addresses. It will be understood that a data packet may be transmitted based on a DMAC address of a destination device, and an SMAC address of the device which issued the data packet may be embedded in the packet. Separate addressing information may be embedded in the packet to designate a particular part of the memory device that is being addressed. In response packets, the SMAC and DMAC addresses swap places, as the destination of a response packet is the sending device for outgoing packets. [00701 A data packet from the host 101 may be received at the host bridge device under a CXL protocol, whereas the switching fabric may operate under an Ethernet protocol. Re-packetisation for transmission of the data packet over the switching fabric 105 may be a responsibility of the host bridge device 103.
[0071] The Ethernet fabric 105 may comprise one or more switch. The fabric further comprises a fabric manager module 113, which is a software module representing an endpoint node for FAM devices 107 to communicate with. CXL-enabled FAM devices 107 generate error reporting and event logs, with errors being reported to the host 101. To support this, error reporting packets are synchronized by the fabric manager 113. A respective fabric manager module 113 may be implemented for each switch in the fabric. The fabric manager(s) 113 understand communications issued under CXL protocols. FAMs respond to read requests by transmitting payloads upstream, back to the host. In the event of errors in the memory device 107, response packets carry error information. Error information may be transmitted back to the requesting node, i.e., the host 101, but in some examples the fabric manager 113 may intervene. The fabric manager 113 may pick up a packet comprising error information, detect and extract error log information, and then assist in forwarding the error information to the host. That is, error reporting may be isolated from the host, but the host still knows about memory failures. More generally, the fabric manager 113 may be configured to assume responsibility for any suitable processing burden that may be offloaded from the host 101.
[0072] Figure 1 shows a connection 115 between the fabric controller 111 and the fabric manager 113. Fabric manager 113 and controller 111 tasks may be synchronized by communication over the connection 115.
[0073] Re-packetisation may comprise modification of CXL read / write requests such that a payload of the CXL request remains intact, but such that the packet conforms to an Ethernet protocol for routing across an Ethernet switching fabric. It will be understood that other data may be incorporated into a packet by the host bridge device during the re-packetisation modifications.
[0074] The term 'data tunneling' refers to wrapping a first data packet with a new packet header to form a second data packet, wherein the first packet is comprised within a payload of the second packet, then transmitting the second data packet and later unwrapping the second packet to retrieve the first data packet. In the present examples, additional routing information may be incorporated into a data packet such that it conforms to a different communication protocol, for subsequent transmission across the switching fabric to a FAM device.
[0075] In examples of the present disclosure, a first packet may be generated by the host 101 under a CXL protocol (with a CXL-compliant header). The host bridge device 103 may modify the first packet to form a second packet, the second packet having additional routing information compliant with a second (Ethernet) communication protocol.
[0076] The second data packet is configured in a format that is interpretable by a fabric switch (e.g., an Ethernet fabric switch) so that the packet can be routed to a correct memory location in the plurality of FAM devices that are connected downstream of the fabric switch.
[0077] Each FAM device 107 may be serviced by a respective memory bridge 109, configured such that when a packet such as the second packet is received (i.e., a packet comprising, in its payload, a tunneled CXL data packet), that CXL data packet may be unpacked, forwarded to the FAM 107, and the operation indicated by the CXL data packet (e.g., a read / write request) performed at the FAM device 107. The memory bridges 109 may comprise an ASIC, though alternative solutions will be appreciated.
[0078] In some examples, a memory bridge device 109 may serve a plurality of FAMs 107. That is, the memory bridge may fan out to multiple memory devices, which may each comprise a memory window. A MAC address may be defined for each memory window, and the memory bridge 109 may be configured to unpack a CXL data packet from an Ethernet packet and forward the unpacked packet to a correct one of the multiple memory devices, based on the MAC address.
[0079] In some examples, the memory bridge may unpack the CXL data packet and transmit to a memory device under a CXL protocol. In other examples, the memory bridge may transmit a data packet to a memory device under a non-CXL protocol. In such an example, the memory bridge 109 may be configured to perform repacketization operations to conform with the non-CXL protocol before transmitting downstream to the memory device. The memory bridge may be configured to convert to a format for forwarding to a DDR local address. By way of example, the memory bridge may convert to a format in accordance with a JEDEC (Joint Electron Device Engineering Council) DDR standard.
[0080] Additional architectures than the example of Figure 1 may be configured. For example, peer-to-peer (P2P) communication may be enabled. In some examples, a first host may have local memory which may comprise a first memory window accessible to a second host across an Ethernet fabric. The second host may equally have local memory comprising a second memory window, accessible to the first host across the fabric. In this example, a first and second host bridge device is configured for each of the first and second hosts respectively. The first host bridge device may present the second host memory windows to the first host device, and the second host bridge device may present the first host memory windows to the second host device. Similarly, the respective host bridge devices act as memory bridges, for unpacking ethernet packets and forwarding the unpacked packet to the respective host. It will be understood that in this example, each host may provide functionalities as a host and as a FAM, and the first and second bridges may act both as host bridges and as memory bridges (i.e., a host bridge and memory bridge may effectively be located on a same port of the Ethernet fabric).
[0081] In the example outlined above, the first host bridge and first host memory bridge may be provided on a same node of the system, i.e., on a same upstream port of an Ethernet switch in the fabric. However, initiator and responder activities associated with the host bridge and host memory bridges respectively may be handled by separate devices on the node, and the devices may be connected to the host via different CXL downstream ports on the host.
[0082] In some examples, a single device, e.g., a Type-2 CXL device, may comprise a combined host bridge and host memory bridge on the node. That is, in some examples, a device may be configured as a merged host bridge and host memory bridge with initiator and responder functionality.
[0083] Furthermore, in some examples a FAM device 107 may comprise a server appliance. The server appliance may comprise a low-cost CPU with an idle operating system. However, the server appliance may comprise one or more memory module (e.g., DIMM). In examples where a FAM device comprises a server appliance, a corresponding memory bridge 109 may be plugged into the server appliance. In such an example, the memory bridge may present as a CXL Type-1 device, using the CXL.cache protocol.
[0084] It will further be understood that the CXL.cache protocol may not be used at the host side. As described above, the host may operate under a CXL.mem protocol. However, a CXL.mem read request (for example) may be translated into a modified Ethernet packet by the host bridge 103 for transmission over the Ethernet fabric. The modified packet also comprises an Ethernet destination address, e.g., in the form of a DMAC address, rather than identifying a memory window. At the memory bridge 109, a data packet in accordance with the CXL.cache protocol may be generated and forwarded to the memory (typically DIMMs) via the local CPU of the server appliance.
[0085] Read / Write requests
[0086] To illustrate the above tunneling process, reference is made to Figures 2a and 2b. Figure 2a shows a highly schematic packet representation of a write request as it travels through the exemplary system of Figure 1.
[0087] Figure 2a shows a write request packet 210 having a header portion 211 and a payload 213. The data packet 210 is configured under a CXL protocol, and represents a data packet that may be transmitted by a host (e.g., host 101) for writing data in an accessible memory window. 'Store' requests are examples of write requests configured under a CXL protocol.
[0088] The CXL header 211 and payload 213 portions are represented simply as blocks in Figure 2a. However, it will be appreciated that in practice, these data packet portions 211, 213 comprise a plurality of machine-interpretable bit values. It will also be understood that different, and / or additional packet portions than those illustrated in Figure 2a may be present in a read request 210. For example, as described later, error mitigation techniques may be implemented by configuring additional portions of the data packet 210. That is, Ethernet packet formats support error bits. Other data packets than write requests may also be modified via the same process, for example Figure 2b (described later) illustrates the process for receiving data transmitted by a FAM device at the host, in response to a read request issued by the host.
[0089] The CXL header 211 may designate a memory window to be accessed, e.g., a memory window at which data in the payload 213 is to be written. In some examples, the host may request information regarding the parameters described above that indicate capacity (size), latency and bandwidth of each memory window. The host may determine a memory window to designate in the packet 210 based on the parameters.
[0090] In the present context, the data packet 210 issued by the host 101 designates a memory window presented to the host 101 by the host bridge device 103. The data packet 210 may further comprise memory offset data indicating an address sub-range within the memory window at which the data is to be written. As above, the memory window has a true memory location on a fabric attached memory device 107 downstream of the host bridge device 103, which is accessible via a switching fabric 105. The true memory location is concealed from the host by the host bridge device 103. Therefore, on its own, the CXL header 211 of the write request packet 210 may provide insufficient detail to precisely designate the true memory location of the designated memory window, as the host is ignorant of the fabric-attached memory structure. For example, the host device 101 does not directly designate an Ethernet destination address of a FAM device 107 in the CXL header 211, as the host does not perceive the memory windows to be provided by FAMs accessible over an Ethernet fabric.
[0091] So that the write request 210 from the host 101 reaches the true memory location of the memory window designated by the host 101, the host bridge device is configured to receive the write request 210, read the header 211 of the write request packet 210 to determine a designated memory window, determine therefrom an Ethernet destination address of a FAM at which the designated memory window is provided, and to construct a second, modified, data packet 220 that is to be transmitted across the switching fabric 105 to the true memory location, i.e., the FAM device 107. Registers in the host bridge device 103 may enable an address translation between memory window understood by the host, and the corresponding Ethernet destination address of a FAM device 107.
[0092] The second data packet 220 comprises a second header portion 221, e.g., an L2 Ethernet header, that designates an Ethernet destination address of a FAM device 107 which corresponds to the memory window designated, for example, in the header 211 of the first packet 210. The second header portion 221 may, where an Ethernet switching fabric 105 is implemented, have a structure that is compliant with Ethernet protocol. This enables an Ethernet switch to successfully route the second packet 220 to a correct FAM device 107, for example via a memory bridge 109 serving the FAM device 107. By way of example, the second header portion 221 may designate a DMAC address corresponding to a FAM 107 that provides the memory window. In the examples of Figures 2a and 2b, a pair of addresses (SMAC, DMAC) may be provided. The request comprises a DMAC address which becomes an SMAC address when a response packet is issued. In some examples, however, a response packet may not be transmitted in response to a write packet.
[0093] The second packet 220 further comprises a second payload 223, which may comprise the entire first packet 210. The second payload 223 may also be augmented with error information. The payload 223 may further comprise the memory offset data of the first data packet, such that the offset may be used at the FAM 107 to designate a particular address range within the FAM 107 that is to be accessed. Further still, where collect buffers are implemented, the payload 223 may comprise data for writing which is coalesced from multiple write requests issued by the host. The second payload 223 of the Ethernet packet 220 may embed a physical address of a recipient memory device. In some examples, a memory bridge 109 may serve a plurality of downstream FAMs 107. That is, the memory bridge 109 may fan out to multiple memory devices. In such an example, the ethernet payload 223 may embed a DMAC address of a particular one of the multiple memory devices, and the memory bridge may be configured to intervene to unpack the original packet 210, and direct the packet to the particular device based on the DMAC address. As described later, a DMAC address is used as an SMAC address when a response packet is issued.
[0094] A FAM device may be configured to operate under CXL protocols, and may comprise CXL logic configured to perform memory operations based on CXL instructions; i.e., instructions that conform to a CXL protocol, such as the CXL.mem protocol. The memory bridge 109 serving each FAM device 107 may comprise logic that is configured to receive the second data packet 220 and unpack the payload 223, which as described above comprises the first packet 210. By unpacking the payload 223 of the second packet 220, the FAM device retrieves the first data packet 210, which comprises the first (CXL) header 211 and the write request payload 213. The memory bridge 109 may be configured to perform the unpacking operation and forward the unpacked first packet 210 to the FAM 107. The FAM 107 may then read data in the header 211 of the unpacked first data packet and perform a write operation in accordance with the instructions provided in the packet 210, at a memory window designated by the header 211. It will be noted thatthe memory bridge 109 does not perform any address translation in respect of packets transmitted from the host to the FAM, because the unpacked packet 210 comprises routing information for reaching the FAM 107.
[0095] Write requests issued by the host may be posted, in that the host does not expect a response that acknowledges receipt of the write request. In the read example provided below, a response packet comprising the data that has been read itself provides acknowledgement of the read request.
[0096] Figure 2b shows a second example of packet transmission across the system of Figure 1, showing transmission of a read response travelling back to the host. The FAM is a CXL device configured to send read responses under a CXL protocol.
[0097] Figure 2b shows a read response data packet 230 comprising a CXL header 231 and a read response payload 233. The payload may comprise data that has been read from a FAM device in response to a request from the host.
[0098] To generate the read response packet 230, the host may issue a read request packet (not shown). A read request packet issued by the host may provide header information for accessing a particular memory window, but may comprise an empty payload field for receiving the requested data, or a payload comprising error information. Read requests issued by the host are subject to the same address translation processes as described with reference to Figure 2a, as they are transmitted by the host which is unaware of the FAM devices that provide the memory windows. The read request may arrive at the FAM device 107 with header information that defines an address of the host bridge device 103 that transmitted the request across the fabric.
[0099] The read response packet 230 may be transmitted by the FAM device to a memory bridge 109 serving that FAM 107. The memory bridge 109 may re-packet the response packet 230 into a modified Ethernet packet 240 for transmission back to the host, the modified Ethernet packet 240 having an Ethernet header 241 and a payload 243. The Ethernet payload 243 may comprise the entire response packet 230, but includes at least the data that has been read from memory at the FAM 107. Additional error information may be incorporated in the modified packet 240. The Ethernet payload 243 may embed a DMAC address of the host bridge device 103, based on an SMAC address embedded in a corresponding read request that prompted generation of the read response packet.
[0100] In some examples, wherein pre-fetch buffers are implemented, a modified Ethernet packet 240 may comprise a larger payload comprising coalesced read data from a plurality of read requests. Fewer packets, having larger payloads, may be transmitted across the fabric in this case.
[0101] The memory bridge 109 may map a local ID for the response to an address of the host bridge 103 that issued the read request. The header 241 of the modified response packet 240 may then be constructed with a DMAC address of the host bridge device 103, such that the packet may be correctly routed across the Ethernet fabric to the host bridge 103 for subsequent unpacking and relaying to the host 101 that initially requested the read data.
[0102] The modified Ethernet packet 240 arrives at the host bridge 103 and is unpacked to retrieve the read response packet 230. The host bridge 103 may access a transaction ID, which may be provided in the header 231 of the response packet 230. The transaction ID may indicate a request in response to which the response packet 230 has been received, and may enable the host bridge device to transmit the unpacked response packet 230 back to the host 101.
[0103] Figure 3 shows a highly schematic diagram of a memory space 300 accessible to a host device. The memory space 300 is segmented into a plurality of memory windows 301a-301f, wherein each window 301 represents a portion of the host-accessible memory space 300. As explained above, the term 'memory window' does not imply any physical structure or location of the memory, but merely denotes an abstract amount of memory that is known to be accessible to the host.
[0104] Each memory window 301 may have an associated bandwidth, latency, capacity and other properties. Two different memory windows 301 may have different such properties, dependent on such factors as a type / specification of memory device by which the memory window 301 is provided, a physical distance between the memory device and the host, and other such factors. It will also be understood that if more memory devices and more memory windows are added, properties such as latency and bandwidth may change.
[0105] Properties of each memory window 301 (e.g., latency, bandwidth, capacity etc.) may be accessible to an operating system of the host device, e.g., via registers of the host bridge device 103 that hold parameters defining size and an indication of expected bandwidth and latency. However, the host remains unaware of the physical devices and architectures that provide the memory windows 301. The properties of each memory window 301 may be accessible to the host via registers of the host bridge device 103. The registers may be accessed via CXL. Bandwidth and latency can be measured. Parameters indicative of bandwidth and latency may be determined at boot, or just after, based on a system topology, clock cycles, and a number of 'hops' packets take in the system. These parameters may be referred to herein as 'static' parameters. Telemetry measurements may also be taken while the system is running, to define values for 'dynamic' latency and bandwidth parameters. In some examples, static and / or dynamic latency and bandwidth parameters may not be stored at the host bridge 103. However, the parameters may be stored such that they are accessible to the host.
[0106] In Figure 3 memory windows 301a-301f are represented as blocks by way of schematic illustration of the concept of memory windows, where the size of each block indicates a memory capacity of each window 301. A memory window 301 may, by way of example, represent an amount of memory of the order of Gigabytes. In the example of Figure 3, each memory window 301a-301f is associated with a respective FAM device 107a-107f which accommodates that memory window 301. Each memory window 301 in Figure 3 is associated with a different FAM device 107. However, it will be understood that a FAM device may provide more than one memory window 301.
[0107] Figure 3 shows a mapping of memory windows 301 to FAMs 107. That is, memory windows are set up to point to a FAM device which may be located anywhere in the Ethernet fabric. In examples where a DCD is implemented, the memory is segmented into a set with crossing pointers P. That is, memory windows may not span a contiguous space as seen from a physical perspective. Each pointer P defines the relationship between a memory window and the FAM locations onto which that memory window may map.
[0108] Known techniques for pre-fetch buffering and collect buffering may optionally be implemented by the host bridge device 103. A buffer may be provided for each window.
[0109] In pre-fetch examples, read requests may be issued to consecutive addresses to bring in data, such as cache lines, from those addresses. Collect buffers operate according to the same principle, wherein data are preemptively collected at the collect buffers before a write request is received. The aim of pre-fetch / collect buffering is to preemptively read or write data in anticipation of that data being requested by the host. Pre-fetching or collecting can help to reduce wait times. The above techniques also allow for bigger, more efficient payloads with fewer headers, which reduces data overheads.
[0110] In some examples, the host bridge device, Ethernet switch, and FAM devices may form a memory subsystem in a broader computer system. For example, the memory subsystem may coexist in a computer system that also comprises a CXL subsystem and the host.
[0111] The pre-fetch and collect buffering examples above may be implemented in cases where a plurality of data packets received at the host bridge device from the host designate a common memory window, and wherein offset data in the plurality of data packets define a consecutive range of addresses within the window.
[0112] Reference is made now to Figure 4, which shows an exemplary computer system in which a host 101 is connected via CXL to plurality of devices. On a left-hand side of Figure 2, the host 101 is connected to a memory subsystem 410 comprising host bridge device 103 and a fabric attached memory system, all in accordance with Figure 1 and the above description thereof.
[0113] On a right-hand side of Figure 4, the host 101 is shown to be connected via CXL to a CXL subsystem 420. The CXL subsystem 420 of Figure 4 is a CXL switching fabric comprising a CXL switch 421 and a plurality of downstream CXL memory devices 423.
[0114] As explained with reference to Figure 1, the host 101 is unaware of the true structure of the memory subsystem beyond the abstraction border 110 (downstream of the host bridge device 103). However, the host does recognise the structure of the CXL subsystem 420.
[0115] Figure 5 shows a second exemplary computer system in which a host is connected to memory subsystems. In the example of Figure 5, the host 101 is connected only to a CXL switch 421. The CXL switch 421 is connected via a horizontal CXL connection 510 to a host bridge device 103 of a memory subsystem 410 in accordance with Figure 4.
[0116] The CXL subsystem 420 of Figures 4 and 5 may be configured to use data striping and redundancy techniques to distribute data across a plurality of memory devices. This is known as 'redundant array of independent memory' (RAIM). Reference is made to the applicant's earlier GB application no. GB2313537.9 (the contents of which are herein incorporated by reference), which describes the use of RAIM in fabric-attached memory. RAIM setup in the present examples may be a responsibility of fabric management software, for example at the fabric controller.
[0117] RAIM may also be implemented in FAMs across an Ethernet fabric in systems according to the examples above. That is, a host may issue a memory access request packet (i.e., a read / write request such as a load / store request) indicating a plurality of memory windows, and indicating that a RAIM striping technique is to be implemented across the plurality of memory windows. In write examples, a payload of the access request comprises a data item to be segmented and striped across the memory windows (though, in practice, across the FAMs).
[0118] In a write request example (e.g., a store operation), the host bridge may receive the data packet comprising an indication of a memory window, and comprising a data item to be written. The host bridge may comprise logic configured to segment the data item into a plurality of segments, and to generate a corresponding parity data segment. In RAIM examples, a memory window may be mapped to a plurality of FAM devices. The data item may be segmented such that the number of data item segments plus the parity data segment total the number of FAM devices mapped to the memory window indicated in the packet. The host bridge may then execute repacketization logic to access registers for mapping the memory window to Ethernet destination addresses corresponding to the plurality of FAMs, and may construct a plurality of modified packets in accordance with the Ethernet protocol. That is, a modified Ethernet packet may be constructed for each segment of the data item and for the parity data segment, wherein each modified Ethernet packet comprises an Ethernet destination address corresponding to a FAM associated with the memory window indicated in the original data packet from the host. The modified data packets may then be transmitted across the Ethernet fabric to the correct DMAC addresses.
[0119] Figure 6 shows a highly schematic diagram of a host bridge and Ethernet switch 1051 according to an example in which RAIM is enabled over the Ethernet fabric. The Ethernet switch 1051 may be comprised within an Ethernet switching fabric 105 and connected via an upstream port 621 of the Ethernet switch 1051 to the host bridge 103. The Ethernet switch may be a vanilla switch configured to perform Ethernet switching functionalities. It will be appreciated that the terms 'upstream' and 'downstream' port in respect of the Ethernet switch refer to a position of the port relative to the host. A distinction is made between ports on CXL switches, which may be programmed differently depending on whether they are configured as downstream ports or upstream ports.
[0120] RAIM logic 624 is implemented in a state machine in the host bridge 103, for example in the form of a simple pipeline. Functions such as parity generation and parity deployment, when writing and reading respectively, may be implemented by the RAIM logic state machine 624 in the host bridge 103. The host bridge 103 further comprises registers 622 for holding information concerning data allocation. Exemplary register structures are shown in Figure 7, which is described later herein. At a hardware level, the RAIM logic 624 may comprise buffers and sequencers.
[0121] The exemplary switch 1051 of Figure 6 comprises a plurality of downstream ports 626, e.g., ports 626a-626e. Each of the downstream ports 626a-626e is communicatively coupled to the upstream port 621 via respective crossbar links 628a-628e.
[0122] RAIM logic 624 at the host bridge 103 may be configured to implement re-packetisation of data packets received from the host, as described above.
[0123] In the example of Figure 6, each downstream port 626 is connected to a respective fabric-attached memory device 107, via a respective memory bridge device 109 in accordance with the description above. The downstream ports 626a-626e are respectively connected to memory devices 107a-107e.
[0124] In the present example, the function of the system is enhanced by enabling memory redundancy. In the example of Figure 6, memory devices 107a-107e are respectively labelled 'A', 'B', 'C', 'D', and 'P'. Each of the memory device 107a-107e may be mapped to a single memory window as viewed from a host perspective. The host may indicate a memory window mapped to the memory devices 107a-107e in examples where RAIM is enabled. The memory devices 107 labelled A-D may be designated for data striping. Memory device 107e, labelled 'P', may be designated for redundancy purposes for data security in the event of memory errors. In general terms, the RAIM logic 624 may be configured to generate parity data (or other error correction data) at the host bridge device 103 for redundancy purposes.
[0125] The memory sub-system of Figure 6 represents an Ethernet switched fabric network, wherein endpoint nodes in a wider system (e.g., the memory devices 107 and the host 101 of Figure 1) are interconnected via memory bridges 109, an Ethernet switch, e.g., switch 1051, and the host bridge 103. The physical links and software protocols by which nodes in the system communicate may enable low-latency memory access by the host. The memory bridges 109, as described previously herein, may be configured to identify memory errors.
[0126] Reference is now made to Figure 7, which shows a highly schematic diagram representing example registers for configuring a RAIM fabric attached memory system. The registers may be comprised within the host bridge 103.
[0127] Figure 7 represents a register for holding a RAIM-enable bit 710, and a plurality of member registers 720, each comprising a MAC address, wherein each port is connected to a corresponding fabric attached member device. For example, member registers 721-727 may respectively correspond to ports connected to fabric-attached members 1-4, and member register 729 corresponds to a port connected to fabric-attached member 'P'.
[0128] Members 1-4 may represent respective fabric-attached memory devices, such as devices 107a-107e as shown in Figure 6. Member P may represent an additional redundancy memory device, such as device 107e of Figure 6.
[0129] The RAIM enable register 710 comprises a single bit to indicated whether or not RAIM logic is to be implemented. Only a single bit is required to indicate this binary state. The RAIM enable register 710 may be a configuration register, configured at boot and typically not updated during runtime.
[0130] The size, in bits, of a member register 720 must offer at least as many states as there are downstream connections on the switch. That is, if a switch comprises 8 downstream ports — for connecting to 8 respective downstream devices — each register (of which there must be 8, one for each downstream device), must comprise at least 3 bits, because 23 = 8. This allows a unique identifier to be provided in each register, for each of the 8 downstream devices. This can be seen in Figure 7, wherein the member registers comprise three bits.
[0131] In some examples, other devices than the member devices 1-4 and P may be connected downstream of the switch. In such an example, additional member registers 720 may be provided in the host bridge 103 to accommodate the additional devices. Again, as the number of downstream connections increases, so too does the minimum size of the member registers, as each register needs to hold a value that can uniquely identify a port connected to a downstream device.
[0132] Figure 7 represents a memory system comprising four memory devices across which a data item is striped, and a single additional memory device to which redundancy data is stored. Such a system may be referred to as a '4+1' system.
[0133] It will be appreciated by those skilled in the art that the registers represented in Figure 7 may be altered to accommodate more, or fewer, fabric attached memory devices, with varying numbers of redundant memory devices. For example, a 3+1, 4+2, or 5+1 system may be configurable.
[0134] As described previously herein, collect buffers may be implemented in some examples to enable larger payloads in modified Ethernet packets. In RAIM examples, collect buffers may also be implemented. For example, segmentation techniques at the host bridge 103 may be applied to buffered data, wherein the data item to be segmented at the host bridge 103 is a combined payload of data packets received at the buffer from the host. As described above, collect buffers may be implemented in examples where plural host packets indicate consecutive addresses in a single memory window.
[0135] It will be appreciated that the above embodiments have been disclosed by way of example only. Other variants or use cases may become apparent to a person skilled in the art once given the disclosure herein. The scope of the present disclosure is not limited by the above-described embodiments, but only by the accompanying claims.
Claims
1. A host bridge device for use in accessing a fabric-attached memory system, the host bridge device comprising:an upstream port connectable to a host device and configured to operate under a CXL communication protocol to transmit and receive data between the host bridge device and the host;a downstream port connectable to a switched fabric network comprising a fabric switch and at least one fabric-attached memory device, the downstream port configured to operate under an Ethernet communication protocol to transmit and receive data between the host bridge device and the switched fabric network;at least one register storing memory address information for the at least one fabric-attached memory device of the switched fabric network, the memory address information mapping memory windows of a plurality of memory windows to Ethernet destination addresses; andpacketization logic configured to:receive, via the upstream port, a first data packet comprising an indication of a memory window;construct, based on the indication of a memory window and the memory address information, a modified data packet in accordance with the Ethernet communication protocol, the modified data packet comprising an indication of an Ethernet destination address corresponding to the memory window of the first packet, and a payload that comprises a payload of the first data packet; andtransmit the modified data packet via the downstream port.
2. The host bridge of claim 1, further comprising registers storing, for each memory window of the plurality of memory windows, one or more of:a dynamic bandwidth parameter indicating a bandwidth associated with a fabric attached memory device, wherein the dynamic bandwidth parameter is determined based on a measurement of bandwidth pertaining to the fabric attached memory device; anda static bandwidth parameter indicating a bandwidth associated with a fabric-attached memory device, wherein the static bandwidth parameter is determined based on an Ethernet fabric topology.
3. The host bridge of any preceding claim, further comprising registers storing, for each memory window of the plurality of memory windows, one or more of:a dynamic latency parameter indicating a latency associated with the fabric attached memory device, wherein the dynamic latency parameter is determined based on a measurement of latency pertaining to the fabric attached memory device; anda static latency parameter indicating a latency of the fabric-attached memory device, wherein the static bandwidth and static latency parameters are determined based on an Ethernet fabric topology.
4. The host bridge device of claim 2 or 3, wherein the host bridge device is configured to receive, from the host via the upstream port, a request for a value of at least one of the static bandwidth parameter and the dynamic bandwidth parameter and, responsive to the request, to transmit the value of the at least one parameter via the upstream port.
5. The host bridge device of claim 3 or 4, wherein the host bridge device is configured to receive, from the host via the upstream port, a request for a value of at least one of the static latency parameter and the dynamic latency parameter associated with a first memory window and, responsive to the request, to transmit the value of the at least one parameter via the upstream port.
6. The host bridge device of any preceding claim, wherein the modified data packet is a layer-2 Ethernet packet, and wherein the indication of an Ethernet destination address comprises a DMAC address.
7. The host bridge device of any preceding claim, further configured to:receive, via the downstream port, a read data packet in accordance with the Ethernet protocol, the read data packet comprising read data that is read from a fabric-attached memory device;unpack the read data from a payload of the read data packet and identify a transaction identifier in a header of the read data packet;construct a host read packet in accordance with the CXL protocol, the host read packet comprising the read data; andbased on the transaction identifier, transmit the packet to the host via the upstream port of the host bridge device.
8. The host bridge device of claim 7, wherein the read data packet comprises error information pertaining to the fabric attached memory device.
9. The host bridge device of any preceding claim, which implements a CXL Dynamic Capacity function.
10. The host bridge device of any preceding claim, comprising a collect buffer and configured to:receive plural first data packets comprising respective indications of a same memory window, wherein respective ones of the plural first data packets comprise consecutive addresses within the same memory window; andconstruct a modified data packet comprising an indication of an Ethernet destination address corresponding to the memory window of the plural first packets, and a payload that comprises respective payloads of the plural first data packets.
11. The host bridge device of any preceding claim, wherein the upstream port is connectable to the host via a CXL switch provided between the host bridge device and the host.
12. A computer system, comprising:a host configured to operate under a CXL protocol;a host bridge device in accordance with any of claims 1-9, connected to the host via an upstream port of the host bridge device;an Ethernet fabric comprising at least one fabric switch, the Ethernet fabric connected to the host bridge device via a downstream port of the host bridge device;at least one fabric-attached memory device connected downstream of the Ethernet fabric; andfor each fabric-attached memory device, a corresponding memory bridge configured to receive a modified data packet from the host bridge device in accordance with the Ethernet protocol, to unpack a first data packet from the modified packet, and to transmit the first data packet to the corresponding fabric-attached memory device.
13. The system of claim 12, wherein a first fabric attached memory device and a first corresponding memory bridge are comprised within a server appliance connected downstream of the Ethernet fabric.
14. The system of claim 12, wherein the host is a first host, and the host bridge device is a first host bridge, the system further comprising:a second host configured to operate under a CXL protocol, the second host configured to operate as a fabric attached memory device comprising host memory;a second host bridge device in accordance with any of claims 1-9, connected to the second host via an upstream port of the second host bridge device, wherein the second host bridge is connected to the Ethernet fabric via a first port in the Ethernet fabric; anda host memory bridge connected to the Ethernet fabric via the first port in the Ethernet fabric and connected to the second host via a downstream port of the second host.
15. The system of claim 14, wherein the second host bridge device and the host memory bridge are connected to respective CXL ports on the second host.
16. The system of claim 14, wherein the second host bridge device and the host memory bridge are implemented by a single initiator-responder device connected to the Ethernet fabric via the first port, the initiator-responder device comprising a Type-2 CXL device.
17. The system of any of claims 12-16, further comprising:a CXL switch provided between the host bridge device and the host, wherein the upstream port of the host bridge device is connectable to the host via the CXL switch.
18. The system of claim 12, wherein the host bridge is configured to receive a first data packet comprising a memory access request, an indication of a memory window, the memory window corresponding to a plurality of downstream fabric attached memory devices,, wherein the Ethernet fabric comprises an Ethernet switch comprising:a first upstream port connected to the host via the host bridge device;a plurality of downstream ports, each of which is connected to a respective downstream fabric-attached memory device; andwherein the host bridge further comprises logic configured to:receive the first data packet from the hostsegment a data item of the first data packet into a plurality of segments, including a parity segment;construct, based on the indication of the memory window, a plurality of modified data packets in accordance with the Ethernet communication protocol, each modified data packet comprising:an indication of an Ethernet destination address corresponding to a respective fabric attached memory device associated with the memory window of the first packet, anda payload that comprises a segment of the first data item; andtransmit the modified data packets via the downstream port to an Ethernet switch of the ethernet fabric.
19. The system of claim 12, wherein each memory bridge is configured to unpack a first data packet from the modified packet, and to transmit the first data packet to the corresponding fabric-attached memory device based on one of a CXL protocol, or a JEDEC protocol.
20. A method comprising:receiving, from a host device via an upstream port of a host bridge device, a first data packet comprising an indication of a memory window, the first data packet configured in accordance with a CXL protocol;accessing at least one register of the host bridge device storing memory address information for the at least one fabric-attached memory device of a switched fabric network, the memory address information mapping memory windows of a plurality ofmemory windows to corresponding Ethernet destination addresses of the switched fabric network;constructing, by packetization logic of the host bridge device based on the indication of a memory window and the memory address information, a modified data packet in accordance with an Ethernet communication protocol, the modified data packet comprising an indication of an Ethernet destination address of a fabric attached memory device of the switched fabric network, the Ethernet destination address corresponding to the memory window indicated in the first data packet, wherein a payload of the modified data packet comprises a payload of the first data packet; andtransmitting the modified data packet via a downstream port of the host bridge device.
21. The method of claim 20, further comprising:receiving, at a memory bridge device via the Ethernet switched fabric network, the modified data packet from the host bridge device in accordance with the Ethernet protocol; andunpacking, by the memory bridge device, a first data packet from the modified packet; andtransmitting the first data packet to a corresponding fabric-attached memory device indicated by the Ethernet destination address in the modified packet.
22. Transitory or non-transitory computer readable media embodying computer readable instructions which, when executed by one or more processor of one or more computer device cause the one or more processor to implement a method in accordance with claim 20 or 21.38
Citation Information
Patent Citations
Data transmission method and device thereof, storage medium, processor and electronic equipment
CN113645258A
Compute express link over ethernet in composable data centers
US11632337B1
Compute Express Link™ (CXL) Over Ethernet (COE)
US20230385223A1