System with Cache-Coherent Memory and Server Link Switch
Through the combined architecture of cache coherent switches and server-link switches, the CXL protocol and RDMA technology are used to solve the problem of low memory resource management efficiency in the server system, efficient memory access and data interaction are achieved, and system performance and bandwidth are improved.
Patent Information
- Application Number
- CN202110585663.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-18
- Filing Date
- 2021-05-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-05-27
AI Technical Summary
In the prior art, the management efficiency of memory resources in the server system is low, making it difficult to achieve efficient memory access and data interaction, especially in data transmission between different servers, where there are delay and bandwidth limitations.
The combined architecture of cache coherent switches and server-linked switches is adopted. Through the computing Quick Links (CXL) protocol, the aggregation and virtualization of memory modules between multiple servers is realized, remote direct memory access (RDMA), and the data routing and address conversion are enhanced by CXL switches, reducing the participation of processing circuits.
It improves the efficiency and latency of memory access, supports high bandwidth and low latency data transmission, realizes efficient memory resource management and data interaction between different servers, and reduces the load on processing resources.
Smart Images

Figure CN113746762B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the priority and benefit of U.S. Provisional Application No. 63 / 031,508, filed on May 28, 2020, entitled "EXTENDING MEMORY ACCESSES WITH NOVEL CACHE COHERENCE CONNECTS"; the priority and right of U.S. Provisional Application No. 63 / 031,509, filed on May 28, 2020, entitled "POOLING SERVER MEMORY RESOURCES FOR COMPUTE EFFICIENCY"; the priority and benefit of U.S. Provisional Application No. 63 / 068,054, filed on August 20, 2020, entitled "SYSTEM WITH CACHE - COHERENT MEMORY AND SERVER - LINKING SWITCH FIELD"; and the priority and benefit of U.S. Provisional Application No. 63 / 057,746, filed on July 28, 2020, entitled "DISAGGREGATED MEMORY ARCHITECTURE WITH NOVEL INTERCONNECTS", the entire contents of all of these applications are incorporated herein by reference. Technical Field
[0003] One or more aspects of embodiments in accordance with the present disclosure relate to computing systems, and more particularly, to systems and methods for managing memory resources in a system including one or more servers. Background Art
[0004] This background art section is intended to provide context only, and the disclosure of any embodiments or concepts in this section does not constitute an admission that the embodiments or concepts are prior art.
[0005] Some server systems may include a collection of servers connected by network protocols. Each server in such a system may include processing resources (e.g., processors) and memory resources (e.g., system memory). In some cases, it may be advantageous for the processing resources of one server to access the memory resources of another server, and it may be advantageous for such access to occur while minimizing the processing resources of either server.
[0006] Accordingly, there is a need for improved systems and methods for managing memory resources in a system that includes one or more servers. SUMMARY OF THE INVENTION
[0007] In some embodiments, a data storage and processing system includes a plurality of servers connected by a server link switch. Each server may include one or more processing circuits, a system memory, and one or more memory modules connected to the processing circuits through a cache coherent switch. The cache coherent switch may be connected to the server link switch, and it may include a controller (e.g., a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC)) that provides enhanced capabilities for it. These capabilities may include virtualizing the memory modules and enabling the switch to store data in the memory modules using an underlying technology that is well-suited to the storage requirements of the data (e.g., latency, bandwidth, or persistence). As a result of storage needs having been sent by the processing circuits to the cache coherent switch, or as a result of monitoring access patterns, the cache coherent switch may receive these needs. The enhanced capabilities may also include enabling a server to interact with the memory of another server without having to access a processor such as a central processing unit (CPU) (e.g., by performing remote direct memory access (RDMA)).
[0008] According to an embodiment of the present invention, there is provided a system including: a first server including a storage program processing circuit, a cache coherent switch, and a first memory module; a second server; and a server link switch connected to the first server and connected to the second server, wherein: the first memory module is connected to the cache coherent switch, the cache coherent switch is connected to the server link switch, and the storage program processing circuit is connected to the cache coherent switch.
[0009] In some embodiments, the server link switch includes a Peripheral Component Interconnect Express (PCIe) switch.
[0010] In some embodiments, the server link switch includes a Compute Express Link (CXL) switch.
[0011] In some embodiments, the server link switch includes a Top-of-Rack (ToR) CXL switch.
[0012] In some embodiments, the server link switch is configured to discover the first server.
[0013] In some embodiments, the server link switch is configured to restart the first server.
[0014] In some embodiments, the server link switch is configured to cause the cache coherent switch to disable a first memory module.
[0015] In some embodiments, the server link switch is configured to send data from a second server to a first server and perform flow control on the data.
[0016] In some embodiments, the system further includes a third server connected to the server link switch, wherein: the server link switch is configured to receive a first packet from the second server, receive a second packet from the third server, and send the first packet and the second packet to the first server.
[0017] In some embodiments, the system further includes a second memory module connected to the cache coherent switch, wherein the first memory module includes volatile memory and the second memory module includes persistent memory.
[0018] In some embodiments, the cache coherent switch is configured to virtualize the first memory module and the second memory module.
[0019] In some embodiments, the first memory module includes flash memory, and the cache coherent switch is configured to provide a flash translation layer to the flash memory.
[0020] In some embodiments, the first server includes an expansion slot adapter connected to an expansion slot of the first server, the expansion slot adapter includes a cache coherent switch and a memory module slot, and the first memory module is connected to the cache coherent switch through the memory module slot.
[0021] In some embodiments, the memory module slot includes an M.2 slot.
[0022] In some embodiments: the cache coherent switch is connected to the server link switch through a connector, and the connector is on the expansion slot adapter.
[0023] According to an embodiment of the present invention, a method for performing remote direct memory access in a computing system is provided. The computing system includes: a first server, a second server, a third server, and a server link switch connected to the first server, connected to the second server, and connected to the third server. The first server includes a storage program processing circuit, a cache coherent switch, and a first memory module. The method includes receiving, by the server link switch, a first packet from the second server, receiving, by the server link switch, a second packet from the third server, and sending the first packet and the second packet to the first server.
[0024] In some embodiments, the method further includes receiving, by a cache coherent switch, a Remote Direct Memory Access (RDMA) request, and sending, by the cache coherent switch, an RDMA response.
[0025] In some embodiments, receiving the RDMA request includes receiving the RDMA request via a server link switch.
[0026] In some embodiments, the method further includes receiving, by a cache coherent switch, a read command for a first memory address from a storage program processing circuit, converting, by the cache coherent switch, the first memory address to a second memory address, and retrieving, by the cache coherent switch, data at the second memory address from a first memory module.
[0027] According to an embodiment of the present invention, a system is provided, which includes: a first server, including a storage program processing circuit, a cache coherent switching mechanism, and a first memory module; a second server; and a server link switch connected to the first server and connected to the second server, wherein: the first memory module is connected to the cache coherent switching mechanism, the cache coherent switching mechanism is connected to the server link switch, and the storage program processing circuit is connected to the cache coherent switching mechanism. Description of the Drawings
[0028] The drawings provided herein are only for the purpose of illustrating certain embodiments; other embodiments that may not be explicitly shown are not excluded from the scope of the present disclosure.
[0029] These and other features and advantages of the present disclosure will be recognized and understood with reference to the specification, claims, and drawings, wherein:
[0030] Figure 1A is a block diagram of a system for attaching a memory resource to a computing resource using a cache coherent connection according to an embodiment of the present disclosure;
[0031] Figure 1B is a block diagram of a system for attaching a memory resource to a computing resource using a cache coherent connection and employing an expansion slot adapter according to an embodiment of the present disclosure.
[0032] Figure 1C is a block diagram of a system for aggregating memories using an Ethernet ToR switch according to an embodiment of the present disclosure;
[0033] Figure 1D is a block diagram of a system for aggregating memories using an Ethernet ToR switch and an expansion slot adapter according to an embodiment of the present disclosure;
[0034] Figure 1EBlock diagram of a system for aggregating memories according to an embodiment of the present disclosure;
[0035] Figure 1F Block diagram of a system for aggregating memories using an expansion slot adapter according to an embodiment of the present disclosure;
[0036] Figure 1G Block diagram of a system for decomposing a server according to an embodiment of the present disclosure;
[0037] Figure 2A According to an embodiment of the present disclosure for Figures 1A - 1G Flowchart of an example method for a bypass processing circuit to perform a Remote Direct Memory Access (RDMA) transfer according to the embodiment shown;
[0038] Figure 2B According to an embodiment of the present disclosure for Figures 1A - 1D Flowchart of an example method for performing an RDMA transfer with the participation of a processing circuit according to the embodiment shown;
[0039] Figure 2C According to an embodiment of the present disclosure for Figure 1E and Figure 1F Flowchart of an example method for performing an RDMA transfer through a Compute Express Link (CXL) switch according to the embodiments shown in
[0040] Figure 2D According to an embodiment of the present disclosure for Figure 1G Flowchart of an example method for performing an RDMA transfer through a CXL switch according to the embodiment shown. Detailed Description of Specific Embodiments
[0041] The detailed description set forth below in connection with the accompanying drawings is intended as a description of exemplary embodiments of systems and methods for managing memory resources provided in accordance with the present disclosure, and is not intended to represent the only forms in which the present disclosure may be constructed or utilized. The description sets forth the features of the present disclosure in connection with the illustrated embodiments. However, it will be understood that the same or equivalent functions and structures may be achieved by different embodiments, which are also intended to be covered within the scope of the present disclosure. As indicated elsewhere herein, like element numbers are intended to indicate like elements or features.
[0042] Peripheral Component Interconnect Express (PCIe) can refer to a computer interface that can have relatively high and variable latency, which may limit the usefulness of the interface in establishing a connection to memory. CXL is an open industry standard for communicating over PCIe 5.0, which can provide a fixed, relatively short packet size, and as a result, can be capable of providing relatively high bandwidth and relatively low fixed latency. Thus, CXL can be capable of supporting cache coherence and can be well-suited for establishing a connection to memory. CXL can also be used to provide connectivity between a host and accelerators, memory devices, and network interface circuits (or "network interface controllers" or "network interface cards" (NICs)) in a server.
[0043] Cache coherence protocols such as CXL can also be used for heterogeneous processing in, for example, scalar, vector, and buffer memory systems. CXL can be used to provide a cache coherence interface by utilizing channels, retimers, the PHY layer of the system, the logical aspects of the interface, and the protocol in PCIe 5.0. The CXL transaction layer can include three multiplexed sub-protocols that run simultaneously on a single link and can be referred to as CXL.io, CXL.cache, and CXL.memory. CXL.io can include I / O semantics that can be similar to PCIe. CXL.cache can include cache semantics, and CXL.memory can include memory semantics; both the cache semantics and the memory semantics can be optional. Like PCIe, CXL can support (i) raw widths of x16, x8, and x4 that can be divisible, (ii) data rates of 32GT / s that can be reduced to 8GT / s and 16GT / s, 128b / 130b, (iii) 300W (75W in an x16 connector), and (iv) plug-and-play. To support plug-and-play, a PCIe or CXL device link can start training in Gen1 in PCIe, negotiate CXL, complete Gen1-5 training, and then start CXL transactions.
[0044] In some embodiments, in a system including multiple servers connected together via a network, as discussed further below in detail, using CXL connections for a collective or "pool" of memories (e.g., a certain amount of memories, including multiple memory units connected together) can provide various advantages. For example, a CXL switch that has further capabilities in addition to providing packet switching functionality for CXL packets (referred to herein as an "enhanced-capability CXL switch") can be used to connect the aggregation of memories to one or more central processing units (CPUs) (or "central processing circuits") and to one or more network interface circuits (which may have enhanced capabilities). Such a configuration can enable the following: (i) the aggregation of memories includes various types of memories with different characteristics; (ii) the enhanced-capability CXL switch virtualizes the aggregation of memories and stores data with different characteristics (e.g., access frequency) in the appropriate type of memory; (iii) the enhanced-capability CXL switch supports remote direct memory access (RDMA), such that RDMA can be performed with little or no participation of the processing circuit of the server. As used herein, "virtualizing" a memory means performing memory address translation between the processing circuit and the memory.
[0045] The CXL switch can (i) support memory and accelerator disaggregation through single-stage switching; (ii) enable resources to go offline and online between domains based on demand, which can enable time-division multiplexing across domains; and (iii) support virtualization of downstream ports. CXL can be employed to implement an aggregated memory, which can enable one-to-many and many-to-one switching (e.g., it can be capable of (i) connecting multiple root ports to one endpoint, (ii) connecting one root port to multiple endpoints, or (iii) connecting multiple root ports to multiple endpoints), and the aggregated devices are divided into multiple logical devices in some embodiments, each logical device having a corresponding LD-ID (logical device identifier). In such embodiments, the physical device can be divided into multiple logical devices, and each logical device is visible to the corresponding initiator. The device can have one physical function (PF) and multiple (e.g., 16) isolated logical devices. In some embodiments, the number of logical devices (e.g., the number of partitions) can be limited (e.g., capped at 16), and there can also be a control partition (which can be for controlling the physical function of the device).
[0046] In some embodiments, a fabric manager may be employed to (i) perform device discovery and virtual CXL software creation, and (ii) bind virtual ports to physical ports. Such a fabric manager may operate over a connection on the SMBus sideband. The fabric manager may be implemented in hardware, software, firmware, or a combination thereof, and it may reside, for example, in a host, in one of the memory modules 135, in the enhanced-capability CXL switch 130, or elsewhere in the network. The fabric manager may issue commands, which may include commands issued over the sideband bus or over the PCIe tree.
[0047] Referring Figure 1A , in some embodiments, a server system includes a plurality of servers 105 connected together via a top-of-rack (ToR) Ethernet switch 110. Although the switch is described as using the Ethernet protocol, any other suitable network protocol may be used. Each server includes one or more processing circuits 115, each of which is connected to (i) system memory 120 (e.g., double data rate (version 4) (DDR4) memory or any other suitable memory), (ii) one or more network interface circuits 125, and (iii) one or more CXL memory modules 135. Each processing circuit 115 may be a stored-program processing circuit, such as a central processing unit (CPU) (e.g., an x86 CPU), a graphics processing unit (GPU), or an ARM processor. In some embodiments, the network interface circuit 125 may be embedded in one of the memory modules 135 (e.g., on the same semiconductor chip as one of the memory modules 135, or in the same module as one of the memory modules 135), or the network interface circuit 125 may be separately packaged from the memory module 135.
[0048] As used herein, a "memory module" is a package that includes one or more memory dies (e.g., a package that includes a printed circuit board and components connected thereto, or an attachment to a printed circuit board), and each memory die includes a plurality of memory cells. Each memory die or each in a set of memory dies can be in a package (e.g., an epoxy molding compound (EMC) package) that is soldered to (or connected to the printed circuit board of the memory module through a connector) the printed circuit board of the memory module. Each memory module 135 can have a CXL interface and can include a controller 137 (e.g., an FPGA, an ASIC, a processor, etc.) that is configured to translate signals, such as signals suitable for a memory technology of the memory in the memory module 135, between CXL packets and a memory interface of the memory die. As used herein, the "memory interface" of a memory die is an interface inherent to the technology of the memory die, e.g., in the case of DRAM, the memory interface can be word lines and bit lines. The memory module can also include a controller 137 that can provide enhanced capabilities as described in further detail below. The controller 137 of each memory module 135 can be connected to the processing circuitry 115 through a cache coherent interface (e.g., through the CXL interface). The controller 137 can also bypass the processing circuitry 115 to facilitate data transfer (e.g., RDMA requests) between different servers 105. The ToR Ethernet switch 110 and the network interface circuitry 125 can include an RDMA interface to facilitate RDMA requests (e.g., the ToR Ethernet switch 110 and the network interface circuitry 125 can provide hardware offloading or hardware acceleration of RDMA over Converged Ethernet (RoCE), InfiniBand, and iWARP packets) between CXL memory devices on different servers.
[0049] The CXL interconnect in the system can conform to a cache coherent protocol, such as the CXL 1.1 standard, or in some embodiments, conform to the CXL 2.0 standard, conform to a future version of CXL, or any other suitable protocol (e.g., a cache coherent protocol). The memory module 135 can be directly attached to the processing circuitry 115 as shown, and the top-of-rack Ethernet switch 110 can be used to scale the system to a larger size (e.g., having a greater number of servers 105).
[0050] In some embodiments, each server can be populated with a plurality of directly attached CXL-attached memory modules 135, as Figure 1AAs shown. Each memory module 135 may expose a set of base address registers (BARs) to the host's basic input / output system (BIOS) as memory ranges. One or more of the memory modules 135 may include firmware to transparently manage its memory space behind the host OS mapping. Each memory module 135 may include one or a combination of memory technologies, including for example (but not limited to) dynamic random access memory (DRAM), non-volatile (NAND) flash memory, high bandwidth memory (HBM), and low power double data rate synchronous dynamic random access memory (LPDDR SDRAM) technology. Each memory module 135 may also include a cache controller or separate respective disaggregation controllers for memory devices of different technologies (memory module 135 for combining several memory devices of different technologies). Each memory module 135 may include different interface widths (x4 - x16) and may be constructed according to various relevant specifications (e.g., U.2, M.2, half-height half-length (HHHL), full-height half-length (FHHL), E1.S, E1.L, E3.S, and E3.H).
[0051] In some embodiments, as described above, the enhanced-capability CXL switch 130 includes an FPGA (or ASIC) controller 137 and provides additional functionality in addition to switching CXL packets. The controller 137 of the enhanced-capability CXL switch 130 may also act as a management device for the memory modules 135 and assist in host control plane processing, and it may enable rich control semantics and statistics. The controller 137 may include an additional "backdoor" (e.g., 100 Gigabit Ethernet (GbE)) network interface circuit 125. In some embodiments, the controller 137 presents itself as a CXL type 2 device to the processing circuit 115, which enables the issuance of cache invalidation instructions to the processing circuit 115 after receiving a remote write request. In some embodiments, the DDIO technology is enabled, and the remote data is first pulled into the last-level cache (LLC) of the processing circuit and later written (from the cache) to the memory module 135. As used herein, a "type 2" CXL device is a device that can initiate transactions and implement optional coherent caches and host-managed device memory, and the transaction types applicable to such a device include all CXL.cache and all CXL.mem transactions.
[0052] As described above, one or more of the memory modules 135 may include persistent memory or "persistent storage" (i.e., storage in which data is not lost when the external power is disconnected). If the memory module 135 is presented as a persistent device, the controller 137 of the memory module 135 may manage the persistent domain. For example, it may store persistent storage data identified by the processing circuitry 115 (e.g., as a result of an application invoking a corresponding operating system function) as requiring persistent storage. In such an embodiment, the software API may cache and flush data to the persistent storage.
[0053] In some embodiments, direct memory transfer from the network interface circuit 125 to the memory module 135 is enabled. Such a transfer may be a one-way transfer to a remote memory for fast communication in a distributed system. In such an embodiment, the memory module 135 may expose the hardware details to the network interface circuit 125 in the system to enable faster RDMA transfers. In such a system, two scenarios may occur depending on whether data direct I / O (DDIO) of the processing circuitry 115 is enabled or disabled. DDIO may enable direct communication between the Ethernet controller or Ethernet adapter and the cache of the processing circuitry 115. If DDIO of the processing circuitry 115 is enabled, the target of the transfer may be the last-level cache of the processing circuitry, and the data may then be automatically flushed from the last-level cache to the memory module 135. If DDIO of the processing circuitry 115 is disabled, the memory module 135 may operate in device bias mode to force accesses that are to be directly received by the destination memory module 135 (without using DDIO). A network interface circuit 125 capable of RDMA with a host channel adapter (HCA), buffers, and other processing may be used to enable such RDMA transfers, which may bypass the target memory buffer transfers that may exist in other modes of RDMA transfers. For example, in such an embodiment, the use of bounce buffers (e.g., buffers in a remote server when the final destination in memory is in an address range not supported by the RDMA protocol) may be avoided. In some embodiments, RDMA uses another physical media option in addition to Ethernet (e.g., for switches configured to handle other network protocols). Examples of server-to-server connections that may enable RDMA include (but are not limited to) InfiniBand, RDMA over Converged Ethernet (RoCE) (which uses the Ethernet User Datagram Protocol (UDP)), and iWARP (which uses the Transmission Control Protocol / Internet Protocol (TCP / IP)).
[0054] Figure 1B is shown in connection with Figure 1AA system similar to the system, where the processing circuit 115 is connected to the network interface circuit 125 through the memory module 135. The memory module 135 and the network interface circuit 125 are on the expansion slot adapter 140. Each expansion slot adapter 140 can be inserted into an expansion slot 145 (e.g., an M.2 connector) on the motherboard of the server 105. In this way, the server can be any suitable (e.g., industry standard) server modified by installing the expansion slot adapter 140 in the expansion slot 145. In such an embodiment, (i) each network interface circuit 125 can be integrated into a corresponding one of the memory modules 135, or (ii) each network interface circuit 125 can have a PCIe interface (the network interface circuit 125 can be a PCIe endpoint (i.e., a PCIe slave device), such that the processing circuit 115 (which can operate as a PCIe master device or "root port") to which it is connected can communicate with it through a root port to endpoint PCIe connection, and the controller 137 of the memory module 135 can communicate with it through a peer-to-peer PCIe connection).
[0055] According to an embodiment of the present invention, a system is provided that includes a first server. The first server includes a storage program processing circuit, a first network interface circuit, and a first memory module. The first memory module includes a first memory die and a controller. The controller is connected to the first memory die through a memory interface, connected to the storage program processing circuit through a cache coherence interface, and connected to the first network interface circuit. In some embodiments: The first memory module further includes a second memory die. The first memory die includes volatile memory, and the second memory die includes persistent memory. In some embodiments, the persistent memory includes NAND flash memory. In some embodiments, the controller is configured to provide a flash translation layer to the persistent memory. In some embodiments, the cache coherence interface includes a Compute Express Link (CXL) interface. In some embodiments, the first server includes an expansion slot adapter connected to an expansion slot of the first server. The expansion slot adapter includes the first memory module and the first network interface circuit. In some embodiments, the controller of the first memory module is connected to the storage program processing circuit through the expansion slot. In some embodiments, the expansion slot includes an M.2 slot. In some embodiments, the controller of the first memory module is connected to the first network interface circuit through a Peer-to-Peer Peripheral Component Interconnect Express (PCIe) connection. In some embodiments, the system further includes a second server and a network switch connected to the first server and connected to the second server. In some embodiments, the network switch includes a Top-of-Rack (ToR) Ethernet switch. In some embodiments, the controller of the first memory module is configured to receive a direct Remote Direct Memory Access (RDMA) request and send a direct RDMA response. In some embodiments, the controller of the first memory module is configured to receive a direct Remote Direct Memory Access (RDMA) request through the network switch and through the first network interface circuit, and send a direct RDMA response through the network switch and through the first network interface circuit. In some embodiments, the controller of the first memory module is configured to: receive data from the second server; store the data in the first memory module; and send a command to invalidate a cache line to the storage program processing circuit. In some embodiments, the controller of the first memory module includes a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC). According to an embodiment of the present invention, a method for performing Remote Direct Memory Access in a computing system is provided. The computing system includes a first server and a second server. The first server includes a storage program processing circuit, a network interface circuit, and a first memory module including a controller. The method includes receiving a direct Remote Direct Memory Access (RDMA) request by the controller of the first memory module and sending a direct RDMA response by the controller of the first memory module.In some embodiments: The computing system further includes an Ethernet switch connected to the first server and connected to the second server, and receiving a direct RDMA request includes receiving a direct RDMA request through the Ethernet switch. In some embodiments, the method further includes a controller of the first memory module receiving a read command for a first memory address from a storage program processing circuit, the controller of the first memory module converting the first memory address to a second memory address, and the controller of the first memory module retrieving data at the second memory address from the first memory module. In some embodiments, the method further includes a controller of the first memory module receiving data, the controller of the first memory module storing the data in the first memory module, and the controller of the first memory module sending a command to the storage program processing circuit to invalidate a cache line. According to an embodiment of the present invention, a system is provided that includes a first server, the first server including a storage program processing circuit, a first network interface circuit, and a first memory module, wherein the first memory module includes a first memory die and a controller mechanism, the controller mechanism being connected to the first memory die through a memory interface, connected to the storage program processing circuit through a cache coherence interface, and connected to the first network interface circuit.
[0056] Referring to Figure 1C , in some embodiments, a server system includes a plurality of servers 105 connected together through a top-of-rack (ToR) Ethernet switch 110. Each server includes one or more processing circuits 115, each processing circuit being connected to (i) a system memory 120 (e.g., DDR4 memory), (ii) one or more network interface circuits 125, and (iii) an enhanced-capability CXL switch 130. The enhanced-capability CXL switch 130 can be connected to a plurality of memory modules 135. That is, Figure 1C the system includes a first server 105, the first server 105 including a storage program processing circuit 115, a network interface circuit 125, a cache coherence switch 130, and a first memory module 135. In Figure 1C the system, the first memory module 135 is connected to the cache coherence switch 130, the cache coherence switch 130 is connected to the network interface circuit 125, and the storage program processing circuit 115 is connected to the cache coherence switch 130.
[0057] Memory modules 135 may be grouped by type, specification, or technology type (e.g., DDR4, DRAM, LDPPR, high bandwidth memory (HBM), or NAND flash or other persistent storage (e.g., solid state drives incorporating NAND flash)). Each memory module may have a CXL interface and include interface circuitry for converting between CXL packets and signals appropriate for the memory in memory module 135. In some embodiments, this interface circuitry is alternatively in the enhanced capability CXL switch 130, and each memory module 135 has an interface that is an inherent interface of the memory in memory module 135. In some embodiments, the enhanced capability CXL switch 130 is integrated into the memory module 135 (e.g., integrated with other components of the memory module 135 in an M.2 form factor package, or integrated with other components of the memory module 135 into a single integrated circuit).
[0058] The ToR Ethernet switch 110 may include interface hardware to facilitate RDMA requests between aggregated memory devices on different servers. The enhanced capability CXL switch 130 may include one or more circuits (e.g., it may include an FPGA or ASIC) to (i) route data to different memory types based on workload, (ii) virtualize host addresses to device addresses, and / or (iii) bypass the processing circuitry 115 to facilitate RDMA requests between different servers.
[0059] The memory module 135 can be in an expansion box (e.g., in the same rack as the motherboard for accommodating accessories), and the expansion box can include a predetermined number (e.g., more than 20 or more than 100) of memory modules 135, with each memory module inserted into a suitable connector. The modules can be in the M.2 specification, and the connector can be an M.2 connector. In some embodiments, the connection between servers is via a different network other than Ethernet. For example, they can be wireless connections such as WiFi or 5G connections. Each processing circuit can be an x86 processor or another processor, such as an ARM processor or a GPU. The PCIe link on which the CXL link is instantiated can be PCIe 5.0 or another version (e.g., an earlier version or a later (e.g., future) version (e.g., PCIe 6.0)). In some embodiments, instead of CXL or in addition to CXL, a different cache coherence protocol is used in the system, and instead of the enhanced-capability CXL switch 130 or in addition to the enhanced-capability CXL switch 130, a different cache coherence switch can be used. Such a cache coherence protocol can be another standard protocol or a cache coherence variant of a standard protocol (in a manner similar to the way CXL is a variant of PCIe 5.0). Examples of standard protocols include, but are not limited to, Non-Volatile Dual In-line Memory Module (version P) (NVDIMM-P), Cache Coherent Interconnect for Accelerators (CCIX), and Open Coherent Accelerator Processor Interface (OpenCAPI).
[0060] The system memory 120 can include, for example, DDR4 memory, DRAM, HBM, or LDPPR memory. The memory module 135 can be partitioned or include a cache controller to handle multiple memory types. The memory module 135 can be in different specifications, and examples of the specifications include, but are not limited to, HHHL, FHHL, M.2, U.2, mezzanine card, daughter card, E1.S, E1.L, E3.L, and E3.S.
[0061] In some embodiments, the system implementation includes an aggregated architecture of multiple servers, where each server aggregates multiple memory modules 135 attached with CXL. Each memory module 135 may include multiple partitions, which may be separately exposed as memory devices to multiple processing circuits 115. Each input port of the enhanced-capability CXL switch 130 can independently access multiple output ports of the enhanced-capability CXL switch 130 and the memory modules 135 connected thereto. As used herein, an "input port" or "upstream port" of the enhanced-capability CXL switch 130 is a port connected to a PCIe root port (or suitable for connection to a PCIe root port), and an "output port" or "downstream port" of the enhanced-capability CXL switch 130 is a port connected to a PCIe endpoint (or suitable for connection to a PCIe endpoint). As in the case of the embodiments of Figure 1A , each memory module 135 may expose the collective of base address registers (BARs) to the host BIOS as a memory range. One or more of the memory modules 135 may include firmware to transparently manage its memory space behind the host OS mapping.
[0062] In some embodiments, as described above, the enhanced-capability CXL switch 130 includes an FPGA (or ASIC) controller 137 and provides additional functionality beyond the switching of CXL packets. For example, it can (as described above) virtualize the memory module 135, i.e., operate as a translation layer that converts between a processing circuit-side address (or "processor-side" address, i.e., the address included in the processor read and write commands issued by the processing circuit 115) and a memory-side address (i.e., the address used by the enhanced-capability CXL switch 130 to address storage locations in the memory module 135), thereby masking the physical address of the memory module 135 and presenting a virtual aggregation of the memory. The controller 137 of the enhanced-capability CXL switch 130 can also act as a management device for the memory module 135 and facilitate host control plane processing. The controller 137 can transparently move data and update the memory map (or "address translation table") accordingly without the participation of the processing circuit 115, such that subsequent accesses function as expected. The controller 137 can include a switch management device that (i) can appropriately bind and unbind upstream and downstream connections during runtime, and (iii) can enable rich control semantics and statistics associated with data transfer to and from the memory module 135. The controller 137 can include an additional "backdoor" 100 GbE or other network interface circuit 125 (in addition to the network interface for connecting to the host) for connecting to other servers 105 or connecting to other networking facilities. In some embodiments, the controller 137 presents itself to the processing circuit 115 as a type 2 device, which enables it to issue cache invalidation instructions to the processing circuit 115 after receiving a remote write request. In some embodiments, the DDIO technology is enabled, and remote data is first pulled into the last-level cache (LLC) of the processing circuit 115 and later written (from the cache) to the memory module 135.
[0063] As described above, one or more of the memory modules 135 can include persistent storage. If the memory module 135 is presented as a persistent device, the controller 137 of the enhanced-capability CXL switch 130 can manage the persistent domain (e.g., it can store data identified by the processing circuit 115 (e.g., by using the corresponding operating system functions) as requiring persistent storage in the persistent storage. In such embodiments, software APIs can flush the cache and data to the persistent storage.
[0064] In some embodiments, it can be in a manner similar to that described above for Figure 1A and Figure 1BExecute direct memory transfers to the memory module 135 in a manner similar to that described in the embodiments, and the operations performed by the controller of the memory module 135 are performed by the controller 137 of the enhanced-capability CXL switch 130.
[0065] As described above, in some embodiments, the memory modules 135 are organized into groups, e.g., one group that is memory-intensive, another group that is HBM-dense, another group that has limited density and performance, and another group that has dense capacity. Such groups may have different specifications or be based on different technologies. The controller 137 of the enhanced-capability CXL switch 130 can intelligently route data and commands based on, e.g., workload, tags, or quality of service (QoS). For read requests, there may be no routing based on such factors.
[0066] The controller 137 of the enhanced CXL switch 130 can also virtualize the processing circuit side address and the memory side address (as described above), so that the controller 137 of the enhanced CXL switch 130 can determine where the data is to be stored. The controller 137 of the enhanced CXL switch 130 can make such a determination based on the information or instructions it can receive from the processing circuit 115. For example, the operating system can provide a memory allocation feature such that an application can specify that low-latency storage or high-bandwidth storage or persistent storage is to be allocated, and then the controller 137 of the enhanced CXL switch 130 can take such requests initiated by the application into account when determining where to allocate the memory (e.g., in which memory module 135). For example, storage for which the application requests high bandwidth can be allocated in a memory module 135 that includes HBM, storage for which the application requests data persistence can be allocated in a memory module 135 that includes NAND flash, and other storage (for which the application has not made a request) can be stored on a memory module 135 that includes relatively inexpensive DRAM. In some embodiments, the controller 137 of the enhanced CXL switch 130 can make a determination about where to store certain data based on network usage patterns. For example, the controller 137 of the enhanced CXL switch 130 can determine that data within a certain range of physical addresses is accessed more frequently than other data by monitoring the usage pattern, and then the controller 137 of the enhanced CXL switch 130 can copy this data to a memory module 135 that includes HBM and modify its address translation table so that the data at the new location is stored at the same range of virtual addresses. In some embodiments, one or more of the memory modules 135 include flash memory (e.g., NAND flash), and the controller 137 of the enhanced CXL switch 130 implements a flash translation layer for this flash memory. The flash translation layer can support overwriting the storage location on the processor side (by moving the data to a different location and marking the previous location of the data as invalid), and it can perform garbage collection (e.g., when the portion of the data marked as invalid in a block exceeds a threshold, erasing the block after moving any valid data in the block to another block).
[0067] In some embodiments, the controller 137 of the enhanced-capability CXL switch 130 may facilitate the transfer of physical function (PF) to PF. For example, if one of the processing circuits 115 needs to move data from one physical address to another physical address (which may have the same virtual address; this fact does not necessarily affect the operation of the processing circuit 115), or if the processing circuit 115 needs to move data between two virtual addresses (the processing circuit 115 would need to have these two virtual addresses), then the controller 137 of the enhanced-capability CXL switch 130 may oversee the transfer without the involvement of the processing circuit 115. For example, the processing circuit 115 may send a CXL request, and the data may be sent from one memory module 135 to another memory module 135 behind the enhanced-capability CXL switch 130 without progressing to the processing circuit 115 (e.g., the data may be copied from one memory module 135 to another memory module 135). In this case, since the processing circuit 115 initiates the CXL request, the processing circuit 115 may need to flush its cache to ensure consistency. If a type 2 memory device (e.g., one of the memory modules 135 or an accelerator that may also be connected to the CXL switch) instead initiates the CXL request and the switch is not virtualized, the type 2 memory device may send a message to the processing circuit 115 to invalidate the cache.
[0068] In some embodiments, the controller 137 of the enhanced-capability CXL switch 130 may facilitate RDMA requests between servers. The remote server 105 may initiate such an RDMA request, and the request may be sent through the ToR Ethernet switch 110 and arrive at the enhanced-capability CXL switch 130 in the server 105 (“local server”) that responds to the RDMA request. The enhanced-capability CXL switch 130 may be configured to receive such an RDMA request, and it may treat a set of memory modules 135 in the receiving server 105 (i.e., the server that receives the RDMA request) as its own memory space. In the local server, the enhanced-capability CXL switch 130 may receive the RDMA request as a direct RDMA request (i.e., an RDMA request that is not routed through the processing circuit 115 in the local server), and it may send a direct response to the RDMA request (i.e., it may send the response without routing the response through the processing circuit 115 in the local server). In the remote server, the response (e.g., data sent by the local server) may be received by the enhanced-capability CXL switch 130 of the remote server and stored in the memory module 135 of the remote server without being routed through the processing circuit 115 in the remote server.
[0069] Figure 1D is shown in connection with Figure 1CA system similar to the system, where the processing circuit 115 is connected to the network interface circuit 125 through an enhanced-capability CXL switch 130. The enhanced-capability CXL switch 130, the memory module 135, and the network interface circuit 125 are on the expansion slot adapter 140. The expansion slot adapter 140 can be a circuit board or module inserted into an expansion slot (e.g., a PCIe connector 145) on the motherboard of the server 105. Thus, the server can be any suitable server modified only by installing the expansion slot adapter 140 in the PCIe connector 145. The memory module 135 can be installed in a connector (e.g., an M.2 connector) on the expansion slot adapter 140. In such an embodiment, (i) the network interface circuit 125 can be integrated into the enhanced-capability CXL switch 130, or (ii) each network interface circuit 125 can have a PCIe interface (the network interface circuit 125 can be a PCIe endpoint), such that the processing circuit 115 to which it is connected can communicate with the network interface circuit 125 through a root port to endpoint PCIe connection. The controller 137 of the enhanced-capability CXL switch 130 (which can have a PCIe input port connected to the processing circuit 115 and connected to the network interface circuit 125) can communicate with the network interface circuit 125 through a peer-to-peer PCIe connection.
[0070] According to an embodiment of the present invention, a system is provided that includes a first server, the first server including a storage program processing circuit, a network interface circuit, a cache coherent switch, and a first memory module, wherein: the first memory module is connected to the cache coherent switch, the cache coherent switch is connected to the network interface circuit, and the storage program processing circuit is connected to the cache coherent switch. In some embodiments, the system further includes a second memory module connected to the cache coherent switch, wherein the first memory module includes volatile memory and the second memory module includes persistent memory. In some embodiments, the cache coherent switch is configured to virtualize the first memory module and the second memory module. In some embodiments, the first memory module includes flash memory and the cache coherent switch is configured to provide a flash translation layer to the flash memory. In some embodiments, the cache coherent switch is configured to: monitor the access frequency of a first memory location in the first memory module; determine that the access frequency exceeds a first threshold; and copy the content of the first memory location to a second memory location in the second memory module. In some embodiments, the second memory module includes high bandwidth memory (HBM). In some embodiments, the cache coherent switch is configured to maintain a table for mapping processor-side addresses to memory-side addresses. In some embodiments, the system further includes a second server and a network switch connected to the first server and the second server. In some embodiments, the network switch includes a top-of-rack (ToR) Ethernet switch. In some embodiments, the cache coherent switch is configured to receive direct remote direct memory access (RDMA) requests and send direct RDMA responses. In some embodiments, the cache coherent switch is configured to receive remote direct memory access (RDMA) requests through the ToR Ethernet switch and through the network interface circuit, and send direct RDMA responses through the ToR Ethernet switch and through the network interface circuit. In some embodiments, the cache coherent switch is configured to support the Compute Express Link (CXL) protocol. In some embodiments, the first server includes an expansion slot adapter connected to an expansion slot of the first server, the expansion slot adapter including a cache coherent switch and a memory module slot, and the first memory module is connected to the cache coherent switch through the memory module slot. In some embodiments, the memory module slot includes an M.2 slot. In some embodiments, the network interface circuit is on the expansion slot adapter.According to an embodiment of the present invention, there is provided a method for performing remote direct memory access in a computing system, the computing system including a first server and a second server, the first server including a storage program processing circuit, a network interface circuit, a cache coherent switch, and a first memory module, the method including receiving, by the cache coherent switch, a direct remote direct memory access (RDMA) request, and sending, by the cache coherent switch, a direct RDMA response. In some embodiments: the computing system further includes an Ethernet switch, and receiving the direct RDMA request includes receiving the direct RDMA request through the Ethernet switch. In some embodiments, the method further includes receiving, by the cache coherent switch, a read command for a first memory address from the storage program processing circuit, converting, by the cache coherent switch, the first memory address to a second memory address, and retrieving, by the cache coherent switch, data at the first memory address from the first memory module. In some embodiments, the method further includes receiving, by the cache coherent switch, data, storing, by the cache coherent switch, the data in the first memory module, and sending, by the cache coherent switch, a command to invalidate a cache line to the storage program processing circuit. According to an embodiment of the present invention, there is provided a system, which includes a first server, the first server including a storage program processing circuit, a network interface circuit, a cache coherent switching mechanism, and a first memory module, wherein: the first memory module is connected to the cache coherent switching mechanism, the cache coherent switching mechanism is connected to the network interface circuit, and the storage program processing circuit is connected to the cache coherent switching mechanism.
[0071] Figure 1E An embodiment is shown in which each of a plurality of servers 105 is connected to the shown ToR server link switch 112, which can be a PCIe 5.0 CXL switch with PCIe capabilities. The server link switch 112 can include an FPGA or an ASIC and can provide performance (in terms of throughput and latency) superior to that of an Ethernet switch. Each server 105 can include a plurality of memory modules 135 connected to the server link switch 112 through an enhanced capability CXL switch 130 and through a plurality of PCIe connectors. As shown, each server 105 can further include one or more processing circuits 115 and system memory 120. The server link switch 112 can operate as a master, and each enhanced capability CXL switch 130 can operate as a slave, as discussed in further detail below.
[0072] In Figure 1EIn an embodiment, the server link switch 112 can group or batch multiple cache requests received from different servers 105, and it can group the groups, thereby reducing control overhead. The enhanced-capability CXL switch 130 can include a controller (e.g., from an FPGA or from an ASIC) to (i) route data to different memory types based on the workload, (ii) virtualize the processor-side address to a memory-side address, and (iii) bypass the processing circuit 115 to facilitate coherent requests between different servers 105. Figure 1E The system shown can be based on CXL 2.0, it can include distributed shared memory within the rack, and it can use the ToR server link switch 112 for local connection to remote nodes.
[0073] The ToR server link switch 112 can have additional network connections (e.g., an Ethernet connection as shown, or another connection, such as a wireless connection, such as a WiFi connection or a 5G connection) for connecting to other servers or to clients. Both the server link switch 112 and the enhanced-capability CXL switch 130 can include a controller, which can be or include a processing circuit such as an ARM processor. The PCIe interface can conform to the PCIe 5.0 standard, or an earlier version of the PCIe standard or a future version of the PCIe standard, or it can employ an interface conforming to a different standard (e.g., NVDIMM-P, CCIX, or OpenCAPI) instead of the PCIe interface. The memory module 135 can include various memory types, including DDR4 DRAM, HBM, LDPPR, NAND flash, or solid-state drives (SSDs). The memory module 135 can be partitioned or include cache controllers to handle multiple memory types, and they can come in different form factors, such as HHHL, FHHL, M.2, U.2, mezzanine cards, daughter cards, E1.S, E1.L, E3.L, or E3.S.
[0074] In Figure 1EIn an embodiment, the enhanced-capability CXL switch 130 can enable one-to-many and many-to-one switching, and it can enable a fine-grained load-store interface at the flit (64-byte) level. Each server can have an aggregated memory device, and each device is divided into multiple logical devices, each with a corresponding LD-ID. The ToR switch 112 (which can be referred to as a "server link switch") enables the one-to-many function, and the enhanced-capability CXL switch 130 in the server 105 enables the many-to-one function. The server link switch 112 can be a PCIe switch or a CXL switch or both. In such a system, the requester can be the processing circuits 115 of multiple servers 105, and the responder can be many aggregated memory modules 135. The hierarchical arrangement of the two switches (as described above, the primary switch is the server link switch 112 and the secondary switch is the enhanced-capability CXL switch 130) enables any communication. Each memory module 135 can have one physical function (PF) and up to 16 isolated logical devices. In some embodiments, the number of logical devices (e.g., the number of partitions) can be limited (e.g., limited to 16), and there can also be a control partition (which can be the physical function for control). Each memory module 135 can be a type 2 device, which has cxl.cache, cxl.mem, and cxl.io as well as an address translation service (ATS) implementation to handle cache line copies that the processing circuits 115 can hold. The enhanced-capability CXL switch 130 and the fabric manager can control the discovery of the memory modules 135, and (i) perform device discovery and virtual CXL software creation, and (ii) bind virtual ports to physical ports. As in the Figures 1A - 1D embodiment, the fabric manager can operate through a connection on the SMBus sideband. The interface to the memory module 135 can implement configurability, and the interface can be an Intelligent Platform Management Interface (IPMI) or an interface compliant with the Redfish standard (and can also provide additional functions not required by the standard).
[0075] As described above, some embodiments implement a hierarchical structure in which the master controller (which can be implemented as an FPGA or ASIC) is part of the server link switch 112 and the slave controller is part of the enhanced-capability CXL switch 130 to provide a load-store interface (i.e., an interface with cache line (e.g., 64-byte) granularity and operating within the coherence domain without involving software drivers). Such a load-store interface can extend the coherence domain beyond a single server, CPU, or host and can involve electrical or optical physical media (e.g., an optical connection with electrical-to-optical transceivers at both ends). In operation, the master controller (in the server link switch 112) boots (or "reboots") and configures all the servers 105 on the rack. The master controller can have visibility across all hosts, and it can (i) discover each server and find out how many servers 105 and memory modules 135 exist in the server cluster, (ii) configure each server 105 independently, (iii) enable or disable some memory blocks on different servers (e.g., enable or disable any memory module 135) based on, for example, the rack configuration, (iv) control access (e.g., which server can control which other server), (v) implement flow control (e.g., since all host and device requests pass through the master, it can send data from one server to another and perform flow control on that data), (vi) group or batch requests or packets (e.g., multiple cache requests received by the master from different servers 105), and (vii) receive remote software updates, broadcast communications, etc. In batch mode, the server link switch 112 can receive multiple packets destined for the same server (e.g., destined for the first server) and send them together (i.e., without pausing between them) to the first server. For example, the server link switch 112 can receive a first packet from a second server and a second packet from a third server and send the first packet and the second packet together to the first server. Each server 105 can expose to the master controller (i) an IPMI network interface, (ii) a system event log (SEL), and (iii) a board management controller (BMC), enabling the master controller to measure performance, measure reliability in operation, and reconfigure the server 105.
[0076] In some embodiments, a software architecture that uses a load-store interface that promotes high availability is used. Such a software architecture can provide reliability, replication, consistency, system coherence, hashing, caching, and persistence. By performing periodic hardware checks on the CXL device components via IPMI, the software architecture can provide reliability (in a system with a large number of servers). For example, the server link switch 112 can query the status of the memory server 150 through the IPMI interface of the memory server 150, such as querying the power status (whether the power of the memory server 150 is working properly), the network status (whether the interface of the server link switch 112 is working properly), and the error check status (whether there is an error condition in any subsystem of the memory server 150). The software architecture can provide replication because the master controller can replicate the data stored in the memory module 135 and maintain data consistency among all copies.
[0077] The software architecture can provide consistency because the master controller can be configured with different consistency levels, and the server link switch 112 can adjust the packet format according to the consistency level to be maintained. For example, if eventual consistency is maintained, the server link switch 112 can reorder requests, while to maintain strict consistency, the server link switch 112 can maintain a scoreboard of all requests with precise timestamps at the switch. The software architecture can provide system coherence because multiple processing circuits 115 can read from or write to the same memory address, and to maintain coherence, the master controller can be responsible for reaching the master node at that address (using directory lookup) or broadcasting requests on a common bus.
[0078] The software architecture can provide hashing because the server link switch 112 and the enhanced-capability CXL switch can maintain a virtual mapping of addresses, which can use consistent hashing with multiple hash functions to evenly map data to all CXL devices on all nodes at startup (or to regulate when a server goes idle or starts). The software architecture can provide caching because the master controller can specify certain memory partitions (e.g., in the memory module 135 including HBM or a technology with similar capabilities) to act as caches (which employ, for example, write-through caches or write-back caches). The software architecture can provide persistence because the master controller and the slave controller can manage persistent domains and flushes.
[0079] In some embodiments, the capabilities of the CXL switch are integrated into the controller of the memory module 135. In such embodiments, the server link switch 112 can still act as the master and have enhanced features as discussed elsewhere herein. The server link switch 112 can also manage other storage devices in the system, and it can have an Ethernet connection (e.g., a 100 GbE connection) for connecting to, for example, a client computer that is not part of the PCIe network formed by the server link switch 112.
[0080] In some embodiments, the server link switch 112 has enhanced capabilities and also includes an integrated CXL controller. In additional embodiments, the server link switch 112 is only a physical routing device, and each server 105 includes a primary CXL controller. In such embodiments, the masters across different servers can negotiate a master-slave architecture. The (i) enhanced-capability CXL switch 130 and (ii) the intelligent functions of the server link switch 112 can be implemented in one or more FPGAs, one or more ASICs, one or more ARM processors, or in one or more SSD devices with computing capabilities. The server link switch 112 can perform flow control, for example, by reordering independent requests. In some embodiments, since the interface is load-store, RDMA is optional, but there may be intermediate RDMA requests using the PCIe physical medium (instead of 100 GbE). In such embodiments, a remote master can initiate an RDMA request that can be sent through the server link switch 112 to the enhanced-capability CXL switch 130. The server link switch 112 and the enhanced-capability CXL switch 130 can prioritize RDMA 4KR requests or CXL microslice (64-byte) requests.
[0081] As in Figure 1C and the embodiment of FIG. 2, the enhanced-capability CXL switch 130 can be configured to receive such RDMA requests, and it can treat a set of memory modules 135 in the receiving server 105 (i.e., the server that receives the RDMA request) as its own memory space. Additionally, the enhanced-capability CXL switch 130 can be virtualized across the processing circuitry 115 and initiate an RDMA request on a remote enhanced-capability CXL switch 130 to move data back and forth between the servers 105 without involving the processing circuitry 115.
[0082] Figure 1F A system similar to Figure 1E is shown, where the processing circuitry 115 is connected to the network interface circuitry 125 through the enhanced-capability CXL switch 130. As in Figure 1Das in the embodiments of Figure 1F In Figure 1F , the enhanced capabilities CXL switch 130, memory module 135, and network interface circuit 125 are on the expansion slot adapter 140. The expansion slot adapter 140 can be a circuit board or module inserted into an expansion slot (e.g., a PCIe connector 145) on the motherboard of the server 105. Thus, the server can be any suitable server modified by simply installing the expansion slot adapter 140 in the PCIe connector 145. The memory module 135 can be installed in a connector (e.g., an M.2 connector) on the expansion slot adapter 140. In such an embodiment, (i) the network interface circuit 125 can be integrated into the enhanced capabilities CXL switch 130, or (ii) each network interface circuit 125 can have a PCIe interface (the network interface circuit 125 can be a PCIe endpoint) such that the processing circuit 115 to which it is connected can communicate with the network interface circuit 125 via a root port to endpoint PCIe connection, and the controller 137 of the enhanced capabilities CXL switch 130 (which can have a PCIe input port connected to the processing circuit 115 and connected to the network interface circuit 125) can communicate with the network interface circuit 125 via a peer-to-peer PCIe connection.
[0083] According to an embodiment of the present invention, a system is provided, which includes: a first server, including a storage program processing circuit, a cache coherent switch, and a first memory module; and a second server; and a server link switch connected to the first server and connected to the second server, wherein: the first memory module is connected to the cache coherent switch, the cache coherent switch is connected to the server link switch, and the storage program processing circuit is connected to the cache coherent switch. In some embodiments, the server link switch includes a Peripheral Component Interconnect Express (PCIe) switch. In some embodiments, the server link switch includes a Compute Express Link (CXL) switch. In some embodiments, the server link switch includes a Top of Rack (ToR) CXL switch. In some embodiments, the server link switch is configured to discover the first server. In some embodiments, the server link switch is configured to restart the first server. In some embodiments, the server link switch is configured to disable the first memory module by the cache coherent switch. In some embodiments, the server link switch is configured to send data from the second server to the first server and perform flow control on the data. In some embodiments, the system further includes a third server connected to the server link switch, wherein the server link switch is configured to: receive a first packet from the second server, receive a second packet from the third server, and send the first packet and the second packet to the first server. In some embodiments, the system further includes a second memory module connected to the cache coherent switch, wherein the first memory module includes volatile memory and the second memory module includes persistent memory. In some embodiments, the cache coherent switch is configured to virtualize the first memory module and the second memory module. In some embodiments, the first memory module includes flash memory, and the cache coherent switch is configured to provide a flash translation layer to the flash memory. In some embodiments, the first server includes an expansion slot adapter connected to the expansion slot of the first server, the expansion slot adapter includes a cache coherent switch and a memory module slot, and the first memory module is connected to the cache coherent switch through the memory module slot. In some embodiments, the memory module slot includes an M.2 slot. In some embodiments: the cache coherent switch is connected to the server link switch through a connector on the expansion slot adapter.According to an embodiment of the present invention, there is provided a method for performing remote direct memory access in a computing system, the computing system including: a first server, a second server, a third server, and a server link switch connected to the first server, connected to the second server, and connected to the third server, the first server including a stored program processing circuit, a cache coherent switch, and a first memory module, the method including receiving, by the server link switch, a first packet from the second server, receiving, by the server link switch, a second packet from the third server, and sending the first packet and the second packet to the first server. In some embodiments, the method further includes receiving, by the cache coherent switch, a direct remote direct memory access (RDMA) request, and sending, by the cache coherent switch, a direct RDMA response. In some embodiments, receiving the direct RDMA request includes receiving the direct RDMA request through the server link switch. In some embodiments, the method further includes receiving, by the cache coherent switch, a read command for a first memory address from the stored program processing circuit, converting, by the cache coherent switch, the first memory address to a second memory address, and retrieving, by the cache coherent switch, data at the second memory address from the first memory module. According to an embodiment of the present invention, there is provided a system, including: a first server, including a stored program processing circuit, a cache coherent switching mechanism, a first memory module; a second server; and a server link switch connected to the first server and connected to the second server, wherein: the first memory module is connected to the cache coherent switching mechanism, the cache coherent switching mechanism is connected to the server link switch, and the stored program processing circuit is connected to the cache coherent switching mechanism.
[0084] Figure 1G An embodiment is shown in which each of a plurality of memory servers 150 is connected to a ToR server link switch 112 as shown, and the ToR server link switch 112 may be a PCIe 5.0 CXL switch. As in Figure 1E and Figure 1F 's embodiment, the server link switch 112 may include an FPGA or an ASIC, and may provide performance superior to that of an Ethernet switch (in terms of throughput and latency). As in Figure 1E and Figure 1F 's embodiment, the memory server 150 may include a plurality of memory modules 135 connected to the server link switch 112 through a plurality of PCIe connectors. In Figure 1G 's embodiment, there may be no processing circuit 115 and system memory 120, and the main purpose of the memory server 150 may be to provide memory for use by other servers 105 having computing resources.
[0085] In Figure 1G the embodiment, the server link switch 112 can group or batch multiple cache requests received from different memory servers 150, and it can group the groups, thereby reducing the control overhead. The enhanced-capability CXL switch 130 can include composable hardware building blocks to (i) route data as different memory types based on the workload, and (ii) virtualize processor-side addresses (convert these addresses to memory-side addresses). Figure 1G The system shown can be based on CXL 2.0, which can include composable and disaggregated shared memory within the rack, and it can use the ToR server link switch 112 to provide pooled (i.e., aggregated) memory to remote devices.
[0086] The ToR server link switch 112 can have additional network connections for connecting to other servers or clients (e.g., an Ethernet connection as shown, or another connection, such as a wireless connection, such as a WiFi connection or a 5G connection). Both the server link switch 112 and the enhanced-capability CXL switch 130 can include a controller, which can be or include processing circuitry such as an ARM processor. The PCIe interface can conform to the PCIe 5.0 standard, or an earlier version of the PCIe standard or a future version of the PCIe standard, or can adopt a different standard (e.g., NVDIMM-P, CCIX, or OpenCAPI) instead of PCIe. The memory module 135 can include various memory types, including DDR4 DRAM, HBM, LDPPR, NAND flash, and solid-state drives (SSDs). The memory module 135 can be partitioned or include cache controllers to handle multiple memory types, and they can come in different form factors, such as HHHL, FHHL, M.2, U.2, mezzanine cards, daughter cards, E1.S, E1.L, E3.L, or E3.S.
[0087] In Figure 1GIn embodiments, the enhanced-capability CXL switch 130 can enable one-to-many and many-to-one switching, and it can implement a fine-grained load-store interface at the microtile (64-byte) level. Each memory server 150 can have an aggregated memory device, with each device divided into multiple logical devices, each logical device having a corresponding LD-ID. The enhanced-capability CXL switch 130 can include a controller 137 (e.g., an ASIC or FPGA) and circuitry for device discovery, enumeration, partitioning, and presentation of physical address ranges (which can be separate from or part of such an ASIC or FPGA). Each memory module 135 can have one physical function (PF) and up to 16 isolated logical devices. In some embodiments, the number of logical devices (e.g., the number of partitions) can be limited (e.g., to 16), and there can also be a control partition (which can be for controlling the physical function of the device). Each memory module 135 can be a type 2 device, having cxl.cache, cxl.mem, and cxl.io, as well as an address translation service (ATS) implementation to handle cache line copies that the processing circuitry 115 can hold.
[0088] The enhanced-capability CXL switch 130 and the fabric manager can control the discovery of the memory modules 135 and (i) perform device discovery and virtual CXL software creation, and (ii) bind virtual ports to physical ports. As in Figures 1A - 1D the embodiments of, the fabric manager can operate over a connection on the SMBus sideband. The interface to the memory module 135 can implement configurability, which can be an intelligent platform management interface (IPMI) or an interface compliant with the Redfish standard (and can also provide additional features not required by the standard).
[0089] For Figure 1G the embodiments, the building blocks can include (as described above) a CXL controller 137 implemented on an FPGA or ASIC for switching to enable aggregation of memory devices (e.g., memory modules 135), SSDs, accelerators (GPUs, NICs), CXL and PCIe5 connectors, and firmware for an advanced configuration and power interface (ACPI) table (such as a heterogeneous memory attributes table (HMAT) or a static resource affinity table SRAT) for exposing device details to the operating system.
[0090] In some embodiments, the system provides composability. The system can provide the ability to bring CXL devices and other accelerators online and offline based on software configuration, and it can be capable of grouping and allocating accelerator, memory, and storage device resources to each memory server 150 in a rack. The system can hide the physical address space and use faster devices such as HBM and SRAM to provide transparent caching.
[0091] In Figure 1G an embodiment, the controller 137 of the enhanced-capability CXL switch 130 can (i) manage the memory module 135, (ii) integrate and control heterogeneous devices such as NICs, SSDs, GPUs, DRAMs, and (iii) achieve dynamic reconfiguration of storage for memory devices through power gating. For example, the ToR server link switch 112 can disable the power supply to one of the memory modules 135 (i.e., cut off the power supply or reduce the power supply) (by instructing the enhanced-capability CXL switch 130 to disable the power supply to the memory module 135). Then, the enhanced-capability CXL switch 130 can disable the power supply to the memory module 135 after being instructed by the server link switch 112 to disable the power supply to the memory module. Such a disablement can save power and can improve the performance (e.g., throughput and latency) of other memory modules 135 in the memory server 150. Each remote server 105 can see a different logical view of the memory module 135 and its connections based on negotiation. The controller 137 of the enhanced-capability CXL switch 130 can maintain the state such that each remote server maintains the allocated resources and connections, and it can perform compression or deduplication of memory to save memory capacity (using a configurable block size). Figure 1G A disaggregated rack can have its own BMC. It can also expose the IPMI network interface and system event log (SEL) to remote devices, enabling a master machine (e.g., a remote server using the storage provided by the memory server 150) to measure performance and reliability in operation and reconfigure the disaggregated rack. Figure 1G A disaggregated rack can be in a manner similar to that described here for Figure 1Ein a manner similar to that described in the embodiments to provide reliability, replicability, consistency, system coherence, hashing, caching, and persistence. For example, coherence is provided by multiple remote servers that read from or write to the same memory address, and each remote server is configured to have a different consistency level. In some embodiments, the server link switch maintains eventual consistency between data stored on a first memory server and data stored on a second memory server. The server link switch 112 can maintain different consistency levels for different pairs of servers; for example, the server link switch can also maintain a consistency level between data stored on a first memory server and data stored on a third memory server, which is strict consistency, sequential consistency, causal consistency, or processor consistency. The system can employ communication in "local bands" (server link switch 112) and "global bands" (disaggregated servers) domains. Writes can be flushed to the "global band" to be visible for new reads from other servers. The controller 137 of the enhanced-capability CXL switch 130 can manage the persistent domain and flush separately for each remote server. For example, the cache coherence switch can monitor the fullness of a first region of memory (volatile memory, used as a cache), and when the fullness level exceeds a threshold, the cache coherence switch can move data from the first region of memory to a second region of memory, and the second region of memory is in persistent memory. Flow control can be handled because priorities can be established between remote servers by the controller 137 of the enhanced-capability CXL switch 130 to present different perceived latencies and bandwidths.
[0092] According to an embodiment of the present invention, a system is provided, which includes: a first memory server including a cache coherent switch and a first memory module; and a second memory server; and a server link switch connected to the first memory server and the second memory server, wherein: the first memory module is connected to the cache coherent switch, and the cache coherent switch is connected to the server link switch. In some embodiments, the server link switch is configured to disable power supply to the first memory module. In some embodiments: the server link switch is configured to disable power supply to the first memory module by instructing the cache coherent switch to disable power supply to the first memory module, and the cache coherent switch is configured to disable power supply to the first memory module after being instructed by the server link switch to disable power supply to the first memory module. In some embodiments, the cache coherent switch is configured to perform deduplication within the first memory module. In some embodiments, the cache coherent switch is configured to compress data and store the compressed data in the first memory module. In some embodiments, the server link switch is configured to query the status of the first memory server. In some embodiments, the server link switch is configured to query the status of the first memory server through the Intelligent Platform Management Interface (IPMI). In some embodiments, querying the status includes querying a status selected from the group consisting of a power state, a network state, and an error check state. In some embodiments, the server link switch is configured to batch cache requests directed to the first memory server. In some embodiments, the system further includes a third memory server connected to the server link switch, wherein the server link switch is configured to maintain a consistency level between the data stored on the first memory server and the data stored on the third memory server, and the consistency level is selected from the group consisting of strict consistency, sequential consistency, causal consistency, and processor consistency. In some embodiments, the cache coherent switch is configured to: monitor the fullness of a first region of the memory, and move data from the first region of the memory to a second region of the memory, wherein: the first region of the memory is in volatile memory, and the second region of the memory is in persistent memory. In some embodiments, the server link switch includes a Peripheral Component Interconnect Express (PCIe) switch. In some embodiments, the server link switch includes a Compute Express Link (CXL) switch. In some embodiments, the server link switch includes a Top-of-Rack (ToR) CXL switch. In some embodiments, the server link switch is configured to send data from the second memory server to the first memory server and perform flow control on the data.In some embodiments, the system further includes a third memory server connected to the server link switch, where: the server link switch is configured to receive a first packet from the second memory server, receive a second packet from the third memory server, and send the first packet and the second packet to the first memory server. According to an embodiment of the present invention, there is provided a method for performing remote direct memory access in a computing system, the computing system including: a first memory server; a first server; a second server; and a server link switch connected to the first memory server, connected to the first server and connected to the second server, the first memory server including a cache coherent switch and a first memory module, the first server including a stored program processing circuit, the second server including a stored program processing circuit, the method including receiving, by the server link switch, a first packet from the first server, receiving, by the server link switch, a second packet from the second server, and sending the first packet and the second packet to the first memory server. In some embodiments, the method further includes compressing data by the cache coherent switch and storing the data in the first memory module. In some embodiments, the method further includes querying, by the server link switch, the status of the first memory server. According to an embodiment of the present invention, there is provided a system, which includes: a first memory server, including a cache coherent switch and a first memory module; and a second memory server; and a server link switching mechanism connected to the first memory server and the second memory server, where: the first memory module is connected to the cache coherent switch, and the cache coherent switch is connected to the server link switching mechanism.
[0093] Figures 2A - 2D are flowcharts of various embodiments. In the embodiments of these flowcharts, the processing circuit 115 is a CPU; in other embodiments, it may be other processing circuits (e.g., GPU). Referring to Figure 2A , Figure 1A and Figure 1B the controller 137 of the memory module 135 of the embodiment of Figures 1C - 1GThe enhanced-capability CXL switch 130 of any embodiment can be virtualized across the processing circuitry 115 and initiate an RDMA request on the enhanced-capability CXL switch 130 in another server 105 to move data back and forth between the servers 105 without involving the processing circuitry 115 in either server (the virtualization is handled by the controller 137 of the enhanced-capability CXL switch 130). For example, at 205, the controller 137 of the memory module 135 or the enhanced-capability CXL switch 130 generates an RDMA request for additional remote memory (e.g., CXL memory or aggregated memory); at 210, the network interface circuit 125 bypasses the processing circuitry and sends the request to the ToR Ethernet switch 110 (which may have an RDMA interface); at 215, the ToR Ethernet switch 110 bypasses the remote processing circuitry 115 and routes the RDMA request to the remote server 105 via an RDMA access to the remote aggregated memory for processing by the controller 137 of the memory module 135 or by the remote enhanced-capability CXL switch 130; at 220, the ToR Ethernet switch 110 receives the processed data and routes the data to the local memory module 135 or to the local enhanced-capability CXL switch 130 via RDMA bypassing the local processing circuitry 115; at 222, Figure 1A and Figure 1B the controller 137 of the memory module 135 or the enhanced-capability CXL switch 130 of the embodiment directly receives an RDMA response (e.g., the response is not forwarded by the processing circuitry 115).
[0094] In such an embodiment, the controller 137 of the remote memory module 135 or the enhanced-capability CXL switch 130 of the remote server 105 is configured to receive a direct Remote Direct Memory Access (RDMA) request and send a direct RDMA response. As used herein, the controller 137 of the remote memory module 135 receives or the enhanced-capability CXL switch 130 receives a "direct RDMA request" (or "directly" receives such a request) means that such a request is received by the controller 137 of the remote memory module 135 or by the enhanced-capability CXL switch 130 without the response being forwarded or otherwise processed by the processing circuitry 115 of the remote server, and the controller 137 of the remote memory module 135 or the enhanced-capability CXL switch 130 sends a "direct RDMA response" (or "directly" sends such a request) means that such a response is sent without the response being forwarded or otherwise processed by the processing circuitry 115 of the remote server.
[0095] Referring to Figure 2B, in another embodiment, RDMA can be performed when the processing circuit of the remote server participates in data processing. For example, at 225, the processing circuit 115 can send data or a workload request via Ethernet; at 230, the ToR Ethernet switch 110 can receive the request and route it to the corresponding server 105 among the multiple servers 105; at 235, the request can be received within the server through the port(s) of the network interface circuit 125 (e.g., a 100GbE-enabled NIC); at 240, the processing circuit 115 (e.g., an x86 processing circuit) can receive the request from the network interface circuit 125; and at 245, the processing circuit 115 can process the request via the CXL 2.0 protocol using DDR and additional memory resources (e.g., together) to share the memory (which in Figure 1A and Figure 1B embodiments can be aggregated memory).
[0096] Referring to Figure 2C , in Figure 1E embodiments, RDMA can be performed when the processing circuit of the remote server participates in data processing. For example, at 225, the processing circuit 115 can send data or a workload request via Ethernet or PCie; at 230, the ToR Ethernet switch 110 can receive the request and route it to the corresponding server 105 among the multiple servers 105; at 235, the request can be received within the server through the port(s) of the PCIe connector; at 240, the processing circuit 115 (e.g., an x86 processing circuit) can receive the request from the network interface circuit 125; and at 245, the processing circuit 115 can process the request via the CXL 2.0 protocol using DDR and additional memory resources (e.g., together) to share the memory (which in Figure 1A and Figure 1BIn an embodiment, it may be an aggregated memory). At 250, the processing circuit 115 may identify the need to access memory contents (e.g., DDR or aggregated memory contents) from different servers; at 252, the processing circuit 115 may send a request for the memory contents (e.g., DDR or aggregated memory contents) from different servers via the CXL protocol (e.g., CXL 1.1 or CXL 2.0); at 254, the request propagates through the local PCIe connector to the server link switch 112, and then the server link switch 112 sends the request to the second PCIe connector of the second server on the rack; at 256, the second processing circuit 115 (e.g., x86 processing circuit) receives the request from the second PCIe connector; at 258, the second processing circuit 115 may process the request (e.g., retrieval of memory contents) using the second DDR and the second additional memory resources via the CXL 2.0 protocol to share the aggregated memory; and at 260, the second processing circuit (e.g., x86 processing circuit) sends the result of the request back to the original processing circuit via their respective PCIe connectors and through the server link switch 112.
[0097] Referring to Figure 2D , in Figure 1G In an embodiment, RDMA may be performed when the processing circuit of the remote server participates in data processing. For example, at 225, the processing circuit 115 may send data or a workload request via Ethernet; at 230, the ToR Ethernet switch 110 may receive the request and route it to the corresponding server 105 among the multiple servers 105; at 235, the request may be received within the server through the (multiple) ports of the network interface circuit 125 (e.g., 100GbE-enabled NIC). At 262, the memory module 135 receives the request from the PCIe connector; at 264, the controller of the memory module 135 processes the request using the local memory; at 250, the controller of the memory module 135 identifies the need to access memory contents (e.g., aggregated memory contents) from different servers; at 252, the controller of the memory module 135 sends a request for the memory contents (e.g., aggregated memory contents) from different servers via the CXL protocol; at 254, the request propagates through the local PCIe connector to the server link switch 112, and then the server link switch 112 sends the request to the second PCIe connector of the second server on the rack; at 266, the second PCIe connector provides access via the CXL protocol to share the aggregated memory, thereby allowing the controller of the memory module 135 to retrieve the memory contents.
[0098] As used herein, a "server" is a computing system that includes at least one program processing circuit (e.g., processing circuit 115), at least one memory resource (e.g., system memory 120), and at least one circuit for providing network connectivity (e.g., network interface circuit 125). As used herein, a "portion" of something means "at least some" of that thing, and thus can mean less than all or all of that thing. Thus, a "portion" of a thing includes the whole thing as a special case, i.e., the whole thing is an example of a portion of the thing.
[0099] The background provided in the "Background" section of this disclosure is included only to set the context, and the content of this section is not admitted to be prior art. Any component or combination of components described (e.g., in any system diagram included herein) can be used to perform one or more operations of any flowchart included herein. Additionally, (i) these operations are example operations and can involve various additional steps not explicitly covered, and (ii) the chronological order of these operations can be changed.
[0100] The term "processing circuit" or "controller mechanism" is used herein to denote any combination of hardware, firmware, and software used to process data or digital signals. Processing circuit hardware can include, for example, application specific integrated circuits (ASICs), general or special purpose central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), and programmable logic devices such as field programmable gate arrays (FPGAs). In a processing circuit as used herein, each function is performed by hardware configured (i.e., hardwired) to perform that function, or by more general hardware (such as a CPU) configured to run instructions stored on a non-transitory storage medium. The processing circuit can be fabricated on a single printed circuit board (PCB) or distributed across several interconnected PCBs. The processing circuit can contain other processing circuits; for example, the processing circuit can include two processing circuits, an FPGA, and a CPU interconnected on a PCB.
[0101] As used herein, "controller" includes circuitry, and a controller may also be referred to as "control circuit" or "controller circuit". Similarly, "memory module" may also be referred to as "memory module circuit" or "memory circuit". As used herein, the term "array" refers to an ordered set of numbers, regardless of how it is stored (e.g., stored in consecutive memory locations or in a linked list). As used herein, when a second number is "within Y% of" a first number, it means the second number is at least (1 - Y / 100) times the first number and at most (1 + Y / 100) times the first number. As used herein, the term "or" shall be construed as "and / or", such that for example, "A or B" means any of "A" or "B" or "A and B".
[0102] As used herein, when a method (e.g., adjustment) or a first quantity (e.g., a first variable) is said to be "based on" a second quantity (e.g., a second variable), it means the second quantity is an input to the method or affects the first quantity. For example, the second quantity may be an input to a function that calculates the first quantity (e.g., the sole input or one of several inputs), or the first quantity may be equal to the second quantity, or the first quantity may be the same as the second quantity (e.g., stored in the same one or more memory locations).
[0103] It will be understood that although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers, and / or portions, these elements, components, regions, layers, and / or portions should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or portion from another element, component, region, layer, or portion. Thus, a first element, component, region, layer, or portion discussed herein may be referred to as a second element, component, region, layer, or portion without departing from the spirit and scope of the inventive concept.
[0104] For ease of description, spatial relationship terms such as "below", "beneath", "under", "below", "above", "on", etc. may be used herein to describe the relationship of one element or feature to another (or others) as shown in the figures. It will be understood that such spatial relationship terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is flipped, the element described as "below" or "beneath" or "under" another element or feature will be oriented "above" that other element or feature. Thus, the example terms "below" and "beneath" can encompass both upward and downward directions. The device may be otherwise oriented (e.g., rotated 90 degrees or oriented in other directions), and the spatial relationship descriptors used herein should be interpreted accordingly. Additionally, it will be understood that when a layer is referred to as being "between" two layers, it may be the only layer between the two layers, or there may be one or more intervening layers.
[0105] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the inventive concept. As used herein, the terms "substantially", "about" and similar terms are used as approximate terms and not as terms of degree, and are intended to account for inherent variations in measured or calculated values that would be recognized by a person of ordinary skill in the art. As used herein, the singular form "a" is also intended to include the plural form, unless the context clearly indicates otherwise. It will also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. When following a list of elements, an expression such as "at least one of" modifies the entire list of elements and not individual elements in the list. Additionally, when describing embodiments of the inventive concept, the use of "may" refers to "one or more embodiments of the present disclosure". Further, the term "exemplary" is intended to refer to an example or illustration. As used herein, the terms "use", "using", and "used" may be considered synonymous with the terms "utilize", "utilizing", and "utilized", respectively.
[0106] It will be understood that when an element or layer is referred to as being “on,” “connected to,” “coupled to,” or “adjacent to” another element or layer, it can be directly on, connected to, coupled to, or adjacent to the other element or layer, or there can be one or more intervening elements or layers. In contrast, when an element or layer is referred to as being “directly on,” “directly connected to,” “directly coupled to,” or “immediately adjacent to” another element or layer, there are no intervening elements or layers.
[0107] Any numerical range recited herein is intended to include all sub-ranges of the same numerical precision subsumed within the recited range. For example, the range “1.0 to 10.0” or “between 1.0 and 10.0” is intended to include all sub-ranges between the recited minimum value 1.0 and the recited maximum value 10.0 (and including the recited minimum value 1.0 and the recited maximum value 10.0), that is, all sub-ranges having a minimum value equal to or greater than 1.0 and a maximum value equal to or less than 10.0, such as, for example, 2.4 to 7.6. Any maximum numerical limitation recited herein is intended to include all lower numerical limitations subsumed therein, and any minimum numerical limitation recited in this specification is intended to include all higher numerical limitations subsumed therein.
[0108] Although exemplary embodiments of systems and methods for managing memory resources have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Accordingly, it will be understood that, except as specifically described herein, systems and methods for managing memory resources constructed in accordance with the principles of the present disclosure can be embodied. The invention is also defined by the appended claims and their equivalents.
Claims
1. A system, comprising: A first server, comprising: A stored program processing circuit, A cache coherent switch, and A first memory module; and A second server; and A server link switch connected to the first server and the second server, Wherein: The first memory module is connected to the cache coherent switch, The cache coherent switch is connected to the server link switch, The stored program processing circuit is connected to the cache coherent switch, and The server link switch is configured to disable the first memory module by the cache coherent switch.
2. The system according to claim 1, wherein the server link switch comprises a Peripheral Component Interconnect Express (PCIe) switch.
3. The system according to claim 1, wherein the server link switch comprises a Compute Express Link (CXL) switch.
4. The system according to claim 3, wherein the server link switch comprises a Top of Rack (ToR) CXL switch.
5. The system according to claim 1, wherein the server link switch is configured to discover the first server.
6. The system according to claim 1, wherein the server link switch is configured to reboot the first server.
7. The system according to claim 1, wherein the server link switch is configured to send data from the second server to the first server and perform flow control on the data.
8. The system according to claim 1, further comprising a third server connected to the server link switch, wherein: The server link switch is configured to: Receive a first packet from the second server, Receive a second packet from the third server, and Send the first packet and the second packet to the first server.
9. The system according to claim 1, further comprising a second memory module connected to the cache coherent switch, wherein the first memory module comprises volatile memory and the second memory module comprises persistent memory.
10. The system according to claim 9, wherein the cache coherent switch is configured to virtualize the first memory module and the second memory module.
11. The system according to claim 10, wherein the first memory module comprises flash memory and the cache coherent switch is configured to provide a flash translation layer to the flash memory.
12. The system according to claim 1, wherein the first server comprises an expansion slot adapter connected to an expansion slot of the first server, the expansion slot adapter comprising: A cache coherent switch; And A memory module slot, The first memory module is connected to the cache coherent switch through the memory module slot.
13. The system according to claim 12, wherein the memory module slot comprises an M.2 slot.
14. The system according to claim 12, wherein: The cache coherent switch is connected to the server link switch through a connector, and The connector is on the expansion slot adapter.
15. A method for performing remote direct memory access in a computing system, wherein, The computing system comprises: A first server, A second server, A third server, and A server link switch connected to the first server, the second server and the third server, The first server comprises: A stored program processing circuit, Cache coherent switch, and a first memory module, The method includes: receiving, by a server link switch, a first packet from a second server, receiving, by the server link switch, a second packet from a third server, and sending the first packet and the second packet to a first server; wherein the server link switch is configured to cause the cache coherent switch to disable the first memory module.
16. The method according to claim 15, further comprising: receiving, by the cache coherent switch, a Remote Direct Memory Access (RDMA) request, and sending, by the cache coherent switch, an RDMA response.
17. The method according to claim 16, wherein receiving the RDMA request includes receiving the RDMA request via the server link switch.
18. The method according to claim 16, further comprising: receiving, by the cache coherent switch, a read command for a first memory address from a storage program processing circuit, converting, by the cache coherent switch, the first memory address to a second memory address, and retrieving, by the cache coherent switch, data at the second memory address from the first memory module.
Citation Information
Patent Citations
Method and system for facilitating high-capacity shared memory using dimm from retired servers
CN110134329A