Multidimensional memory cluster

By constructing a multidimensional memory cluster and adopting full mesh connectivity and CXL connectivity, the challenges of DDR memory in the assembly of large memory pools are solved, achieving efficient node discovery and fault isolation, and improving the scalability and resource management efficiency of the memory pool.

CN116069243BActive Publication Date: 2026-04-21SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2022-10-25
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing DDR memory is difficult to assemble efficiently when building large memory pools, leading to challenging problems.

Method used

A multidimensional memory cluster is constructed using an N-dimensional system with kN nodes. Each node is connected via a full mesh. The host allocates memory according to application requirements and uses compute fast link (CXL) connections and topology identifiers for node management, supporting fault isolation and QoS/CoS features.

Benefits of technology

It achieves efficient node discovery and fault isolation, supports fast access and flexible resource management for large memory pools, reduces complexity and improves the scalability of memory pools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116069243B_ABST
    Figure CN116069243B_ABST
Patent Text Reader

Abstract

A multidimensional storage cluster. In some embodiments, the storage cluster includes: a first node having an external port for connecting to a host; a second node connected to the first node via a first storage-centric connection; the second node storing a service level descriptor; the first node being configured to: receive from the host a first request addressed to the second node for the service level descriptor; and forward the first request to the second node, the second node being configured to: receive the first request; and send a first response to the first node, the first response including the service level descriptor of the second node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority and benefit to U.S. Provisional Application No. 63 / 274,381, filed November 1, 2021, entitled “ADAPTIVE TOPOLOGY DISCOVERYPROTOCOL FOR DISAGREGATED MEMORY ARCHI TECTURE WITH NOVEL INTERCONNECTS”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] One or more aspects of embodiments of this disclosure relate to memory, and more specifically to multidimensional memory clusters. Background Technology

[0004] As computing systems become more powerful and require greater computing resources, larger memory can be used. Dual Data Rate (DDR) memory with a parallel interface may not be easy to assemble; therefore, building large memory pools using DDR memory can be challenging.

[0005] Therefore, an improved architecture is needed for large memory pools. Summary of the Invention

[0006] In some embodiments, the memory cluster is constructed to have a value equal to k. N An N-dimensional system of nodes, where k is the number of nodes in the basic building block cluster. Each node may include a certain amount of memory. These nodes can be connected along each dimension in a fully meshed manner, such that reaching any node from any other node requires at most N hops. Each node may have one or more service level characteristics; for example, it may be a relatively fast node (e.g., measured by latency and throughput) or a relatively slow node. Hosts connected to the memory cluster can allocate memory from the cluster to applications running on the host based on the service level requirements of each application.

[0007] According to one embodiment of this disclosure, a system is provided, comprising: a first node having an external port for connecting to a host; a second node connected to the first node via a first memory-centric connection; the second node storing a service level descriptor; the first node being configured to: receive from the host a first request for the service level descriptor addressed to the second node; and forward the first request to the second node, the second node being configured to: receive the first request; and send a first response to the first node, the first response including the service level descriptor of the second node.

[0008] In some embodiments, the first memory-centric connection is a compute fast link (CXL) connection.

[0009] In some embodiments, the second node also stores a topology identifier, which includes the address of the second node along the first dimension and the address of the second node along the second dimension.

[0010] In some embodiments, the topology identifier further includes: a port identifier for a first port of the second node, and a port identifier for a second port of the second node, the first port of the second node being connected to a third node via a second memory-centric connection, the third node being separated from the second node along a first dimension; and the second port of the second node being connected to a fourth node via a third memory-centric connection, the fourth node being separated from the second node along a second dimension.

[0011] In some embodiments, the bandwidth of the third memory-centric connection is at least twice that of the second memory-centric connection.

[0012] In some embodiments, the second node is configured to disconnect from the third node upon receiving a disconnect command identifying a first port of the second node.

[0013] In some embodiments, the system further includes a host connected to the first node, wherein the host is configured to detect a fault affecting the third node and, in response to detecting the fault, send a disconnect command.

[0014] In some embodiments: the first node is further configured to receive from the host a second request for a topology identifier addressed to the second node; and forward the second request to the second node, the second node being further configured to: receive the second request; and send a second response to the first node, the second response including the topology identifier of the second node.

[0015] In some embodiments, the first node is connected to M other nodes along a first dimension, where M is an integer greater than or equal to 4.

[0016] In some embodiments, the first node is connected to M other nodes along a first dimension, where M is equal to 1 or 2.

[0017] According to one embodiment of this disclosure, a method is provided comprising: receiving a first request via a first node having an external port for connecting to a host, the first request addressing to a second node connected to the first node via a first memory-centric connection, the first request being a request for a service level descriptor stored by the second node; forwarding the first request to the second node via the first node; receiving the first request via the second node; and sending a first response via the second node to the first node, the first response including the service level descriptor of the second node.

[0018] In some embodiments, the first memory-centric connection is a compute fast link (CXL) connection.

[0019] In some embodiments, the second node also stores a topology identifier, which includes the address of the second node along the first dimension and the address of the second node along the second dimension.

[0020] In some embodiments, the topology identifier further includes: a port identifier for a first port of the second node, and a port identifier for a second port of the second node, the first port of the second node being connected to a third node via a second memory-centric connection, the third node being separated from the second node along a first dimension; and the second port of the second node being connected to a fourth node via a third memory-centric connection, the fourth node being separated from the second node along a second dimension.

[0021] In some embodiments, the bandwidth of the third memory-centric connection is at least twice that of the second memory-centric connection.

[0022] In some embodiments, the method further includes: receiving a disconnect command through the second node that identifies a first port of the second node; and disconnecting the second node from the third node.

[0023] In some embodiments, the method further includes: detecting a fault affecting a third node via a host connected to the first node; and sending a disconnect command via the host in response to detecting the fault.

[0024] In some embodiments, the method further includes: receiving a second request for a topology identifier addressed to a second node from a host via a first node; forwarding the second request to the second node via the first node; receiving the second request via the second node; and sending a second response to the first node via the second node, the second response including the topology identifier of the second node.

[0025] In some embodiments, the first node is connected to M other nodes along a first dimension, where M is an integer greater than or equal to 4.

[0026] According to one embodiment of this disclosure, a system is provided, comprising: a first node including a first means for processing and having an external port for connecting to a host; a second node connected to the first node via a first memory-centric connection, the second node including a second means for processing; and the second node storing a service level descriptor; the first means for processing being configured to: receive from the host a first request addressing to the second node for the service level descriptor; and forward the first request to the second node; the second means for processing being configured to: receive the first request; and send a first response to the first node, the first response including the service level descriptor of the second node. Attached Figure Description

[0027] These and other features and advantages of this disclosure will be recognized and understood by referring to the specification, claims and drawings, wherein:

[0028] Figure 1 This is an address word diagram according to an embodiment of the present disclosure;

[0029] Figure 2A This is a memory cluster structure diagram according to an embodiment of the present disclosure;

[0030] Figure 2B This is a memory cluster structure diagram according to an embodiment of the present disclosure;

[0031] Figure 2C This is a memory cluster structure diagram according to an embodiment of the present disclosure;

[0032] Figure 2D This is a memory cluster structure diagram according to an embodiment of the present disclosure;

[0033] Figure 3 It is a table of cluster characteristics according to embodiments of this disclosure;

[0034] Figure 4A This is a schematic diagram of a memory cluster having segments for resource management according to an embodiment of the present disclosure;

[0035] Figure 4B This is a schematic diagram of a memory cluster segmented by service level characteristics according to an embodiment of the present disclosure;

[0036] Figure 4C This is a schematic diagram of a memory cluster having segments for fault isolation according to an embodiment of the present disclosure;

[0037] Figure 5 This is a schematic diagram of a node with 15 routing port groups according to an embodiment of the present disclosure, which covers up to the 5th dimension of 1K nodes;

[0038] Figure 6 It is a table of possible node positions in up to three dimensions according to embodiments of this disclosure;

[0039] Figure 7 This is an illustration of a two-dimensional cluster according to an embodiment of the present disclosure;

[0040] Figure 8 This is a flowchart of a method according to an embodiment of the present disclosure; and

[0041] Figure 9 This is a hybrid block diagram and flowchart illustrating process steps overlaid on a block diagram including a host and a cluster of four nodes, according to embodiments of the present disclosure. Detailed Implementation

[0042] The detailed description set forth below with reference to the accompanying drawings is intended as a description of exemplary embodiments of the multidimensional memory clusters provided in accordance with this disclosure, and is not intended to represent the only form in which this disclosure can be constructed or utilized. The description illustrates features of this disclosure in conjunction with the illustrated embodiments. However, it should be understood that the same or equivalent functionality and structure may be implemented through different embodiments, which are also intended to be included within the scope of this disclosure. As shown elsewhere herein, similar element designations are intended to indicate similar elements or features.

[0043] In some embodiments, the systems disclosed herein contribute novel and efficient addressing schemes to cluster architectures that scale upwards or outwards in a dimension-oriented (or “layer-oriented”) manner. The systems disclosed herein provide a method for abstracting resource pools with a large number of memory or storage nodes in an exceptionally efficient manner. By applying QoS / CoS and fault isolation capabilities, these methods can lead to faster node discovery. Therefore, latency control is possible, and efficient hardware design (e.g., receive (Rx) buffer and retry buffer size determination) is permitted. Some embodiments can keep complexity as low as possible, even as capacity grows with higher layers or dimensions (e.g., as the cluster size scales upwards or outwards). Some embodiments involve recursively applying the same addressing rules (e.g., hypertoroidal full mesh) as the cluster size grows in layers or dimensions. This approach enables the most efficient node discovery protocol for large, distributed memory or storage resource pools.

[0044] The topology information provided in some embodiments enables deterministic latency control for each layer or dimension of the full mesh cluster; this allows memory or storage pool access designs for each segment or region to be synchronized if needed.

[0045] Topology information from some embodiments allows for QoS / CoS engineering of memory or storage regions at each layer or dimension (e.g., hot, warm, and cold layers). This topology information enables disaster control; effective fault isolation can be used in layer-oriented or dimension-oriented cluster architectures.

[0046] In some embodiments, the addressing scheme combines several feature components, including: node network topology, Quality of Service (QoS) or Class of Service (CoS) characteristics, and memory addresses for memory operations. This addressing scheme can provide an efficient and simplified method for managing large pools of memory nodes that scale up or out. As used herein, a "node" or "memory node" is a circuit that includes memory (e.g., dynamic random access memory (DRAM) or flash memory) and has one or more ports (e.g., electrical ports or optical ports) for connecting to other nodes or one or more hosts.

[0047] Some embodiments may use a four-node (or “quadruple-node”) full mesh and dimension-driven hypertoroidal architecture. The dimensional information of the nodes and the port information linked on the nodes can be part of the addressing scheme (discussed in further detail below). In some embodiments, (i) dimensional information, (ii) the node ID in a given dimension, and (iii) the link ID, for example, for a four-node full mesh connection, are components of the addressing scheme for the cluster architecture. Along with these components, QoS / CoS features are added to provide more useful information, enabling a richer range of services to be available for the data center (DC).

[0048] For node addressing schemes, dimension and node ID information can be considered together. In a four-node full-mesh cluster, which serves as a basic building block in some embodiments, two bits can generate four unique IDs that can identify each node in the four-node full mesh. As the cluster grows, this basic building block rule is recursively applied to each additional dimension (see...). Figures 2A-2D (This will be discussed in more detail below). Each dimension's address space uses a 2-bit field, therefore each increasing dimension requires an additional 2-bit field, concatenated with the preceding dimensions. Dimension information is embedded in the 2-bit field of the node's address space (see...). Figure 1 (This will be discussed in more detail below). The first 2-bit field, starting from the least significant bit (LSB) side, represents the lowest dimension (see...). Figure 1 The node address space is divided into dimension 1 (D1), and a 2-bit field is added for each additional dimension in the direction of the most significant bit (MSB). Some embodiments use up to 10 dimensions, which covers up to approximately 1 million nodes. Although in Figure 1In this embodiment, 32 bits are allocated to the node address space, but only 20 bits are used for the 10-dimensional node ID. Therefore, 12 bits are reserved for future use.

[0049] Each routing port (or simply "port") on each node can be addressed individually. In the routing port addressing scheme, dimension information and port identifier (ID) information are considered together. In a four-node full-mesh cluster, which serves as a basic building block in some embodiments, three bits in each node can represent three unique IDs, which can be assigned to each port using a one-hot encoding scheme. The first dimension contains links to port IDs 0, 1, and 2; the second dimension contains links to port IDs 3, 4, and 5, and so on. In some embodiments, this pattern persists for up to 10 dimensions. Each node is connected to three links that extend to the other three nodes within the dimension. The three ports of each node act as link connectors. One-hot encoding techniques can be employed to provide efficient link management (e.g., fault isolation or data isolation).

[0050] Each dimension's routing port address space consists of a 3-bit field, and an additional 3-bit field is added for each dimension, concatenated with the preceding dimensions. Dimension information is embedded in the 3-bit field of the routing port address space (see [link to documentation]). Figure 1 The first 3-bit field starting from the LSB side represents the lowest dimension (see...). Figure 1 The routing port address space is divided into dimension 1 (D1), and a 3-bit field is added for each additional dimension in the direction of the most significant bit (MSB). Some embodiments include up to 10 dimensions, which covers up to an estimated 30 million ports. Although in Figure 1 In this embodiment, 32 bits are allocated to the routing port address space, but 30 bits are used for the 10-dimensional port ID. Therefore, 2 bits are reserved for future use.

[0051] Figure 1 This illustrates a high-level overview of the addressing scheme for the adaptive topology discovery protocol. The entire addressing scheme consists of a 32-byte address word divided into four sub-address spans (each span containing 8 bytes or 64 bits). The topology ID sub-address span internally comprises a node address space (32 bits) and a routing port space (32 bits). For a cluster architecture based on a four-node full-mesh basic building block, 2 bits can be used to represent the four nodes in each dimension.

[0052] In the node address space within the topology ID sub-address span, "Dn" and "bb" represent the node ID in the dimension and the given dimension, respectively. Thus, a node address can include the address of a node along the first dimension, the address of a node along the second dimension, and so on. For the routing port space within the topology ID sub-address span, "Dn" and "bbb" represent the port ID in the dimension and the given dimension, respectively. One-hot encoding is used for port IDs. As described above, the addressing scheme can cover up to 10 dimensions, involving 1 million nodes and 30 million ports.

[0053] The 32-byte address word also includes a QoS / CoS addressing scheme (also known as a service level descriptor) with a 64-bit address space, which can be used to classify the memory pool into cold, warm, and hot tiers based on service level characteristics. For example, a node's temperature can refer to the read latency or write latency of the node's storage medium, or throughput, or a quality factor based on read latency and write latency combined with throughput, where hot nodes are generally faster than cold nodes (e.g., with lower latency and higher throughput). Reserved values ​​(e.g., 0) in the word allocated to each tier can indicate that a node is not in that tier. Multiple sub-tiers are available within each tier. For example, an 8-bit hot tier includes 255 possible sub-tiers (in addition to the reserved values), and for example, 0 can mean the node is not hot, 1 can mean the node is almost not hot (only slightly warmer than the warmest warm node), and 255 can mean the node is very hot. The last sub-address span of the 32-byte address word carries the memory address used for memory operations.

[0054] In some embodiments, each node stores its topology identifier and service level descriptor. This information can be written to non-volatile memory within the node when the cluster is assembled (e.g., when physical connections between nodes are established). When a host sends a request for a topology identifier or service level descriptor to a node (which may be relayed by other nodes), the node can send a response with the topology identifier or service level descriptor (also relayed by other nodes), or it can disable one or more ports of the node (as discussed in further detail below).

[0055] Figures 2A-2D This illustrates how the number of dimensions in a memory cluster increases from n to n+1 dimensions. Figure 2AA series of clusters are shown, illustrating how cluster size can be increased by increasing the number of dimensions. First, nodes 205 (each of which can be considered a zero-dimensional cluster) are interconnected via a full mesh to create cluster 210 (which can be referred to as a “basic building block cluster”), a one-dimensional cluster. The connections or “links” linking nodes 205 together can be memory-centric connections (discussed in further detail below). Then, layer 215 is created by increasing the number of basic building blocks to the number of nodes in the initial blocks. The building block clusters 210 in layer 215 are interconnected via the same node ID; for example, all nodes with node ID 0 are connected together via a full mesh, forming a partially connected two-dimensional cluster 220. The partially connected two-dimensional cluster 225 is identical to the partially connected two-dimensional cluster 220, but it is drawn as one of the basic building block clusters 210 with translation and rotation, making the full mesh connections along the second dimension between the nodes with node ID 0 of the four basic building block clusters 210 more easily identifiable. At least one of the nodes may have (in addition to routing ports for connecting to other nodes) an external port for connecting to host 207. Figure 2C ).

[0056] Figure 2B The diagram shows: (i) a partially connected two-dimensional cluster 220, (ii) a partially connected two-dimensional cluster 230, wherein nodes 205 with node ID1 are connected together by a full mesh, (iii) a partially connected two-dimensional cluster 235, wherein nodes 205 with node ID2 are connected together by a full mesh, and (iv) a partially connected two-dimensional cluster 230, wherein nodes 205 with node ID3 are connected together by a full mesh. Figure 2C A fully connected 2D cluster 245 is shown, where nodes 205 with node ID 0 are connected together via a full mesh, nodes 205 with node ID 1 are connected together via a full mesh, nodes 205 with node ID 2 are connected together via a full mesh, and nodes 205 with node ID 3 are connected together via a full mesh. In this example (which uses 4-node full mesh connections as the basic building block cluster), every three links extending from a node in a given dimension create a full mesh together with other nodes along that dimension. Thus, by adding three links to each node of a cluster with n dimensions, another layer and dimension can be created to expand upwards (or outwards) towards a larger cluster.

[0057] Figure 2D This shows another instance of extending to higher dimensions. Figure 2DThe diagram shows: (i) node 205 (zero-dimensional cluster), (ii) one-dimensional cluster 210, (iii) partially connected two-dimensional cluster 220, and (iv) partially connected three-dimensional cluster 250. The two-dimensional cluster 220 and the three-dimensional cluster 250 are shown in a partially connected form to make the connections easier to identify; in some embodiments, each of these clusters has a full mesh connection along each basic building block cluster 210 in each dimension. The process of increasing the number of dimensions, one example of which is... Figure 2D As shown, this can be repeated an arbitrary number of times to produce arbitrarily large clusters. Each additional dimension increases the number of nodes in the new cluster by four times the number of nodes in the lower-dimensional cluster. For example, from one dimension to two dimensions, the number of nodes increases from 4 to 16, and from two dimensions to three dimensions, the number of nodes increases from 16 to 64. Furthermore, in this example, each additional dimension increases the cluster's storage capacity by a factor of four. The multiplication factor depends on the number of nodes in the basic building block cluster; for example, if the basic building block cluster includes three nodes instead of four, then each additional dimension increases both the number of nodes and the total storage capacity by a factor of three.

[0058] The number of dimensions corresponds to the maximum number of hops in communication from any node to any other node. This means that a one-dimensional cluster can have a maximum of one hop, a two-dimensional cluster can have a maximum of two hops, and a three-dimensional cluster can have a maximum of three hops.

[0059] In operation, host 207 may (e.g., at startup) send a request for a service level descriptor (SSD) for each of nodes 205, and each node 205 may send a response to host 207 including the requested SSD. If the node 205 addressed by the request is not directly connected to host 207, both the request and the response may be forwarded by other nodes 205, including those directly connected to host 207 ((i) to the node 205 addressed by the request, and (ii) back to host 207 from the node 205 addressed by the request). The cluster may be densely distributed, with most or all nodes supported by the addressing scheme present. Thus, when a request is issued at startup, the host can traverse all supported addresses. If a node is absent or inactive, a request from the host may time out, or if another node along the path to the addressed node knows that the addressed node is absent or inactive, that other node may send a response notifying the host that the addressed node is absent or inactive.

[0060] Figure 3This is a table showing the cluster properties as a function of the number of dimensions. The table displays the number of dimensions, the number of nodes, the number of full mesh networks, the maximum number of hops (the maximum number of hops for communication from any node in the N-dimensional cluster to any other node), the number of bits in the node address, and the number of links (or ports) per node. The table below summarizes the properties of a four-node full mesh cluster architecture as the number of dimensions increases. For example, in the case of five dimensions, the possible number of nodes is 1024. These 1024 nodes comprise (e.g., constitute) 256 four-node full mesh networks along any of the five dimensions. Within these 1024 nodes, communication between any two nodes can be accomplished with a maximum of 5 hops.

[0061] Some embodiments provide efficient solutions for managing large pools of memory or storage resources through adaptive topology discovery protocols. For example, Figure 4A This illustrates a region of memory or storage segmentation for efficient resource management. This can be implemented, for example, using logical segmentation, where regions are created within a memory or storage resource pool. Segmentation can be accomplished using an adaptive topology discovery protocol. In this example, 64 nodes are divided into 4 regions. Figure 4A The regions are designated as A, B, C, and D. Furthermore, sub-regions can be created within each region. In this case, there are four sub-regions within each region. Figure 4A Segmentation up to the third layer or dimension is illustrated. In some embodiments, logical segmentation is performed considering the link bandwidth of each layer or dimension. For example, links with higher bandwidth or lower latency can be used along a particular dimension (e.g., the bandwidth of a link along a second dimension may be at least twice the bandwidth of a link along a first dimension), and a memory region comprising a cluster connected along that dimension can be used to provide high-throughput, low-latency storage. For example, in a two-dimensional network where connections along the second dimension have higher bandwidth than connections along the first dimension, if a host is directly connected to a first node, and second, third, and fourth nodes are connected to the first node along the second dimension, then a region comprising the first, second, third, and fourth nodes can be used by the host as a high-throughput, low-latency storage region.

[0062] Figure 4B The QoS / CoS area memory or storage segment is shown. Figure 4BAn example of CoS allocation for each region is shown. This segmentation can be accomplished using an adaptive topology discovery protocol, where CoS allocation for each region can be performed via a layered or dimensional approach (e.g., for hot, warm, and cold layers). In this case, three categories are assigned to four regions. For example, the hot, warm, and cold data categories could be assigned to (i) Dynamic Random Access Memory (DRAM), (ii) Storage Class Memory / Persistent Memory (SCM / PM), or (iii) Solid State Drive (SSD) resource pools, respectively. QoS (e.g., memory type with load status - CXL Type 3 memory extension solution) can also be applied via the adaptive topology discovery protocol. In some embodiments, categories (e.g., data type or memory type) are allocated based on link bandwidth.

[0063] In some embodiments, a region memory or storage segment is used for fault isolation. Figure 4C An example of fault isolation in a memory or storage resource pool is illustrated. Fault isolation, or data isolation, can be accomplished using one-hot encoding of port IDs via an adaptive topology discovery protocol. Assuming region D becomes a faulty region, it can be isolated by disabling the three links shown by the dashed lines. This can be achieved by the host (which can detect the fault in region D) sending a disconnect command to each node with a connection to region D, the disconnect command identifying each port through which the node is connected to region D. Each node receiving such a disconnect command can then disconnect from every node in region D. Nodes in other regions are not affected by the fault and can continue to provide service. Figure 4C In this example, the solution is implemented in the third dimension. However, this fault isolation mechanism can be effectively applied to every layer or dimension of the cluster hierarchy.

[0064] Figure 5 The diagram illustrates the number of links in a node of a five-dimensional, four-node full mesh cluster. Each node has three links along each of the five dimensions (in order to be part of a full mesh network along each of the five dimensions), and Figure 5 A total of 15 links are shown. In the case of memory (or storage) nodes, these links can be used for large-scale memory (or storage) pools.

[0065] Figure 6 The table displays possible node locations in up to three dimensions. Depending on the node location, the link medium (electric or optical) can be applied differently. Dimensions are embedded in a 2-bit field of the node address. One-hot encoding is used to identify routing ports, and a 3-bit field is assigned to each node, with one bit mapped to each port (as shown above). Figure 1 (as explained in the context).

[0066] Figure 7 A memory pool with a four-node full-mesh hypertoroidal cluster architecture is shown. Figure 7 This illustrates scaling up to a cluster comprising 16 nodes 205 using two layers or dimensions, with a maximum of two hops required for communication between any two nodes 205. Thus, the 16 nodes 205 can be a larger memory object or pool, and all sub-memory objects or pools can be synchronously accessed with a maximum latency of two hops. Figure 7 In the embodiments, each of the nodes 205 includes a controller 705 (which may be processing circuitry or means for processing, (discussed in further detail below)), one or more ports (links) 710 for connection to other nodes 205, and a memory 715.

[0067] exist Figure 7 Links drawn with thin lines represent links along the first dimension, while links drawn with thick lines represent links along the second dimension. The link speed for each layer or dimension can vary depending on the design or implementation, and the choice of link speed can be used as a CoS characteristic of the memory. Various approaches to link engineering can be employed. For example, based on bandwidth and latency requirements, the link medium can be electrical or optical (e.g., silicon photonics). Alternatively, the link medium can be determined based on link distance and latency (this approach can also be used to implement QoS / CoS characteristics). Nodes can be assembled in various ways—a single node on a single board (e.g., dimension 0), multiple nodes on a single sled (e.g., dimension 1), multiple nodes on multiple sleds (e.g., dimensions 2 or 3), or multiple sleds in a pod (e.g., dimensions 3 and above), and the link medium (electric / optical) can be chosen depending on the size of the cluster.

[0068] Pluggable methods can be used to integrate adaptive topology discovery protocols with other existing protocols. Adaptive topology discovery protocols according to embodiments described herein, such as those for a four-node full-mesh hypertoroidal cluster architecture, can be applied to existing protocols for efficient topology information discovery purposes. Furthermore, existing protocols can leverage features of some embodiments described herein. Existing protocols can employ some embodiments through pluggable methods, resulting in reduced complexity, lower overhead costs, and the ability to use modular approaches (with increased flexibility and composability).

[0069] For example, the following existing protocols can be considered for use in a pluggable manner to adopt this protocol and for providing scalable solutions. Some embodiments may utilize PCIe through extensions to the Physical Function (PF) and Virtual Function (VF) of the Fast Peripheral Component Interconnect (PCIe). Some embodiments may utilize Compute Fast Link (CXL) through extensions to Input-Output (CXL.io), PF (Physical Function), and VF (Virtual Function). Some embodiments may utilize Gen-Z links, or Non-Volatile Memory Fast Fiber (NVMe-oF) links, or other types of links (e.g., NVLink). TM This includes Open Coherent Accelerator Processor Interface (OpenCAPI) or Cache Coherent Interconnect for Accelerators (CCIX). As used herein, "memory-centric connectivity" refers to connectivity suitable for connecting the host to memory (e.g., with sufficiently low latency and sufficiently high throughput), such as PCIe, CXL, Gen-Z, NME-oF, NVLink, etc. TM Or OpenCAPI.

[0070] Some embodiments described herein use a cluster of basic building blocks comprising four nodes; however, this disclosure is not limited to this size of basic building blocks, and typically, the number of nodes in a basic building block can be any positive integer k greater than 1. The number of nodes in such a cluster may be k. N And the total number of ports can be N(k–1)k N .

[0071] Figure 8 This is a flowchart of a method according to some embodiments. The method includes: at 805, receiving a first request via a first node having an external port for connecting to a host, the first request addressing to a second node connected to the first node via a first memory-centric connection, the first request being a request for a service level descriptor stored by the second node; at 810, forwarding the first request to the second node via the first node; at 815, receiving the first request via the second node; and at 820, sending a first response via the second node to the first node, the first response including the service level descriptor of the second node. Figure 9A set of processing steps is shown overlaid on a block diagram of a cluster including host 207 and four nodes 205. This set of processing steps comprises three main phases: (i) a node discovery phase (which may be performed according to a PCIe / CXL enumeration procedure or a similar procedure), (ii) a detection phase, including the detection of Service Level Descriptors (SLDs) and formal topology, and (iii) a connectivity phase, where dynamic memory configuration may be performed using Set_Config(), Get_Config(), Add_Cap(), or Remove_Cap() operations (or other such operations). Each of these three phases may include a request 825 sent by host 207 (which may be a request to discover or detect cluster attributes), the request being forwarded at 830, and a response to the request by one of the nodes 205 at 835.

[0072] As used in this article, a “service level descriptor” is a measure of node performance (e.g., latency or throughput). As used in this article, a “topology identifier” is a node’s location (e.g., coordinates) within an N-dimensional network and may include other information, such as a number of bits corresponding to each routing port of the node.

[0073] As used herein, “a portion” of something means “at least some” of that thing, and can also mean less than or all of that thing. Similarly, “a portion” of a thing includes the whole thing as a special case, i.e., the whole thing is an instance of a portion of that thing. As used herein, when a second quantity is “within Y” of a first quantity X, it means that the second quantity is at least XY and at most X+Y. As used herein, when a second number is “within Y%” of a first number, it means that the second number is at least (1-Y / 100) times the first number and at most (1+Y / 100) times the first number. As used herein, the term “or” should be interpreted as “and / or”, such that, for example, “A or B” means any one of “A” or “B” or “A and B”.

[0074] The background provided in the background section of this disclosure is included only to set the context, and the content of that section is not considered prior art. Any component or combination of components described (e.g., in any system diagram included herein) may be used to perform one or more operations of any flowchart included herein. Furthermore, (i) the operations are exemplary and may include various additional steps not explicitly covered, and (ii) the temporal order of the operations may vary.

[0075] The terms “processing circuitry” and “means for processing” are used herein to mean any combination of hardware, firmware, and software used for processing data or digital signals. Processing circuitry hardware may include, for example, application-specific integrated circuits (ASICs), general-purpose or special-purpose central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), and programmable logic devices such as field-programmable gate arrays (FPGAs). As used herein, in processing circuitry, each function is performed by hardware configured to perform that function (i.e., hardwired), or by more general-purpose hardware (such as a CPU) configured to execute instructions stored in a non-transitory storage medium. Processing circuitry may be fabricated on a single printed circuit board (PCB) or distributed across several interconnected PCBs. Processing circuitry may include other processing circuitry; for example, processing circuitry may include two processing circuits, an FPGA and a CPU, interconnected on a PCB.

[0076] As used herein, when a method (e.g., adjustment) or a first quantity (e.g., a first variable) is referred to as “based on” a second quantity (e.g., a second variable), it means that the second quantity is an input to the method or affects the first quantity. For example, the second quantity may be an input to a function that computes the first quantity (e.g., a unique input, or one of several inputs), or the first quantity may be equal to the second quantity, or the first quantity may be the same as the second quantity (e.g., stored in one or more locations in memory that are the same as the second quantity).

[0077] It should be understood that although the terms "first," "second," "third," etc., may be used herein to describe various elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are used only to distinguish one element, component, region, layer, or section from another element, component, region, layer, or section. Therefore, the first element, component, region, layer, or section discussed below may be referred to as the second element, component, region, layer, or section without departing from the spirit and scope of the invention.

[0078] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the inventive concept. As used herein, the terms “substantially,” “about,” and similar terms are used as approximations rather than terms of degree and are intended to describe inherent variations in measured or calculated values ​​that would be recognized by one of ordinary skill in the art.

[0079] As used herein, the singular forms “a” and “an” are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that, when used in this specification, the terms “comprise” and / or “comprising” specify the presence of stated features, integrals, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or combinations thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items. When expressions such as “at least one” are used after a list of elements, the entire list of elements is modified and no individual element in the list is modified. Furthermore, when describing embodiments of the inventive concept, the use of “may” means “one or more embodiments of this disclosure.” Additionally, the term “exemplary” is intended to refer to an instance or illustration. As used herein, the terms “use,” “using,” and “used” may be considered synonymous with the terms “utilize,” “utilizing,” and “utilized,” respectively.

[0080] It should be understood that when an element or layer is referred to as being "on," "connected to," "linked to," or "adjacent to" another element or layer, it may be directly "on," "connected to," "linked to," or "adjacent to" another element or layer, or one or more intermediate elements or layers may exist. Conversely, when an element or layer is referred to as being "directly on," "directly connected to," "directly linked to," or "directly adjacent to" another element or layer, no intermediate elements or layers exist.

[0081] Any numerical ranges listed herein are intended to include all subranges with the same numerical precision contained within the listed range. For example, ranges “1.0 to 10.0” or “between 1.0 and 10.0” are intended to include all subranges between (and inclusive) the listed minimum value of 1.0 and the listed maximum value of 10.0, i.e., minimum values ​​equal to or greater than 1.0 and maximum values ​​equal to or less than 10.0, such as 2.4 to 7.6. Similarly, ranges described as “within 35% of 10” are intended to include all subranges between (and inclusive) the listed minimum value of 6.5 (i.e., (1-35 / 100) multiplied by 10) and the listed maximum value of 13.5 (i.e., (1+35 / 100) multiplied by 10), i.e., minimum values ​​equal to or greater than 6.5 and maximum values ​​equal to or less than 13.5, such as 7.4 to 10.6. Any maximum numerical limits listed herein are intended to include all lower numerical limits contained therein, and any minimum numerical limits listed in this specification are intended to include all higher numerical limits contained therein.

[0082] It should be understood that when a component is referred to as “directly connected” or “directly coupled” to another component, there is no intermediate component. As used herein, “generally connected” means a connection via an electrical path that may contain any intermediate component, including those whose presence qualitatively alters the behavior of the circuit. As used herein, “connection” means: (i) “directly connected” or (ii) connected to an intermediate component, which is a component that qualitatively affects the behavior of the circuit (e.g., a low-value resistor or inductor, or a short section of a transmission line).

[0083] Although exemplary embodiments of multidimensional memory clusters have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Therefore, it should be understood that multidimensional memory clusters constructed in accordance with the principles of this disclosure can be embodied in addition to those specifically described herein. The invention is also defined in the appended claims and their equivalents.

Claims

1. A system comprising: The first node has an external port for connecting to the host; The second node is connected to the first node via a first memory-centric connection; The second node provides a storage service level descriptor; The first node is configured as follows: The host receives a first request for the service level descriptor addressed to the second node; and Forward the first request to the second node. The second node is configured as follows: Receive the first request; and A first response is sent to the host via the first node, the first response including the service level descriptor of the second node. The second node also stores a topology identifier, which includes the address of the second node along the first dimension and the address of the second node along the second dimension. The topology identifier further includes: The port identifier of the first port of the second node, and The port identifier of the second port of the second node. The first port of the second node is connected to the third node via a second memory-centric connection, the third node being separated from the second node along the first dimension; and The second port of the second node is connected to the fourth node via a third memory-centric connection, the fourth node being separated from the second node along the second dimension.

2. The system of claim 1, wherein the bandwidth of the third memory-centric connection is at least twice the bandwidth of the second memory-centric connection.

3. The system of claim 2, wherein the second node is configured to disconnect the second node from the third node upon receiving a disconnect command identifying the first port of the second node.

4. The system of claim 3, further comprising a host connected to the first node, wherein the host is configured to: Detect the fault affecting the third node, and In response to the detection of the fault, the disconnect command is sent.

5. The system as claimed in claim 1, wherein: The first node is also configured as Receive a second request from the host for the topology identifier addressed to the second node; and Forward the second request to the second node. The second node is also configured as follows: Receive the second request; and A second response is sent to the first node, the second response including the topology identifier of the second node.

6. The system of claim 1, wherein the first node is connected to M other nodes along the first dimension, where M is an integer greater than or equal to 4.

7. The system of claim 1, wherein the first node is connected to M other nodes along the first dimension, where M is equal to 1 or 2.

8. The system of claim 1, wherein the host is configured to allocate memory from the first node or the second node to an application running on the host.

9. The system of claim 1, wherein: The first node is configured to receive the first request from the host via the external port; and The second node is configured to receive the first request from the first node via the first memory-centric connection.

10. The system of claim 9, wherein: The first memory-centric connection is a compute fast link (CXL) connection; and The external port includes an electrical port or an optical port.

11. A method comprising: A first request is received by a first node having an external port for connecting to a host, the first request addressing to a second node connected to the first node via a first memory-centric connection, the first request being a request for a service level descriptor stored by the second node. The first request is forwarded to the second node through the first node; The first request is received through the second node; as well as The second node sends a first response to the host via the first node, the first response including the service level descriptor of the second node. The second node also stores a topology identifier, which includes the address of the second node along the first dimension and the address of the second node along the second dimension. The topology identifier further includes: The port identifier of the first port of the second node, and The port identifier of the second port of the second node. The first port of the second node is connected to the third node via a second memory-centric connection, the third node being separated from the second node along the first dimension; and The second port of the second node is connected to the fourth node via a third memory-centric connection, the fourth node being separated from the second node along the second dimension.

12. The method of claim 11, wherein the bandwidth of the third memory-centric connection is at least twice the bandwidth of the second memory-centric connection.

13. The method of claim 12, further comprising: The second node receives a disconnect command that identifies the first port of the second node; as well as Disconnect the second node from the third node.

14. The method of claim 13, further comprising: The fault affecting the third node is detected by the host connected to the first node; as well as In response to the detection of the fault, the disconnect command is sent through the host.

15. The method of claim 11, further comprising: The first node receives a second request from the host for the topology identifier, addressed to the second node; as well as The second request is forwarded to the second node through the first node; The second request is received through the second node; as well as The second node sends a second response to the first node, the second response including the topology identifier of the second node.

16. A system comprising: The first node includes a first device for processing and has an external port for connecting to a host. The second node is connected to the first node via a first memory-centric connection, and the second node includes a second means for processing. The second node provides a storage service level descriptor; The first device for processing is configured as follows: The host receives a first request for the service level descriptor addressed to the second node; and Forward the first request to the second node. The second device for processing is configured as follows: Receive the first request; and A first response is sent to the host via the first node, the first response including the service level descriptor of the second node. The second node also stores a topology identifier, which includes the address of the second node along the first dimension and the address of the second node along the second dimension. The topology identifier further includes: The port identifier of the first port of the second node, and The port identifier of the second port of the second node. The first port of the second node is connected to the third node via a second memory-centric connection, the third node being separated from the second node along the first dimension; and The second port of the second node is connected to the fourth node via a third memory-centric connection, the fourth node being separated from the second node along the second dimension.

Citation Information

Patent Citations

  • Storing and retrieving training data for models in a data center

    US20190034763A1

  • High Availability Storage Access Using Quality Of Service Based Path Selection In A Storage Area Network Environment

    US20190102093A1