Mirrored and segregated memory in a clustered environment
By allocating and mirroring remote memory on alternative nodes within a cluster using a hypervisor and disjointed memory manager, the system addresses the reliability and availability issues in disaggregated memory environments, ensuring redundancy and improved performance.
Patent Information
- Application Number
- JP2025530690
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-04
- Publication Date
- 2026-01-06
AI Technical Summary
Existing computer processing systems in clustered environments face challenges with reduced reliability and availability due to the use of disaggregated memory, where the loss of remote memory can be catastrophic, and current solutions either rely solely on local memory or tolerate reduced reliability, leading to undesirable configuration overhead or management complexity.
Implementing a hypervisor and disjointed memory manager to allocate and mirror remote memory on alternative nodes within a cluster, ensuring redundancy and resilience by maintaining mirrored memory and adjusting allocations to improve performance and reliability.
Enhances the resilience and performance of clustered environments by providing mirrored memory, protecting against catastrophic memory loss and improving reliability and availability without losing the benefits of disaggregated memory.
Smart Images

Figure 2026500111000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The embodiments described herein relate generally to computer processing systems, and more particularly to computer processing systems that implement mirrored, disjointed memory in a clustered environment.
[0002] Cloud computing and cloud storage provide users with the ability to store and process their data in third-party data centers. Cloud computing facilitates the ability to quickly and easily provision virtual machines for customers without requiring the customers to purchase hardware or provide floor space for physical servers. Generally, virtual machines running as guests under the control of a host hypervisor rely on the hypervisor to transparently provide virtualization services for the guest. These services include memory management, instruction emulation, and interrupt handling.
[0003] The term "hypervisor" refers to a processing environment or platform service that manages and allows one or more virtual machines to run using multiple (and possibly different) operating systems on the same host machine. It should be understood that deploying a virtual machine includes the process of installing the virtual machine and the process of activating (or starting) the virtual machine. In another example, deploying a virtual machine includes the process of activating (or starting) the virtual machine (e.g., if the virtual machine was previously installed or already exists).
[0004] Some computer processing systems (or simply "processing systems"), such as nodes in cloud computing systems, include the ability to dynamically share memory across multiple processing systems via a dedicated network fabric. Direct cabling of a processor chip in one processing system to a processor chip in another processing system may link the processing systems together, allowing them to share each other's physical memory. This capability is known as "memory inception," "memory disaggregation," or "memory clustering." This technique may be used for multiple use cases, including in composable data centers, where resources, including memory, may be dynamically allocated and shared across systems. Summary of the Invention
[0005] In one exemplary embodiment, an exemplary computer-implemented method for mirroring memory in a disjointed memory clustering environment is provided. The method includes a hypervisor allocating disjointed memory to a virtual machine including remote disjointed memory, the virtual machine being a node of a cluster in the disjointed memory clustering environment. The method further includes a disjointed memory manager allocating mirrored memory to the remote disjointed memory and mirroring the remote disjointed memory on an alternative node of the cluster in the disjointed memory clustering environment. The method further includes the disjointed memory manager maintaining the mirrored memory in response to a memory access occurring. The method further includes the disjointed memory manager modifying memory utilization across the cluster in response to detecting a memory allocation adjustment. The method further includes implementing corrective action in response to detecting a failure that results in a loss of access to the remote disjointed memory. The method improves the resilience and / or performance of a cluster and / or node. Furthermore, the method improves cluster functionality by providing mirrored memory. For example, the method provides for preventing catastrophic memory loss and improving reliability and availability within a segregated memory environment without losing the benefits of having segregated memory available.
[0006] In another exemplary embodiment, a system is provided that includes a memory having computer-readable instructions and a processing device for executing the computer-readable instructions. The computer-readable instructions control the processing device to perform operations for mirroring memory in an isolated memory clustering environment. The operations include a hypervisor allocating isolated memory to a virtual machine that includes remote isolated memory, the virtual machine being a node of a cluster in the isolated memory clustering environment. The operations further include a isolated memory manager allocating mirrored memory to the remote isolated memory and mirroring the remote isolated memory on an alternative node of the cluster in the isolated memory clustering environment. The operations further include the isolated memory manager maintaining the mirrored memory in response to a memory access occurring. The operations further include the isolated memory manager modifying memory utilization across the cluster in response to detecting a memory allocation adjustment. The operations further include implementing a corrective action in response to detecting a failure that results in a loss of access to the remote isolated memory. The system improves the resilience and / or performance of a cluster and / or nodes. Additionally, the system improves cluster functionality by providing mirrored memory. For example, the system provides protection against catastrophic memory loss and improved reliability and availability within a disaggregated memory environment without losing the benefits of disaggregated memory.
[0007] In another exemplary embodiment, a computer program product is provided comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions executable by a processor to cause the processor to perform operations for mirroring memory in an isolated memory clustering environment. The operations include a hypervisor allocating isolated memory to a virtual machine including remote isolated memory, the virtual machine being a node of a cluster in the isolated memory clustering environment. The operations further include a isolated memory manager allocating mirrored memory to the remote isolated memory and mirroring the remote isolated memory on an alternative node of the cluster in the isolated memory clustering environment. The operations further include the isolated memory manager maintaining the mirrored memory in response to a memory access occurring. The operations further include the isolated memory manager modifying memory utilization across the cluster in response to detecting a memory allocation adjustment. The operations further include implementing a corrective action in response to detecting a failure resulting in a loss of access to the remote isolated memory. The computer program product improves the resilience and / or performance of a cluster and / or nodes. Furthermore, the computer program product improves the functionality of a cluster by providing mirrored memory. For example, the computer program product provides protection against catastrophic memory loss and improved reliability and availability within a disaggregated memory environment without losing the benefits of disaggregated memory.
[0008] In another exemplary embodiment, a computer-implemented method for mirroring memory in a segregated memory clustering environment is provided. The method includes a hypervisor allocating segregated memory to virtual machines in a cluster. The method further includes a segregated memory manager allocating the mirrored memory to a remote segregated memory allocation. The method further includes the segregated memory manager maintaining the mirrored memory when memory accesses occur. The method further includes the segregated memory manager adjusting the allocation to maintain mirroring in response to memory allocation adjustments. The method further includes the segregated memory manager determining whether to make an allocation change to improve performance and resilience in response to changes to the cluster and adjusting the allocation accordingly. The method further includes determining that a primary node of the segregated memory has failed. The method further includes switching to use a secondary node as the primary memory source and re-establishing an alternative secondary node for the mirrored memory in response to determining that the primary node has failed. The method improves the resilience and / or performance of the cluster and / or node. Additionally, the method improves cluster functionality by providing mirrored memory, for example, the method provides protection against catastrophic memory loss and improved reliability and availability within a disaggregated memory environment without losing the benefits of disaggregated memory.
[0009] In another exemplary embodiment, a system is provided that includes a memory having computer-readable instructions and a processing device for executing the computer-readable instructions. The computer-readable instructions control the processing device to perform operations for mirroring memory in a disjointed memory clustering environment. The operations include a hypervisor allocating disjointed memory to virtual machines in a cluster. The operations further include a disjointed memory manager allocating the mirrored memory to a remote disjointed memory allocation. The operations further include the disjointed memory manager maintaining the mirrored memory when memory accesses occur. The operations further include the disjointed memory manager adjusting the allocation to maintain mirroring in response to memory allocation adjustments. The operations further include the disjointed memory manager determining whether allocation changes should be made to improve performance and resilience in response to changes to the cluster and adjusting the allocation accordingly. The operations further include determining that a primary node of the disjointed memory has failed. The operations further include, in response to determining that the primary node has failed, switching to use the secondary node as the primary memory source and re-establishing an alternative secondary node for the mirrored memory. The system improves the resilience and / or performance of the cluster and / or nodes. Furthermore, the system improves the functionality of the cluster by providing mirrored memory. For example, the system provides protection against catastrophic memory loss and improved reliability and availability within a disaggregated memory environment without losing the benefits of disaggregated memory.
[0010] The above and other features and advantages of the present disclosure will become readily apparent from the following detailed description when taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0011] The particulars of the exclusive rights set forth herein are particularly pointed out and distinctly claimed in the claims at the conclusion of this specification. The above-discussed and other features and advantages of embodiments of the present invention will become apparent from the following detailed description when taken in conjunction with the accompanying drawings.
[0012] [Figure 1] 1 shows a block diagram of a processing system for implementing one or more embodiments described herein.
[0013] [Figure 2] 1 illustrates a hybrid cloud with nodes in a cluster, as shown in accordance with one or more embodiments described herein.
[0014] [Figure 3-1] FIG. 3A illustrates a cluster for memory servers in the hybrid cloud of FIG. 2 according to one or more embodiments described herein.
[0015] [Figure 3-2] FIG. 3B illustrates a memory for one of the nodes of the cluster of FIG. 3A, according to one or more embodiments described herein.
[0016] FIG. 3C illustrates one of the nodes of FIGS. 2 and 3A according to one or more embodiments described herein.
[0017] [Figure 4] 1 illustrates a flow diagram of a method for mirrored disjoint memory in a clustered environment, according to one or more embodiments described herein.
[0018] [Figure 5] 1 illustrates a flow diagram of a method for mirrored disjoint memory in a clustered environment, according to one or more embodiments described herein.
[0019] [Figure 6A]1 illustrates a flow diagram of a method for mirrored disjoint memory in a clustered environment, according to one or more embodiments described herein. [Figure 6B] 1 illustrates a flow diagram of a method for mirrored disjoint memory in a clustered environment, according to one or more embodiments described herein. [Figure 6C] 1 illustrates a flow diagram of a method for mirrored disjoint memory in a clustered environment, according to one or more embodiments described herein.
[0020] The diagrams shown herein are exemplary. There may be many variations on the diagrams or operations described therein without departing from the scope of the present invention. For example, actions may be performed in a different order, or actions may be added, deleted, or modified. Also, the term "coupled" and variations thereof describe having a communication path between two elements and do not imply a direct connection between the elements without an intervening element / connection between them. All of these variations are considered to be part of this specification. DETAILED DESCRIPTION OF THE INVENTION
[0021] One or more embodiments described herein provide mirrored, disjointed memory in a clustered environment.
[0022] Various aspects of the present disclosure are described by text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in the respective flowchart. For example, depending also on the technology involved, two operations shown in successive flowchart blocks may be performed in the reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.
[0023] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media"), collectively contained in one or more storage devices, that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media.As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device temporary because the data is not temporary while it is stored.
[0024] Computing environment 100 includes an example of an environment for execution of at least a portion of computer code involved in performing the methodology of the present invention, such as memory mirroring 150 in a disaggregated memory clustering environment. In addition to block 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes processor set 110 (including processing circuitry 120 and cache 121), communications fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 150, as identified above), peripheral device set 114 (including user interface (UI), device set 123, storage 124, and Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0025] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or multiple locations. However, in this presentation of computing environment 100, to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not shown in FIG. 1 within a cloud, it may be located within a cloud. However, computer 101 is not required to reside within a cloud except to any extent that may be expressly indicated.
[0026] Processor set 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple tailored integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computing.
[0027] Computer-readable program instructions are typically loaded onto computer 101 and cause processor set 110 of computer 101 to execute a series of operational steps, thereby enabling a computer-implemented method, such that the instructions so executed instantiate the methods specified in the computer-implemented method flowcharts and / or descriptions contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for executing the methods of the present invention may be stored in block 150 within persistent storage 113.
[0028] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive pathways, such as switches and conductive pathways that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication pathways may be used, such as fiber optic communication pathways and / or wireless communication pathways.
[0029] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, although this is not required unless expressly stated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0030] Persistent storage 113 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data is maintained regardless of whether power is supplied to computer 101 and / or to persistent storage 113 directly. Persistent storage 113 can be read-only memory (ROM), but typically at least a portion of persistent storage allows data to be written, data to be deleted, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 122 can take several forms, including various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems that employ a kernel. The code contained in block 150 typically includes at least a portion of the computer code involved in performing the methods of the present invention.
[0031] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as Bluetooth connections, Near-Field Communication (NFC) connections, connections made by cable (such as a universal serial bus (USB)-type cable), insertion-type connections (e.g., a secure digital (SD) card), connections made through a local area communication network, and even connections made through a wide area network such as the Internet. In various embodiments, the UI device set 123 can include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage 124 can be external storage, such as an external hard drive, or insertable storage, such as an SD card. The storage 124 can be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (e.g., computer 101 stores and manages large databases locally), then this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple, geographically distributed computers. IoT sensor set 125 is made up of sensors that may be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0032] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages multiple different network hardware devices. Computer-readable program instructions for implementing the methods of the present invention may be downloaded to computer 101 from an external computer or external storage device, typically through a network adapter card or network interface included in network module 115.
[0033] WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances using any technology for communicating computer data now known or later developed. In some embodiments, a WAN may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.
[0034] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and may take any of the forms described above with respect to computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to the end user, the recommendations would typically be communicated from computer 101's network module 115 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be a client device, such as a thin client, a heavy client, a mainframe computer, a desktop computer, and the like.
[0035] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on past data, then this past data may be provided to computer 101 from remote database 130 of remote server 104.
[0036] A public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer capabilities, particularly data storage (cloud storage) and computing power, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 105 computing resources is performed by computer hardware and / or software in cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers comprising host physical machine set 142, which is the universe of physical computers within and / or available in public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It is understood that these VCEs may be stored as images and may be transferred among and between various physical machine hosts, either as images or after instantiation of the VCEs. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.
[0037] Some further description of virtualized computing environments (VCEs) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running within them. A computer program running on a typical operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices allocated to the container; this feature is known as containerization.
[0038] Private cloud 106 is similar to public cloud 105, except that its computing resources are available only for use by a single enterprise. While private cloud 106 is shown in communication with WAN 102, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often each implemented by a different vendor. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both public cloud 105 and private cloud 106 are part of a larger hybrid cloud.
[0039] One or more embodiments described herein provide mirrored disjointed memory in a clustered environment. As described herein, disjointed memory is the ability to dynamically share memory across multiple processing systems via a dedicated network fabric or any other interface mechanism that can provide memory sharing. The use of disjointed memory in a clustered environment may reduce the reliability and availability of remote disjointed memory compared to local memory due to the additional cabling, connections, distance, and systems involved in such a configuration. If an operating system or virtual machine utilizes disjointed memory located on a remote system and that memory becomes unavailable, it may be as catastrophic to the workload as losing all of the memory due to the possibility that both local and remote memory are presented to virtual memory as contiguous blocks.
[0040] Several modern approaches attempt to address the challenges associated with disaggregated memory in clustered environments, but are inadequate. For example, one modern approach utilizes only local memory for allocation to virtual machines. However, this approach removes the benefits of utilizing disaggregated memory in a clustered environment. Another modern approach utilizes disaggregated memory only for specific memory volumes for which loss of access may be tolerable. However, this approach requires undesirable configuration and management overhead. Yet another modern approach accepts / tolerates the reduced reliability and availability brought about by the added complexity of a disaggregated memory environment. However, this approach is unacceptable in many environments.
[0041] One or more embodiments are provided to address these and other shortcomings by providing mirroring of remote isolated memory in an isolated memory environment to another node in a cluster accessible to the system hosting the virtual machine. According to one embodiment, the mirroring is achieved using a unique physical path, which may provide redundancy for failure events in the physical path. If a failure occurs in the cluster such that the virtual machine's isolated memory is no longer available to the virtual machine, the unique path and memory replica may be automatically utilized, allowing the failure in the cluster to be tolerated.
[0042] Turning now to FIG. 2 , a hybrid cloud 200 is shown having nodes 210-215 in a cluster in accordance with one or more embodiments described herein. In particular, hybrid cloud 200 includes nodes 210, 211, 212, 213, 214, and 215. Each of the nodes includes a memory and a storage device. For example, node 210 includes memory 220 and storage device 230, node 211 includes memory 221 and storage device 231, node 212 includes memory 222 and storage device 232, node 213 includes memory 223 and storage device 233, node 214 includes memory 224 and storage device 234, and node 215 includes memory 225 and storage device 235. It should be understood that each of nodes 210-215 may include multiple memories and / or multiple storage devices.
[0043] FIG. 3A illustrates a cluster 301 for memory servers in a hybrid cloud 200, according to one or more embodiments described herein. Similar to FIG. 2, each of the nodes 210-213 includes a memory and a storage device. Each of the memories 220-223 for the nodes 210-213 may include a local memory portion and a remote memory portion, as shown in FIG. 3B. As an example, FIG. 3B illustrates memory 220 for node 210 of the cluster 301, according to one or more embodiments described herein. The memory 220 includes a local memory portion 351 and a remote memory portion 352. The local memory portion 351 is local memory for the node 210 (e.g., “node 0”). The remote memory portion 352 is remote memory from the other nodes of the cluster 301, including nodes 211, 212, and 213 (e.g., “node 1,” “node 2,” and “node 3”). This provides disaggregated memory in a clustered environment.
[0044] Each of nodes 210-213 may be a requesting node that may request data stored in memory on one or more of the other nodes. For example, node 210 may be a requesting node and may request data stored in memory on one or more of nodes 211, 212, and 213. A requesting node may have an identifier "Cx:Ny:Mz," where "C" is the cluster, "N" is the node, and "M" is the node's memory. If node 210 is a requesting node, node 210 may be identified as "cluster 310:node 210:memory 222." The method used to locate a memory block may not be limited to the above forms of memory addressing of memory blocks or units, and may also include various methods of addressing memory blocks or units in a clustered environment.
[0045] 3C illustrates node 210 of FIGS. 2 and 3A in accordance with one or more embodiments described herein. As shown, node 210 includes memory 220 and storage device 230. Node 210 also includes hypervisor 360, virtual machine 361, and disaggregated memory manager (DMM) 362. Hypervisor 360 manages and allows one or more virtual machines (e.g., virtual machine 361) to run using multiple (and possibly different) operating systems on the same host machine (e.g., node 210). DMM 362 facilitates mirrored disaggregated memory in a clustered environment, as further described herein with reference to FIGS. 4, 5, and 6A-6C.
[0046] 4 shows a flow diagram of a method 400 for mirrored disjoint memory in a clustered environment according to one or more embodiments described herein. Method 400 may be implemented using any suitable system or device, such as computing environment 100, one or more of nodes 210-215, and / or the like, including combinations and / or multiples thereof. For example, method 400 may be implemented using node 210, which includes hypervisor 360, virtual machine 361, and DMM 362, as shown in FIG. 3C.
[0047] In block 402, hypervisor 360 allocates disjointed memory to virtual machine 361, which includes remote disjointed memory (e.g., remote memory portion 352 of memory 220). Virtual machine 361 is included within one node of a cluster (e.g., cluster 301) of the disjointed memory clustering environment. According to one or more embodiments described herein, virtual machine 361 further includes local memory (e.g., local memory portion 351 of memory 220).
[0048] In block 404, the DMM 362 allocates the mirrored memory to the remote isolated memory and mirrors the remote isolated memory on an alternative node (e.g., one of the nodes 211-213) of a cluster (e.g., the cluster 301) in the isolated memory clustering environment. The alternative node may be accessible through an independent path.
[0049] In block 406, DMM 362 maintains mirrored memory in response to memory accesses occurring, for example, when a memory access occurs, DMM 362 ensures that memory writes occur to both the primary and alternate blocks of memory on unique nodes of cluster 301.
[0050] At block 408, DMM 362 modifies memory utilization across the cluster in response to detecting a memory allocation adjustment. For example, when a memory allocation adjustment is made by hypervisor 360, DMM 362 may modify memory utilization of virtual machines 361 across cluster 301 to maintain a mirror of remote allocations of memory within cluster 301. According to one example, when configuration changes occur within cluster 301 (e.g., removing a node, adding a node, and / or the like, including combinations and / or multiples thereof), DMM 362 may respond to these changes by adjusting memory distribution / allocation. This improves the resilience and / or performance of cluster 301 and / or nodes 210.
[0051] In block 410, in response to detecting a failure that results in a loss of access to the remotely isolated memory, a corrective action may be implemented. The corrective action may be selected to improve processing system performance, such as to improve the resilience and / or performance of the cluster 301. For example, if a failure occurs that results in a loss of access to the remotely isolated memory, the DMM 362 may respond to the failure in the cluster 301, such as by selecting a new secondary node to provide mirroring, thereby improving the functionality of the processing system (e.g., cluster 301) by providing mirrored memory.
[0052] Additionally, the method 400 improves the performance of the nodes 210-213 of the cluster 301 by providing increased reliability and availability within a disaggregated memory environment without losing the benefits of having disaggregated memory.
[0053] It should be understood that additional processes may be included, and that the processes shown in FIG. 4 represent examples, and that other processes may be added, or existing processes may be deleted, modified, or rearranged, without departing from the scope of the present disclosure.
[0054] 5 shows a flow diagram of a method 500 for mirrored disjoint memory in a clustered environment according to one or more embodiments described herein. Method 500 may be implemented using any suitable system or device, such as computing environment 100, one or more of nodes 210-215, and / or the like, including combinations and / or multiples thereof. For example, method 500 may be implemented using node 210, including hypervisor 360, virtual machine 361, and DMM 362, as shown in FIG. 3C.
[0055] At block 502, hypervisor 360 allocates segregated memory to virtual machine 361. This may be done either as a user configuration decision or as a memory utilization optimization technique with segregated memory allocated from across the cluster of nodes by virtual machine 361. According to one embodiment, virtual machine 361 includes both local and remote segregated memory (see, e.g., FIG. 3B). This combination of local and remote segregated memory may be presented to the operating system as a single contiguous block of memory within the virtual address space.
[0056] In block 504, the DMM 362 allocates the mirrored memory to the remote unified memory allocation. For example, the DMM 362 also allocates memory and mirrors the remote unified memory on an alternative node (e.g., one or more of nodes 211-214) in the cluster 301 that is accessible through an independent physical path. The DMM may have knowledge of various characteristics of the cluster 301, such as the capacity of the nodes in the cluster (e.g., overall memory capacity, available free memory), the topology of the cluster (e.g., the physical layout of the nodes and the connections between the nodes), the connection status between the nodes (e.g., link speeds may vary depending on cabling or connectors, error or degradation conditions), the power usage of the cluster nodes (e.g., nodes 210-213) and the data center power usage target, the location and characteristics of other DMMs in the cluster (e.g., location, current memory usage), and / or the like, including combinations and / or multiples thereof. The DMM 362 may utilize this knowledge to optimize the selection of the mirrored remote unified memory. Examples of how DMM 362 may use characteristics of the cluster include: the capabilities of nodes 210-213 in cluster 301 may be used to select nodes with sufficient memory capacity; a customer's topology may be used to achieve the goal of having independent physical paths to memory nodes for redundancy or to minimize data transfer distance; connection status between nodes may be used to select error-free connections or connections trained with higher bandwidth; power usage of cluster nodes and data center power usage goals may be used to optimally select backup nodes and help achieve power usage goals by consolidating usage and allowing nodes to be powered off; the location and characteristics of other DMMs in the cluster may be used to optimally select backup nodes that may allow other DMMs to achieve their own goals; and / or the like, including combinations and / or multiples thereof.
[0057] At block 506, DMM 362 maintains mirrored memory when a memory access occurs. For example, when a memory access occurs, DMM 362 ensures that a memory write is made to both the primary block and the alternate block of memory on a unique node of cluster 301. Various techniques can be used to maintain memory consistency across the primary and alternate memory. One example is that a write to memory is considered complete only when both copies of the memory have completed the operation. However, other techniques, such as caching or journaling of changes, can be used to ensure coherency to a desired level of guarantee.
[0058] At block 508, the DMM 362 adjusts allocations to maintain mirroring when memory allocation adjustments are made by the hypervisor 360. For example, when memory allocation adjustments are made by the hypervisor 360, the DMM 362 can modify memory usage of virtual machines 361 across the cluster 301 to maintain mirrors of remote allocations of memory within the cluster 301. These adjustments can involve changes such as if a segment of memory is relocated from a local system to a disjointed primary node, the DMM 362 can allocate and mirror the memory to an alternative node. As another example, if additional remote memory allocations are made, either a new memory allocation will be established on an alternative remote node, or the entire remote memory can be allocated to a new remote node. As another example, characteristics of disjointed memory cluster usage can be used to optimize the use of disjointed memory within the cluster 301.
[0059] In block 510, when changes across the cluster 301 occur, the DMM 362 checks whether allocation / topology changes should be made to improve performance and / or resilience and adjusts the allocation / topology if desired. For example, when configuration changes occur within the cluster 301 (e.g., node removal, node addition), the DMM 362 may adjust memory distribution / allocation in response to these changes to improve resilience and / or performance. According to one embodiment, if a node is to be removed, the DMM 362 may select an alternative node to host memory currently used as a primary source or mirrored backup. Controlled node removal may allow new mirrors to be established before the node is removed to eliminate exposure when redundancy is not available. According to one embodiment, when a node is added to the cluster 301, the DMM 362 may reevaluate the network to determine whether the new node can be utilized to improve the resilience and / or performance of the memory topology. Redundancy may be improved if the addition of a node provides a unique physical path to the primary and backup mirrors of the segregated memory. Performance may improve if the addition of nodes allows for shorter paths (e.g., fewer hops, shorter distances) or the use of higher bandwidth links to the primary and backup mirrors of the split memory.
[0060] At block 512, it is determined whether a failure has occurred in the primary node of the disjointed memory (e.g., node 210). If it is determined that a failure has occurred in the primary node, method 500 proceeds to block 514. For example, if it is determined that a failure has occurred that results in a loss of access to the remote disjointed memory, DMM 362 may respond to the failure in cluster 301. If a failure has occurred in the primary node, at block 514, DMM 362 may switch to utilizing an alternate node disjointed memory to allow virtual machines to continue running without interruption from local memory while the mirror is re-established. This becomes the primary memory node, and at block 516, a new alternate mirrored memory allocation may be established. If it is determined at block 512 that a failure has not occurred in the primary node, method 500 proceeds to block 518, where it is determined whether a failure has occurred in a secondary node (e.g., one or more of nodes 211-213). For example, if the secondary node has failed, as determined at block 518, DMM 362 may find an alternative secondary node in cluster 301 as a mirror for the isolated memory at block 516, and / or DMM 362 may wait until the failed memory becomes accessible and then re-establish the mirror. Until the memory is recovered, the virtual machine may run without the memory redundancy of the remote isolated memory. If at block 518 it is determined that the secondary node has not failed, method 500 proceeds to block 506, where method 500 is at least partially repeated.
[0061] Additional processes may also be included, such as the following: For example, one or more embodiments may provide for remote mirroring of local memory allocation utilized by virtual machine 361, which may be managed by DMM 362 in a manner similar to remote memory allocation, as previously described. This may provide additional resilience by providing a mirror for local memory and utilizing logic in DMM 362 to dynamically select the optimal node to provide the mirrored data.
[0062] One or more embodiments may utilize multiple nodes in cluster 301 (e.g., multiple of nodes 211-213) to provide a full mirror for the partitioned memory in the primary node (e.g., node 210). This may be done if no single node has the capacity to provide a full mirror of the partitioned memory.
[0063] One or more embodiments may use knowledge of accesses to "hot" memory to automatically move memory allocations between local and remote memory within cluster 301. Keeping frequently accessed memory local may provide improved performance and may eliminate the need for mirroring.
[0064] According to one or more embodiments, as an alternative to directly mapping redundant copies of remotely isolated memory to alternate nodes as described herein, DMM 362 may utilize various memory management strategies that involve spreading redundant data across multiple nodes or alternative memory management techniques to provide data redundancy.
[0065] One or more embodiments may use alternative data protection algorithms, such as utilizing known RAID algorithms, to spread data across multiple nodes, or even establish multiple mirrors of the data.
[0066] It should be understood that the processes shown in FIG. 5 represent examples, and that other processes may be added, or existing processes may be deleted, modified, or rearranged, without departing from the scope of the present disclosure.
[0067] 6A-6C together illustrate a flow diagram of a method 600 for mirrored disjointed memory in a clustered environment, according to one or more embodiments described herein. According to one or more embodiments described herein, the method 600 provides for the DMM 362 to allocate the mirrored memory to a remote disjointed memory allocation. The method 600 may be implemented using any suitable system or device, such as the computing environment 100, one or more of the nodes 210-215, and / or the like, including combinations and / or multiples thereof. For example, the method 600 may be implemented using the node 210, which includes the hypervisor 360, the virtual machine 361, and the DMM 362, as shown in FIG. 3C.
[0068] In block 602, DMM 362 examines the available memory capacity of each node in cluster 301 (e.g., nodes 210-213) and identifies a set of nodes that have sufficient capacity to provide mirroring for the current remote memory allocations being handled.
[0069] At block 604, it is determined whether the DMM 362 can identify a set of candidate nodes having a target node set size that has the capacity to provide the memory necessary to mirror the current remote memory allocation. If so, the method 600 proceeds to block 606, where it is determined whether the DMM 362 can determine a set of candidate nodes where each node in the set has an independent path from the system hosting the virtual machine 361 to the set of candidate nodes, relative to the current allocation of mirrored remote isolated memory. If so, the method 600 proceeds to block 614 (see FIG. 6C ).
[0070] If either block 604 or 606 is "no," method 600 proceeds to block 608 (see FIG. 6B). Referring to FIG. 6B, at block 608, DMM 362 increases the target node set size by "1" for the set of nodes that have the necessary capacity to mirror the current remote memory allocation. At block 610, it is determined whether the target node set size is greater than the total number of remote nodes in cluster 301. If "yes," then at block 612, DMM 362 determines that it is unable to determine a sufficient set of mirrored memory for the remote disjointed memory allocation. However, if at block 610 it is determined that the target node set size is not greater than the total number of remote nodes in cluster 301, then method 600 proceeds to block 604 (see FIG. 6A).
[0071] 6C , for each remaining set of candidate nodes, the DMM 362 examines the link speed, link state, and number of hops for the links between the system hosting the virtual machine 361 and the set of candidate nodes that provide the mirror. At block 616, the DMM 362 ranks the set of candidate nodes based on link speed, link state, and number of hops, where the highest rank represents the optimal path considering each factor (e.g., minimizing the number of hops to reduce points of failure, maximizing path link speed for optimal bandwidth, minimizing the use of links that are in a degraded state or have high error rates, and / or the like, including combinations and / or multiples thereof). At block 618, for each remaining set of candidate nodes ranked in terms of link speed / state / hops, the DMM 362 examines the power consumption for the set of nodes in the candidate set and selects the set of candidate nodes with the highest link speed / state / hop rank and the lowest power consumption. In block 620, the DMM 362 outputs an optimal set of nodes in the cluster 301 to provide mirroring for the current remote memory allocation being handled.
[0072] It should be understood that additional processes may be included, and that the processes shown in Figures 6A-6C represent examples, and that other processes may be added, or existing processes may be deleted, modified, or rearranged, without departing from the scope of the present disclosure.
[0073] Exemplary embodiments of the present disclosure include or provide improvements to various technical features, technical effects, and / or techniques. Exemplary embodiments of the present disclosure provide mirrored, isolated memory in a clustered environment by allocating isolated memory to a virtual machine having the isolated memory and mirroring the remote isolated memory on an alternate node of a cluster in the isolated memory clustered environment, such as by using a unique physical path that provides redundancy against failure events in the physical path, and assigning mirrored memory to the remote isolated memory. Such embodiments further provide for maintaining the mirrored memory in response to an access to the memory occurring; modifying memory utilization across the cluster in response to detecting a memory allocation adjustment; and implementing corrective action in response to detecting a failure that results in loss of access to the remote isolated memory. These aspects of the present disclosure constitute technical features that provide the technical effect of allowing failures within a cluster to be tolerated using unique paths and memory duplication when a failure occurs within the cluster such that the virtual machine's isolated memory is no longer available to the virtual machine. This improves processing system functionality by providing protection against catastrophic memory loss without losing the benefits of having disaggregated memory available, and by improving reliability and availability within the disaggregated memory environment. As a result of these technical features and technical effects, a processing system, such as a node or cluster of nodes, according to exemplary embodiments of the present disclosure that uses the techniques for mirrored disaggregated memory in a clustered environment described herein represents an improvement over existing techniques for memory management in clustered environments. It should be understood that the above examples of technical features, technical effects, and improvements to techniques of exemplary embodiments of the present disclosure are merely illustrative and not comprehensive.
[0074] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions, that implements a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order other than that noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may possibly be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or that executes a combination of dedicated hardware and computer instructions.
[0075] While the descriptions of various embodiments of the present invention have been presented for illustrative purposes, they are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles, practical applications, or technical improvements of the embodiments over technologies found in the marketplace, or to enable others skilled in the art to understand the embodiments described herein.
[0076] In a preferred embodiment of the invention described herein, an exemplary computer-implemented method for mirroring memory in a disjointed memory clustering environment is provided, the computer-implemented method comprising: a hypervisor allocating disjointed memory to virtual machines in the cluster; a disjointed memory manager allocating the mirrored memory to a remote disjointed memory allocation; the disjointed memory manager maintaining the mirrored memory when memory accesses occur; the disjointed memory manager adjusting the allocation to maintain the mirroring in response to memory allocation adjustments; the disjointed memory manager determining whether allocation changes should be made to improve performance and resilience in response to changes to the cluster and adjusting the allocation accordingly; determining that a failure has occurred in a primary node of the disjointed memory; and, in response to determining that a failure has occurred in the primary node, switching to use a secondary node as the primary memory source and re-establishing an alternative secondary node for the mirrored memory. Preferably, the method further comprises determining that a failure has occurred in a secondary node of the isolated memory; and, in response to determining that a failure has occurred in the secondary node, re-establishing another alternative secondary node for the mirrored memory.
[0077] In a preferred embodiment of the invention described herein, a system is provided comprising: a memory having computer-readable instructions; and a processing device for executing the computer-readable instructions, the computer-readable instructions controlling the processing device to perform operations for mirroring memory in a disjointed memory clustered environment, the operations comprising: a hypervisor allocating disjointed memory to virtual machines in the cluster; a disjointed memory manager allocating the mirrored memory to a remote disjointed memory allocation; the disjointed memory manager maintaining the mirrored memory when memory accesses occur; the disjointed memory manager adjusting the allocation to maintain the mirroring in response to memory allocation adjustments; the disjointed memory manager determining whether to make allocation changes to improve performance and resilience in response to changes to the cluster and adjusting the allocation accordingly; determining that a failure has occurred in a primary node of the disjointed memory; and, in response to determining that a failure has occurred in the primary node, switching to use a secondary node as the primary memory source and re-establishing an alternative secondary node for the mirrored memory. Preferably, the operations further include determining that a failure has occurred in a secondary node of the split memory; and, in response to determining that a failure has occurred in the secondary node, re-establishing another alternative secondary node for the mirrored memory.
Claims
1. 1. A computer-implemented method for mirroring memory in a segregated memory clustering environment, comprising: a hypervisor allocating isolated memory to a virtual machine including remote isolated memory, the virtual machine being one node of a cluster in the isolated memory clustering environment; a segregated memory manager assigning mirrored memory to the remote segregated memory and mirroring the remote segregated memory on an alternate node of the cluster in the segregated memory clustering environment; the separate memory manager maintaining the mirrored memory in response to memory accesses occurring; modifying memory utilization across the cluster in response to detecting a memory allocation adjustment by the separate memory manager; and implementing corrective action in response to detecting a failure that results in a loss of access to the remote isolated memory.
1. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the virtual machine further comprises a local memory.
3. The computer-implemented method of claim 1 , wherein the alternate nodes are accessible through independent paths.
4. 2. The computer-implemented method of claim 1, wherein the failure is a loss of access to the remote isolated memory and the corrective action is for the isolated memory manager to respond to the failure in the cluster.
5. detecting a configuration change within the cluster; and the separate memory manager modifying memory utilization across the clusters. The computer-implemented method of claim 1 further comprising:
6. The computer-implemented method of claim 5 , wherein the configuration change is the addition of a node from the cluster.
7. The computer-implemented method of claim 5 , wherein the configuration change is the removal of a node from the cluster.
8. a memory having computer-readable instructions; and A processing device for executing the computer readable instructions, the computer readable instructions controlling the processing device to perform the method for memory mirroring in an isolated memory clustered environment according to any one of the preceding claims. A system comprising:
9. 8. A computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by a processor to cause the processor to perform the method for memory mirroring in a decoupled memory clustering environment of any one of claims 1 to 7.