Mirror decomposition memory in cluster environment
By mirroring memory in a decomposed memory cluster environment, mirroring remote decomposed memory to the cluster's standby node using a management program and decomposed memory manager, dynamically adjusting memory allocation and performing correction actions in the event of a failure, the memory reliability and availability problems in the cluster environment are solved, and memory management with high reliability and high availability is achieved.
Patent Information
- Application Number
- CN202380085793.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-12-04
- Publication Date
- 2025-07-22
AI Technical Summary
In clustered environments, prior art is difficult to effectively utilize decomposed memory while improving reliability and availability, and existing methods can lead to catastrophic memory loss or increase management overhead.
By mirroring memory in a decomposed memory cluster environment, mirroring remote decomposed memory to the cluster's backup nodes using a manager and decomposed memory manager, dynamically adjusting memory allocations, and performing correction actions when a failure occurs, providing redundancy and backup of memory.
Improves the resilience and performance of clusters and nodes, prevents catastrophic memory loss, and improves the reliability and availability of decomposed memory environments.
Smart Images

Figure CN120359498A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Embodiments described herein generally relate to computer processing systems, and more particularly, to computer processing systems that implement mirrored split memories in a clustered environment.
[0002] Cloud computing and cloud storage devices provide users with the ability to store and process their data in a third-party data center. Cloud computing helps to quickly and easily provide customers with the ability to run virtual machines without the customer having to purchase hardware or provide floor space for physical servers. Generally, virtual machines running as guests under the control of a hypervisor rely on the hypervisor to transparently provide virtualization services to the guest. These services include memory management, instruction emulation, and interrupt handling.
[0003] The term "hypervisor" refers to a processing environment or platform service that manages and allows one or more virtual machines to execute using multiple (sometimes different) operating systems on the same host. It should be understood that deploying a virtual machine includes the installation process of the virtual machine and the activation (or startup) process of the virtual machine. In another example, deploying a virtual machine includes the activation (or startup) process of the virtual machine (e.g., in the case where the virtual machine has been previously installed or already exists).
[0004] Some computer processing systems, such as nodes of a cloud computing system (or simply referred to as "processing systems"), include the ability to dynamically share memory across multiple processing systems on a dedicated network architecture. A processor chip in one processing system can be directly connected to a processor chip in another processing system, linking these processing systems together and enabling them to share each other's physical memory. This functionality is referred to as "memory starting", "memory splitting", or "memory clustering". This technology can be used in a variety of use cases, including within a composable data center, where resources (including memory) can be dynamically allocated and shared between systems. SUMMARY OF THE INVENTION
[0005] In an exemplary embodiment, there is provided an example computer-implemented method for mirroring memory in a disaggregated memory cluster environment. The method includes a hypervisor assigning disaggregated memory to a virtual machine including remote disaggregated memory, the virtual machine being a node of a cluster in the disaggregated memory cluster environment. The method further includes a disaggregated memory manager allocating memory for a mirror of the remote disaggregated memory to mirror the remote disaggregated memory to a standby node of the cluster in the disaggregated memory cluster environment. The method further includes, in response to a memory access occurring, the disaggregated memory manager maintaining the mirrored memory. The method further includes, in response to detecting a memory allocation adjustment, the disaggregated memory manager modifying the memory usage across the cluster. The method further includes, in response to detecting a failure that results in a loss of access to the remote disaggregated memory, performing a corrective action. The method improves the resiliency and / or performance of the cluster and / or nodes. Additionally, the method improves the functionality of the cluster by providing the mirrored memory. For example, the method provides for preventing catastrophic memory loss and for increasing reliability and availability in a disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory.
[0006] In another example embodiment, there is provided a system that includes a memory having computer-readable instructions and a processing device for executing the computer-readable instructions. The computer-readable instructions control the processing device to perform operations for mirroring memory in a disaggregated memory cluster environment. The operations include a hypervisor assigning disaggregated memory to a virtual machine including remote disaggregated memory, the virtual machine being a node of a cluster in the disaggregated memory cluster environment. The operations further include a disaggregated memory manager allocating memory for a mirror of the remote disaggregated memory to mirror the remote disaggregated memory to a standby node of the cluster in the disaggregated memory cluster environment. The operations further include, in response to a memory access occurring, the disaggregated memory manager maintaining the mirrored memory. The operations further include, in response to detecting a memory allocation adjustment, the disaggregated memory manager modifying the memory usage across the cluster. The operations further include, in response to detecting a failure that results in a loss of access to the remote disaggregated memory, performing a corrective action. The system improves the resiliency and / or performance of the cluster and / or nodes. Additionally, the system improves the functionality of the cluster by providing the mirrored memory. For example, the system is used to prevent catastrophic memory loss and to increase reliability and availability in a disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory.
[0007] In another exemplary embodiment, a computer program product is provided that includes a computer-readable storage medium having program instructions embodied therein that are executable by a processor to cause the processor to perform operations for mirroring memory in a disaggregated memory cluster environment. The operations include the hypervisor assigning disaggregated memory to a virtual machine that includes remote disaggregated memory, the virtual machine being a node of a cluster in the disaggregated memory cluster environment. The operations also include the disaggregated memory manager allocating memory for a mirror of the remote disaggregated memory to mirror the remote disaggregated memory to a standby node of the cluster in the disaggregated memory cluster environment. The operations also include the disaggregated memory manager maintaining the mirrored memory in response to a memory access occurring. The operations also include the disaggregated memory manager modifying memory usage across the cluster in response to detecting a memory allocation adjustment. The operations also include performing a corrective action in response to detecting a failure that results in a loss of access to the remote disaggregated memory. The computer program product improves the resiliency and / or performance of the cluster and / or nodes. Additionally, the computer program product improves the functionality of the cluster by providing mirrored memory. For example, the computer program product is used to prevent catastrophic memory loss and to improve reliability and availability in a disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory.
[0008] In another exemplary embodiment, a computer implementation for mirroring memory in a disaggregated memory cluster environment is provided. The method includes the hypervisor assigning disaggregated memory to a virtual machine of a cluster. The method also includes the disaggregated memory manager allocating memory for a mirror of the remote disaggregated memory. The method also includes the disaggregated memory manager maintaining the mirrored memory when a memory access occurs. The method also includes the disaggregated memory manager adjusting the allocation to maintain the mirror in response to a memory allocation adjustment. The method also includes the disaggregated memory manager determining whether an allocation change is to be made to improve performance and resiliency in response to a change to the cluster and adjusting the allocation accordingly. The method also includes determining that a failure has occurred in the primary node of the disaggregated memory. The method also includes, in response to determining that a failure has occurred in the primary node, switching to using the secondary node as the primary storage source and rebuilding the standby node for the mirrored memory. The method improves the resiliency and / or performance of the cluster and / or nodes. Additionally, the method improves the functionality of the cluster by providing the mirrored memory. For example, the method provides for preventing catastrophic memory loss and improving reliability and availability in a disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory.
[0009] In another exemplary embodiment, a system is provided that includes a memory having computer-readable instructions and a processing device for executing the computer-readable instructions. The computer-readable instructions control the processing device to perform operations for mirroring a memory in a disaggregated memory cluster environment. The operations include the hypervisor assigning disaggregated memory to a virtual machine of the cluster. The operations also include the disaggregated memory manager allocating memory for a mirror of the remote disaggregated memory. The operations also include the disaggregated memory manager maintaining the mirrored memory when a memory access occurs. The operations also include the disaggregated memory manager adjusting the allocation to maintain the mirror in response to a memory allocation adjustment. The operations also include the disaggregated memory manager determining whether an allocation change is to be made to improve performance and resiliency in response to a change to the cluster and adjusting the allocation accordingly. The operations also include determining that a failure has occurred in the primary node of the disaggregated memory. The operations also include, in response to determining that the failure has occurred in the primary node, switching to using a secondary node as the primary storage source and reconstructing a standby secondary node for the mirrored memory. The system improves the resiliency and / or performance of the cluster and / or nodes. Additionally, the system improves the functionality of the cluster by providing mirrored memory. For example, the system is used to prevent catastrophic memory loss and to improve reliability and availability in a disaggregated memory environment without sacrificing the advantages of being able to utilize disaggregated memory.
[0010] The foregoing features and advantages of the present disclosure, as well as other features and advantages, are apparent from the following detailed description in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The details of the exclusive rights described herein are particularly pointed out and distinctly claimed in the claims at the end of the specification. The foregoing and other features and advantages of the embodiments of the invention are apparent from the following detailed description in conjunction with the accompanying drawings, in which:
[0012] Figure 1 A block diagram depicting a processing system for implementing one or more embodiments described herein;
[0013] Figure 2 A hybrid cloud having nodes in a cluster is depicted in accordance with one or more embodiments described herein;
[0014] Figure 3A A cluster of memory servers in a hybrid cloud is depicted in accordance with one or more embodiments described herein Figure 2 ;
[0015] Figure 3B A memory of one of the nodes of a cluster is depicted in accordance with one or more embodiments described herein Figure 3A ;
[0016] Figure 3Cdepicts a node of Figure 2 Figure 2 and Figure 3A in accordance with one or more embodiments described herein;
[0017] Figure 4 depicts a flowchart of a method for mirroring disaggregated memory in a cluster environment in accordance with one or more embodiments described herein;
[0018] Figure 5 depicts a flowchart of a method for mirroring disaggregated memory in a cluster environment in accordance with one or more embodiments described herein; and
[0019] Figures 6A to 6C collectively depict a flowchart of a method for mirroring disaggregated memory in a cluster environment in accordance with one or more embodiments described herein.
[0020] The figures described herein are illustrative. Many variations to the figures or operations described therein may be made without departing from the scope of the present invention. For example, acts may be performed in a different order, or acts may be added, deleted, or modified. Additionally, the term "coupled" and its variants describe having a communication path between two elements and does not imply a direct connection between the elements, with no intermediate element / connection therebetween. All such variations are considered to be part of this specification. DETAILED DESCRIPTION
[0021] One or more embodiments described herein provide for mirroring disaggregated memory in a cluster environment.
[0022] Aspects of the present disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in each flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.
[0023] A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in one or more storage device sets, the storage device sets collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP statement. A “storage device” is any tangible device that can retain and store instructions for use by a computer processor. By way of non-limitation, computer-readable storage media can be electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding devices (such as punched cards or pits / lands formed in the major surface of a disc), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in the present disclosure, is not to be construed as storage in the form of a transitory signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, optical pulses passing through an optical fiber cable, electrical signals transmitted through a wire, and / or other transmission media. As will be understood by those skilled in the art, during normal operation of a storage device, for example during access, defragmentation, or garbage collection, data is typically moved at some occasional points in time, but this does not render the storage device transitory because the data is not transitory when it is stored.
[0024] Computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the methods of the present invention, such as mirroring memory in a disaggregated memory cluster environment 150. In addition to block 150, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a set of processors 110 (including processing circuitry 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 150, as described above), a set of peripherals 114 (including user interface (UI), set of devices 123, storage device 124, and Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes remote database 130. Public cloud 105 includes gateway 140, cloud orchestration module 141, set of host physical machines 142, set of virtual machines 143, and set of containers 144.
[0025] Computer 101 may take the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future that is capable of running a program, accessing a network, or querying a database such as remote database 130. As is well known in the computer technology field and depending on the technology, the performance of computer-implemented methods may be distributed among multiple computers and / or among multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 may be located in the cloud, although it is not shown in the cloud in Figure 1 On the other hand, computer 101 does not need to be in the cloud, except to any degree that is affirmatively indicated.
[0026] Processor set 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, such as multiple coordinated integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is a memory located within the (multiple) processor chip packages and is typically used for data or code that should be available for rapid access by the threads or cores running on processor set 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the caches in the cache of the processor set may be located "off-chip". In some computing environments, processor set 110 may be designed to work with qubits and perform quantum computing.
[0027] Computer-readable program instructions are generally loaded onto computer 101 to cause the processor set 110 of computer 101 to execute a series of operational steps to implement a computer-implemented method such that the instructions so executed will exemplify the methods specified in the flowcharts and / or narrative descriptions of the computer-implemented methods included in this document (collectively referred to as "the methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for executing the methods of the present invention may be stored in persistent storage device 113 in block 150.
[0028] The communication structure 111 is a signal conduction path that allows the various components of the computer 101 to communicate with each other. Typically, this structure is made up of switches and conductive paths, such as those that make up a bus, a bridge, a physical input / output port, and the like. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.
[0029] The volatile memory 112 is any type of volatile memory known now or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory is characterized by random access, but this is not required unless explicitly stated. In the computer 101, the volatile memory 112 is located in a single package and inside the computer 101, but alternatively or additionally, the volatile memory can be distributed across multiple packages and / or be located external to the computer 101.
[0030] The persistent storage device 113 is any form of non-volatile storage device for a computer known now or to be developed in the future. The non-volatility of this storage device means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 can be read-only memory (ROM), but typically at least part of the persistent storage device allows for the writing of data, the deletion of data, and the rewriting of data. Some common forms of persistent storage devices include magnetic disks and solid-state storage devices. The operating system 122 can take many forms, such as various known proprietary operating systems or open-source portable operating system interface-type operating systems that use a kernel. The code included in block 150 typically includes at least some of the computer code involved in performing the method of the present invention.
[0031] The peripheral device set 114 includes the peripheral device set of the computer 101. The data communication connection between the peripheral devices and other components of the computer 101 can be implemented in various ways, such as a Bluetooth connection, a near-field communication (NFC) connection, a connection by a cable (such as a universal serial bus (USB) type cable), a plug-in connection (e.g., a Secure Digital (SD) card), a connection through a local communication network, and even a connection through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, a speaker, a microphone, wearable devices (such as goggles and smartwatches), a keyboard, a mouse, a printer, a touchpad, a game controller, and a haptic device. The storage device 124 is an external storage device, such as an external hard disk drive, or a plug-in memory, such as an SD card. The storage device 124 can be persistent and / or volatile. In some embodiments, the storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where the computer 101 is required to have a large amount of storage (e.g., in the case where the computer 101 locally stores and manages a large database), the storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer and another sensor can be a motion detector.
[0032] The network module 115 is a collection of computer software, hardware, and firmware that allows the computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 115 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the forwarding function of the network module 115 are executed on physically separate devices, such that the control function manages several different network hardware devices. The computer-readable program instructions for performing the methods of the present invention can generally be downloaded to the computer 101 from an external computer or an external storage device through a network adapter card or a network interface included in the network module 115.
[0033] The WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances by any technology known now or developed in the future for transmitting computer data. In some embodiments, the WAN may be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area network such as a Wi-Fi network. The WAN and / or LAN typically includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0034] The end user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating the computer 101) and may take any form discussed above in connection with the computer 101. The EUD 103 typically receives helpful and useful data from the operation of the computer 101. For example, in the hypothetical case where the computer 101 is designed to provide recommendations to an end user, the recommendation will typically be transmitted from the network module 115 of the computer 101 through the WAN 102 to the EUD 103. Thus, the EUD 103 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 103 may be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, etc.
[0035] The remote server 104 is any computer system that provides at least some data and / or functionality to the computer 101. The remote server 104 may be controlled and used by the same entity operating the computer 101. The remote server 104 represents the (multiple) machines that collect and store useful and helpful data used by other computers such as the computer 101. For example, in the hypothetical case where the computer 101 is designed and programmed to provide recommendations based on historical data, the historical data may be provided to the computer 101 from the remote database 130 of the remote server 104.
[0036] A public cloud 105 is any computer system that can be used by multiple entities, which provides computer system resources and / or other computing capabilities, especially the on-demand availability of data storage (cloud storage) and computing power, without the direct active management of the user. Cloud computing typically utilizes resource sharing to achieve consistency and economies of scale. The direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers of the set of host physical machines 142, and the set of host physical machines 142 is the totality of physical computers in and / or available for the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the set of virtual machines 143 and / or containers from the set of containers 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the instantiation of the VCE. The cloud orchestration module 141 manages the transfer and storage of images, deploys new instances of the VCE, and manages the active instances of the VCE deployment. The gateway 140 is a collection of computer software, hardware, and firmware that allows the public cloud 105 to communicate via the WAN 102.
[0037] Some further explanations of the virtualized computing environment (VCE) will now be provided. The VCE can be stored as an "image". New active instances of the VCE can be instantiated from the image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature in which the kernel allows the existence of multiple isolated user space instances called containers. From the perspective of the programs running therein, these isolated user space instances typically appear as real computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and the devices assigned to the container, which is a feature called containerization.
[0038] The private cloud 106 is similar to the public cloud 105, except that the computing resources are only available for use by a single enterprise. Although the private cloud 106 is described as communicating with the WAN 102, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) that are typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is tied together by standardized or proprietary technologies that enable the organization, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of the larger hybrid cloud.
[0039] One or more embodiments described herein provide disaggregated memory of images in a cluster environment. As described herein, disaggregated memory is the ability to dynamically share memory across multiple processing systems through a dedicated network fabric or any other interface mechanism that can provide memory sharing. Due to the additional cables, connections, distances, and systems involved in such a configuration, using disaggregated memory in a cluster environment can reduce the reliability and availability of remote disaggregated memory compared to local memory. If the operating system or virtual machine is utilizing disaggregated memory on a remote system and that memory becomes unavailable, it can be as catastrophic for the workload as losing all memory due to the possibility of local and remote memory being presented to virtual memory as a contiguous block.
[0040] Some current methods have attempted to address the problems associated with disaggregated memory in a cluster environment, but these methods are insufficient. For example, one current method is to only utilize local memory to assign to virtual machines. However, this method eliminates the advantages of using disaggregated memory in a cluster environment. Another current method is to only use disaggregated memory for a specific capacity of memory for which the loss of access is tolerable. However, this method requires undesirable configuration and management overhead. Another current method is to accept / tolerate the reduced reliability and reduced availability introduced by the additional complexity of the disaggregated memory environment. However, this method is unacceptable in many environments.
[0041] One or more embodiments are provided to address these and other drawbacks by providing a mirror of remote disaggregated memory within a disaggregated memory environment to another node within the cluster that is accessible from the system hosting the virtual machine. According to an embodiment, a unique physical path can be used to implement the mirroring to provide redundancy for fault events in the physical path. If a fault occurs in the cluster such that the disaggregated memory of the virtual machine is no longer available to the virtual machine, the unique path and copy of the memory can be automatically utilized to allow tolerance of faults in the cluster.
[0042] Now turning to Figure 2 ,a hybrid cloud 200 with nodes 210 to 215 in a cluster is depicted according to one or more embodiments described herein. Specifically, the hybrid cloud 200 includes nodes 210, 211, 212, 213, 214, 215. Each of the nodes includes a memory and a storage device. For example, node 210 includes memory 220 and storage device 230, node 211 includes memory 221 and storage device 231, node 212 includes memory 222 and storage device 232, node 213 includes memory 223 and storage device 233, node 214 includes memory 224 and storage device 234, and node 215 includes memory 225 and storage device 235. It should be understood that each of the nodes 210 to 215 may include multiple memories and / or multiple storage devices.
[0043] Figure 3A A cluster 301 of memory servers in a hybrid cloud 200 is depicted according to one or more embodiments described herein. Similar to Figure 2 ,each of the nodes 210 to 213 includes a memory and a storage device. Each of the memories 2200 to 223 in the memories of the nodes 210 to 213 may include a local memory portion and a remote memory portion, as Figure 3B shown. As an example, Figure 3B depicts the memory 220 of node 210 of the cluster 301 according to one or more embodiments described herein. The memory 220 includes a local memory portion 351 and a remote memory portion 352. The local memory portion 351 is the local memory of node 210 (e.g., "node 0"). The remote memory portion 352 is the remote memory from other nodes (e.g., "node 1", "node 2", and "node 3") of the cluster 301 including nodes 211, 212, 213. This provides decomposed memory in a cluster environment.
[0044] Each of the nodes 210 to 213 may be a requesting node that may request data stored in the memories of one or more other nodes among other nodes. For example, node 210 may be a requesting node and may request data stored in the memories on one or more of the nodes 211, 212, 213. The requesting node may have an identifier "Cx:Ny:Mz", where "C" is the cluster, "N" is the node, and "M" is the memory of the node. In the case where node 210 is a requesting node, node 210 may be identified as "cluster 310:node 210:memory 222". The method for locating a memory block may not be limited to the above format of memory addressing of the memory block or cell, but may also include various ways of addressing a memory block or cell in a cluster environment.
[0045] Figure 3C Depicting a node 210 according to one or more embodiments described herein Figure 2 and Figure 3A as shown. As shown, node 210 includes a memory 220 and a storage device 230. Node 210 also includes a hypervisor 360, a virtual machine 361, and a disaggregated memory manager (DMM) 362. The hypervisor 360 manages and allows one or more virtual machines (e.g., virtual machine 361) to execute using multiple (and sometimes different) operating systems on the same host (e.g., node 210). The DMM 362 facilitates disaggregated memory mirroring in a cluster environment, as further described herein with reference to Figure 4 , Figure 5 and Figures 6A to 6C .
[0046] Figure 4 Depicts a flowchart of a method 400 for disaggregating mirrored memory in a cluster environment according to one or more embodiments described herein. The method 400 can be implemented using any suitable system or device, such as a computing environment 100, one or more of nodes 210 to 215, etc., including combinations and / or multiples thereof. For example, the method 400 can be implemented using node 210, which includes a hypervisor 360, a virtual machine 361, and a DMM 362, as Figure 3C shown.
[0047] At block 402, the hypervisor 360 assigns disaggregated memory including remote disaggregated memory (e.g., the remote memory portion 352 of memory 220) to the virtual machine 361. The virtual machine 361 is included within a node of a cluster (e.g., cluster 301) of a disaggregated memory cluster environment. According to one or more embodiments described herein, the virtual machine 361 also includes local memory (e.g., the local memory portion 351 of memory 220).
[0048] At block 404, the DMM 362 allocates memory for a mirror of the remote disaggregated memory to mirror the remote disaggregated memory onto a spare node (e.g., one of nodes 211 to 213) of a cluster (e.g., cluster 301) of a disaggregated memory cluster environment. The spare node can be accessed via an independent path.
[0049] At block 406, in response to a memory access occurring, the DMM 362 maintains the mirrored memory. For example, when a memory access is made, the DMM 362 ensures that memory writes are made to primary and spare memory blocks on a single node of cluster 301.
[0050] At block 408, in response to detecting a memory allocation adjustment, the DMM 362 modifies the memory usage across the clusters. For example, when the hypervisor 360 makes a memory allocation adjustment, the DMM 362 can modify the memory usage of the virtual machines 361 across the cluster 301 to maintain a mirror of the remote allocation of memory in the cluster 301. According to an example, when a configuration change occurs within the cluster 301 (e.g., removing a node, adding a node, etc., including combinations and / or multiples thereof), the DMM 362 can respond to these changes by adjusting the memory allocation / assignments. This improves the resiliency and / or performance of the cluster 301 and / or the node 210.
[0051] At block 410, in response to detecting a failure that results in a loss of access to the remote disaggregated memory, corrective actions can be implemented. The corrective actions are selected to improve the processing system performance, such as improving the resiliency and / or performance of the cluster 301. For example, in the case of a failure that results in a loss of access to the remote disaggregated memory, the DMM 362 can respond to the failure within the cluster 301, such as selecting a new secondary node to provide the mirror, thereby improving the functionality of the processing system (e.g., the cluster 301) by providing the mirrored memory.
[0052] Furthermore, the method 400 improves the performance of the nodes 210 to 213 of the cluster 301 by providing improved reliability and availability in a disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory.
[0053] Additional processes may also be included, and it should be understood that Figure 4 the processes depicted in
[0054] Figure 5 FIG. 500 is a flowchart of a method 500 for mirroring disaggregated memory in a cluster environment according to one or more embodiments described herein. The method 500 can be implemented using any suitable system or device, such as the computing environment 100, one or more of the nodes 210 to 215, etc., including combinations and / or multiples thereof. For example, the method 500 can be implemented using the node 210, which includes a hypervisor 360, a virtual machine 361, and a DMM 362, as Figure 3C shown.
[0055] At block 502, the hypervisor 360 assigns disaggregated memory to the virtual machine 361. This can be done as a configuration decision by the user or as a method for optimizing the memory usage with disaggregated memory assigned by the virtual machine 361 across a cluster of nodes. According to an embodiment, the virtual machine 361 includes local and remote disaggregated memory (see, for example Figure 3B)。This combination of local and remote disaggregated memory can be presented as a single contiguous block of memory within the virtual address space of the operating system.
[0056] At block 504, the DMM 362 allocates memory for the mirror of the remotely disaggregated memory allocation. For example, the DMM 362 also allocates memory to mirror the remotely disaggregated memory on a standby node (e.g., one or more of nodes 211 to 214) in the cluster 301 that is accessible via a physically independent path. The DMM can know various characteristics of the cluster 301, such as: the capacity of the nodes in the cluster (e.g., total storage capacity, available free memory), the topology of the cluster (e.g., the physical layout of the nodes and the connections between the nodes), the connection state between the nodes (e.g., the link speed can vary depending on the cable or connector, error or degraded state), the power usage of the cluster nodes (e.g., nodes 210 to 213) and the power usage target of the data center, the location and characteristics of other DMMs in the cluster (e.g., location, current memory usage), etc., including their combinations and / or multiplicities. The DMM 362 can utilize this knowledge to optimize the selection of the mirrored remotely aggregated memory. Examples of how the DMM 362 can use the characteristics of the cluster include: the capabilities of nodes 210 to 213 in the cluster 301 can be used to select a node with sufficient memory capacity, the topology of the client can be used to achieve the goal of having independent physical paths to the memory nodes for redundancy or to minimize data transfer distance, the connection state between the nodes can be used to select a connection that is error - free or trained on a higher bandwidth, the power usage of the cluster nodes and the power usage target of the data center can be used to optimally select a backup node to help achieve the power usage target by allowing nodes to power off through combined usage. The location and characteristics of other DMMs in the cluster can be used to optimally select a backup node that can allow other DMMs to achieve their own goals, and / or the like including their combinations and / or multiplicities.
[0057] At block 506, the DMM 362 maintains the mirrored memory when a memory access occurs. For example, when a memory access occurs, the DMM 362 ensures that memory writes are made to both the primary and standby memory blocks on the unique nodes of the cluster 301. Various techniques can be utilized to maintain memory consistency across the primary and standby memories. One example is to consider a memory write complete only when both copies of the memory have completed the operation. However, other techniques such as caching or change logging can be used to ensure the desired level of consistency.
[0058] At block 508, DMM 362 adjusts the allocation to maintain the mirror when the hypervisor 360 makes a memory allocation adjustment. For example, when the hypervisor 360 makes a memory allocation adjustment, DMM 362 can modify the use of the memory of virtual machine 361 on cluster 301 to maintain the mirror of the remote allocation of memory in cluster 301. These adjustments can involve changes such as if a memory segment is relocated from the local system to a disaggregated master node, DMM 362 can allocate and mirror the memory to a standby node. As another example, if additional remote memory allocation is made, a new memory allocation will be established on the standby remote node, or the entire remote memory allocation can be assigned to a new remote node. As another example, the characteristics of the disaggregated memory cluster usage can be used to optimize the use of disaggregated memory within cluster 301.
[0059] At block 510, when changes are made across cluster 301, DMM 362 checks whether allocation / topology changes should be made to improve performance and / or resiliency, and adjusts the allocation / topology if needed. For example, when configuration changes (e.g., remove node, add node) are made within cluster 301, DMM 362 can adjust the memory allocation / allocation in response to these changes to improve resiliency and / or performance. According to an embodiment, if a node is to be removed, DMM 362 can select a (plural) standby node to host the memory that is currently used as the primary source or mirror backup. Controlled node removal can allow a new mirror to be established before the node is removed to eliminate the exposure of having no redundancy available. According to an embodiment, if a node is added to cluster 301, DMM 362 can re-evaluate the network to determine whether the new node can be used to improve the resiliency and / or performance of the memory topology. If the addition of the node provides the only physical path to the primary and backup mirrors of the disaggregated memory, redundancy can be improved. If the addition of the node allows for a short path (e.g., fewer hops, short distance) or utilization of a higher bandwidth link to the primary and backup mirrors of the disaggregated memory, performance can be improved.
[0060] At block 512, it is determined whether a failure has occurred in the primary node of the split memory (e.g., node 210). If it is determined that a failure has occurred in the primary node, method 500 proceeds to block 514. For example, in the case where it is determined that a failure has occurred that results in a loss of access to the remote split memory, DMM 362 may respond to the failure within cluster 301. If the failure occurs in the primary node, then at block 514, DMM 362 may switch to using a spare node to split the memory, allowing the virtual machine to continue running without interruption to the local memory while rebuilding the image. This will become the primary storage node, and a new spare image memory allocation may be established at block 516. If at block 512 it is determined that no failure has occurred in the primary node, method 500 advances to block 518, where it is determined whether a failure has occurred in a secondary node (e.g., one or more of nodes 211 to 213). For example, if at block 518 it is determined that the failure has occurred in a secondary node, then at block 516, DMM 362 may find a spare secondary node in cluster 301 as a mirror for splitting the memory, and / or DMM 362 may wait for the failed memory to be accessible and then re-establish the mirror. The virtual machine may run without memory redundancy of the remote split memory before the memory is restored. If at block 518 it is determined that no failure has occurred in the secondary node, method 500 advances to block 506, and method 500 is at least partially repeated.
[0061] Additional processes may also be included, such as the following. For example, one or more embodiments may provide a remote mirror of the local memory allocation used by virtual machine 361, which may be managed by DMM 362 in a manner similar to the remote memory allocation described previously. This may provide additional resiliency by providing a mirror for the local memory and leveraging the logic of DMM 362 to dynamically select the best node for providing the mirrored data.
[0062] One or more embodiments may utilize multiple nodes in cluster 301 (e.g., multiple nodes 211 to 213) to provide a complete mirror for the split memory in the primary node (e.g., node 210). This may be done if there is no single node with the capacity to provide a full mirror of the split memory.
[0063] One or more embodiments may use knowledge of "hot" memory accesses to automatically move memory allocations between local and remote memory within cluster 301. Keeping frequently accessed memory local provides improved performance and may not require a mirror.
[0064] According to one or more embodiments, as an alternative to directly mapping redundant copies of remotely decomposed memory to standby nodes as described herein, DMM 362 may utilize various memory management strategies, including spreading redundant data across multiple nodes or alternative memory management methods to provide redundancy for the data.
[0065] One or more embodiments may use a standby data protection algorithm in such a way that data is spread across multiple nodes using a known RAID algorithm or even multiple mirrors of the data are established.
[0066] It should be understood that Figure 5 the processes depicted in are illustrative, and other processes may be added or existing processes may be removed, modified, or rearranged without departing from the scope of the present invention.
[0067] Figures 6A to 6C A flowchart of a method 600 for mirroring decomposed memory in a cluster environment according to one or more embodiments described herein is depicted together. According to one or more embodiments described herein, method 600 provides for DMM 362 to allocate memory for a mirror of a remotely decomposed memory allocation. Method 600 may be implemented using any suitable system or device, such as computing environment 100, one or more of nodes 210 to 215, etc., including combinations and / or multiples thereof. For example, method 600 may be implemented using node 210, which includes hypervisor 360, virtual machine 361, and DMM 362, as Figure 3C shown.
[0068] At block 602, DMM 362 checks the available memory capacity of each node (e.g., nodes 210 213) in cluster 301 and identifies a set of nodes having sufficient capacity to provide a mirror for the current remotely memory allocation being processed.
[0069] At block 604, it is determined whether DMM 362 can identify a candidate set of nodes having a target node set size that has the capacity to provide the memory required for mirroring the current remotely memory allocation. If so, method 600 proceeds to block 606, where it is determined whether DMM 362 can determine a candidate set of nodes for which each node in the set has an independent path from the system hosting virtual machine 361 to the candidate set of nodes other than the current allocation of the remotely decomposed memory being mirrored. If so, method 600 proceeds to block 614 (see Figure 6C ).
[0070] If either of blocks 604 or 606 is "no", method 600 proceeds to block 608 (see Figure 6B ). Refer to Figure 6B, at block 608, DMM 362 increases the target node set size by “1” for the set of nodes having the capacity required to mirror the current remote memory allocation. At block 610, it is determined whether the target node set size is greater than the total number of remote nodes in cluster 301. If so, at block 612, it is determined that DMM 362 cannot determine a sufficient set of mirrored memories for remote split memory allocation. However, if at block 610 it is determined that the target node set size is not greater than the total number of remote nodes in cluster 301, method 600 proceeds to block 604 (see Figure 6A ).
[0071] Refer to Figure 6C , at block 614, for each remaining candidate node set, DMM 362 checks the link speed, link status, and number of hops of the link between the system that is the host of virtual machine 361 and the candidate node set to provide mirroring. At block 616, DMM 362 ranks the candidate node sets based on the link speed, link status, and number of hops, where the highest rank represents the optimal path considering each factor (e.g., minimizing the number of hops to reduce failure points, maximizing the link speed of the path with optimal bandwidth, minimizing the use of links in a degraded state or having a high error rate, and / or the like, including their combinations and / or multiplicities). At block 618, for each remaining candidate node set in the sorted order of link speed / status / hops, DMM 362 checks the power consumption of the node sets in the candidate set and selects the candidate node set having the highest link speed / status / hops ranking and the lowest power consumption. At block 620, DMM 362 outputs the best node set in cluster 301 to provide mirroring for the current remote memory allocation being processed.
[0072] Additional processes may also be included, and it should be understood that Figures 6A to 6C the processes depicted in
[0073] Example embodiments of the present disclosure include or generate various technical features, technical effects, and / or improvements to technology. Example embodiments of the present disclosure provide mirrored disaggregated memory in a cluster environment by assigning disaggregated memory to a virtual machine having remote disaggregated memory and allocating mirrored memory for the remote disaggregated memory to mirror the remote disaggregated memory on a standby node of a cluster in a disaggregated memory cluster environment, for example, by using a unique physical path to provide redundancy for a fault event in the physical path. Such embodiments also provide for maintaining the mirrored memory in response to the occurrence of a memory access, modifying memory usage across the cluster in response to detecting a memory allocation adjustment, and implementing corrective actions in response to detecting a fault that results in a loss of access to the remote disaggregated memory. These aspects of the present disclosure constitute technical features that produce the technical effect of using a unique path and copy of memory to allow for tolerance of faults in the cluster when a fault occurs within the cluster such that the disaggregated memory of the virtual machine is no longer available to the virtual machine. This improves the processing system functionality by providing protection against catastrophic memory loss and by improving reliability and availability within the disaggregated memory environment without sacrificing the benefits of being able to utilize disaggregated memory. As a result of these technical features and technical effects, a processing system (such as a node or a cluster of nodes according to an example embodiment of the present disclosure) using the techniques described herein for mirrored disaggregated memory in a cluster environment represents an improvement over the prior art for memory management in a cluster environment. It should be understood that the above examples of the technical features, technical effects, and improvements to technology of the example embodiments of the present disclosure are merely illustrative and not exhaustive.
[0074] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a system based on dedicated hardware that performs the specified functions or acts, or combinations of dedicated hardware and computer instructions.
[0075] Descriptions of various embodiments of the present invention have been given for illustrative purposes, but these descriptions are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technical improvements to the technologies found in the market, or to enable those skilled in the art to understand the embodiments described herein.
[0076] In a preferred embodiment of the present invention described herein, a computer-implemented method for mirroring memory in a disaggregated memory cluster environment is provided. The computer-implemented method includes: assigning disaggregated memory to a virtual machine of a cluster by a hypervisor; allocating memory for a mirror of a remote disaggregated memory allocation by a disaggregated memory manager; maintaining the mirrored memory by the disaggregated memory manager as memory access occurs; adjusting the allocation by the disaggregated memory manager to maintain the mirror in response to a memory allocation adjustment; determining by the disaggregated memory manager whether an allocation change is to be made to improve performance and resiliency in response to a change in the cluster, and adjusting the allocation accordingly; determining that a failure has occurred in the primary node of the disaggregated memory; and in response to determining that the failure has occurred in the primary node, switching to use a secondary node as the primary storage source and rebuilding a standby secondary node for the memory of the mirror. Preferably, the method further includes: determining that a failure has occurred in the secondary node of the disaggregated memory; and in response to determining that the failure has occurred in the secondary node, rebuilding another standby secondary node for the memory of the mirror.
[0077] In a preferred embodiment of the present invention herein, a system is provided that includes: a memory including computer-readable instructions; and a processing device for executing the computer-readable instructions, the computer-readable instructions controlling the processing device to perform operations for mirroring memory in a disaggregated memory cluster environment. The operations include: assigning disaggregated memory to a virtual machine of a cluster by a hypervisor; allocating memory for a mirror of a remote disaggregated memory allocation by a disaggregated memory manager; maintaining the mirrored memory by the disaggregated memory manager as memory access occurs; adjusting the allocation by the disaggregated memory manager to maintain the mirror in response to a memory allocation adjustment; determining by the disaggregated memory manager whether an allocation change is to be made to improve performance and resiliency in response to a change in the cluster, and adjusting the allocation accordingly; determining that a failure has occurred in the primary node of the disaggregated memory; and in response to determining that the failure has occurred in the primary node, switching to use a secondary node as the primary storage source and rebuilding a standby secondary node for the memory of the mirror. Preferably, the operations further include: determining that a failure has occurred in the secondary node of the disaggregated memory; and in response to determining that the failure has occurred in the secondary node, rebuilding another standby secondary node for the memory of the mirror.
Claims
1. A computer-implemented method for mirroring memory in a disaggregated memory cluster environment, the computer-implemented method comprising: assigning, by a hypervisor, disaggregated memory to a virtual machine including remote disaggregated memory, the virtual machine being a node of a cluster of the disaggregated memory cluster environment; allocating, by a disaggregated memory manager, mirrored memory for the remote disaggregated memory to mirror the remote disaggregated memory onto a spare node of the cluster of the disaggregated memory cluster environment; maintaining, by the disaggregated memory manager, the mirrored memory in response to a memory access occurring; modifying, by the disaggregated memory manager, memory usage across the cluster in response to detecting a memory allocation adjustment; and performing a corrective action in response to detecting a failure that results in a loss of access to the remote disaggregated memory.
2. The computer-implemented method according to claim 1, wherein the virtual machine further includes local memory.
3. The computer-implemented method according to claim 1, wherein the spare node is accessible via an independent path.
4. The computer-implemented method according to claim 1, wherein the failure is a loss of access to the remote disaggregated memory, and wherein the corrective action is the disaggregated memory manager responding to the failure within the cluster.
5. The computer-implemented method according to claim 1, further comprising: detecting a configuration change within the cluster; and modifying, by the disaggregated memory manager, memory usage across the cluster.
6. The computer-implemented method according to claim 5, wherein the configuration change is adding a node to the cluster.
7. The computer-implemented method according to claim 5, wherein the configuration change is removing a node from the cluster.
8. A system, comprising: memory including computer-readable instructions; and a processing device for executing the computer-readable instructions, the computer-readable instructions controlling the processing device to execute the method for mirroring memory in a disaggregated memory cluster environment according to any one of the preceding claims.
9. A computer program product, comprising a computer-readable storage medium having program instructions included therein, the program instructions being executable by a processor to cause the processor to execute the method for mirroring memory in a disaggregated memory cluster environment according to any one of claims 1 to 7.