Prevent memory errors with proactive memory poisoning recovery
By using a scanner to detect and isolate uncorrectable memory errors in a cloud computing environment, generating MCE and performing virtual machine migration, the crash problem caused by uncorrectable memory errors in cloud computing is solved, and active error containment and graceful recovery are achieved.
Patent Information
- Application Number
- CN202280067853.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-15
- Filing Date
- 2022-10-28
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-10-28
AI Technical Summary
In cloud computing environments, uncorrectable memory errors can cause host crashes, impacting a large number of virtual machines and applications, leading to downtime and data loss. Existing technologies make it difficult to proactively detect and contain the spread of these errors before they occur.
By implementing a scanner on the host to actively scan the memory, uncorrectable memory errors are detected and a machine check exception (MCE) is generated. The kernel handler sends a signal to the hypervisor to isolate the poisoned memory page, realize the migration and failover of the virtual machine, and avoid error propagation.
Effectively detect and isolate uncorrectable storage errors, reduce downtime, maintain data integrity and stability of the virtual machine environment, avoid the impact of errors on other parts, and achieve graceful failover and recovery.
Smart Images

Figure CN118076946B_ABST
Abstract
Description
[0001] Cross-reference to related applications:
[0002] This application is a continuation of U.S. patent application No. 17 / 695,406, filed on March 15, 2022, entitled “Preventing Memory Errors Through Active Memory Poison Recovery,” the disclosure of which is incorporated herein by reference. Background Art
[0003] Cloud computing has impacted the way businesses manage their computing needs by cost-effectively providing reliable, flexible, scalable, and redundant computing resources. For example, cloud computing enables businesses to manage their information technology needs without the traditional capital investment and maintenance issues associated with managing and maintaining computer equipment. Furthermore, as more computing moves to cloud systems, the ability of these cloud systems to store, process, and output data has increased to levels that were once unimaginable.
[0004] The impact of this shift to cloud systems is that memory errors occurring in cloud systems, if not contained and / or recovered, can have an impact on customer and user experience comparable to the enterprise's footprint in the cloud. For example, it's not uncommon for the detection of an uncorrectable memory error on a host to cause the host to shut down, resulting in the abrupt termination of all virtual machines (VMs) and applications hosted by the host. Because cloud systems have memory sizes in the gigabyte or terabyte range, such memory errors can affect a large number of VMs and applications, leading to significant downtime, data loss, and customer disapproval.
[0005] When physical memory experiences a memory failure (e.g., an "uncorrectable error"), there are often other undetected memory errors, and the memory is likely "permanently" damaged. In this case, migrating the VM while preserving its running state can reduce downtime while controlling the number and severity of memory error propagation. Summary of the Invention
[0006] Aspects of the disclosed technology may include a method or system implemented in a cloud computing environment that allows for proactive detection, containment (e.g., preventing corrupted data from propagating to a target host during migration), and recovery from uncorrectable memory errors.
[0007] One aspect of the present disclosure relates to a method for proactively detecting memory errors in a cloud computing environment. The method may include scanning a memory of a host for errors by a scanner of the host; detecting a memory error in the memory of the host by the scanner; generating a machine check exception (MCE) by one or more processors of the host; and providing the MCE to a kernel executing on the host by the one or more processors.
[0008] In some cases, the scanning is performed continuously by the scanner. In some examples, the scanning is a read-only scan. In some examples, the memory error is an uncorrectable memory error. In some examples, the MCE includes an indication of a location in the memory where the memory error was detected by the scanner.
[0009] In some examples, the method further includes identifying one or more memory pages determined to be associated with the memory error as one or more poisoned memory pages based on a location of the memory at which a scanner detected the memory error.
[0010] In some examples, the method further includes isolating access by the host to the one or more poisoned memory pages.
[0011] In some examples, the method further includes receiving a page fault associated with a read request made by a guest of a virtual machine executing on the host; and sending, by the kernel, a SIGBUS signal to a hypervisor of the virtual machine.
[0012] In some examples, the method further includes generating, by the hypervisor, a machine check exception; and sending the machine check exception to the guest.
[0013] Another aspect of the technology relates to a system. The system may include a host computer capable of supporting one or more virtual machines; and one or more processing devices coupled to a memory containing instructions. The instructions may cause the one or more processing devices to: scan a memory of the host computer for errors; detect a memory error in the memory of the host computer; generate a machine check exception (MCE); and send the MCE to a kernel of the host computer, the MCE including information associated with the memory error.
[0014] In some cases, the scanning is performed continuously by the scanner. In some examples, the scanning is a read-only scan. In some examples, the memory error is an uncorrectable memory error. In some examples, the MCE includes an indication of a location in the memory where the memory error was detected by the scanner.
[0015] In some examples, the instructions further cause the one or more processors to: identify one or more memory pages determined to be associated with the memory error as one or more poisoned memory pages based on a location of the memory where the memory error was detected.
[0016] In some examples, the instructions further cause the one or more processors to isolate access by the host to the one or more poisoned memory pages.
[0017] In some examples, the instructions further cause the one or more processors to receive a page fault associated with a read request made by a guest of a virtual machine executing on the host; and send a SIGBUS signal to a hypervisor of the virtual machine.
[0018] In some examples, the instructions further cause the one or more processors to generate a machine check exception; and send the machine check exception to the client.
[0019] Another aspect of the present disclosure relates to a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: scan a memory of the host for errors; detect a memory error in the memory of the host; generate a machine check exception (MCE); and send the MCE to a kernel of the host, the MCE including information associated with the memory error. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A block diagram of an example system or environment is illustratively depicted in accordance with aspects of the disclosed technology.
[0021] Figure 2 A block diagram of an example system or environment is illustratively depicted in accordance with aspects of the disclosed technology.
[0022] Figure 3A Flowcharts or swim-lane diagrams illustratively depict example processes or methods according to aspects of the disclosed technology.
[0023] Figure 3B Flowcharts or swim-lane diagrams illustratively depict example processes or methods according to aspects of the disclosed technology.
[0024] Figure 3C Flowcharts or swim-lane diagrams illustratively depict example processes or methods according to aspects of the disclosed technology.
[0025] Figure 4 Flowcharts depicting example processes or methods in accordance with aspects of the disclosed technology.
[0026] Figures 5A to 5D Aspects of example processes or methods and sub-processes or sub-methods according to aspects of the disclosed technology are shown. DETAILED DESCRIPTION
[0027] Overview
[0028] Memory errors are generally categorized as correctable and uncorrectable. Correctable errors typically do not affect the normal operation of hosts in a cloud environment, and thus the normal operation of the host computing system. Uncorrectable errors are often fatal to the entire host computing system, causing the host to crash or shut down, for example. In a cloud-based virtual machine environment, this means that all virtual machines (VMs) supported by the host will crash or shut down along with the host, with little or no clue as to the cause of the crash and virtually no chance of recovery for the VMs / users. In modern cloud computing systems, the impact of uncorrectable memory errors is often significant. Cloud computing systems typically utilize relatively large amounts of memory on each host. For example, a cloud computing engine may support a single VM with 12 terabytes or more of memory. These larger hosts typically experience higher rates of uncorrectable memory errors than smaller hosts. The larger the memory capacity, the greater the chance of memory errors. Downtime caused by memory errors is often costly, especially for large hosts.
[0029] The presence of uncorrectable errors creates additional complexity in managing the expected behavior of VMs because these uncorrectable errors often indicate additional corruption of the underlying physical memory, which may contain additional hidden or unknown errors. Furthermore, correctable errors may become uncorrectable due to the presence of underlying hardware, which may degrade over time. If left unchecked, as the number of uncorrectable errors continues to increase, the physical host running one or more VMs is likely to experience a severe crash, bringing down all VMs on the host. Therefore, mitigation techniques (such as migrating virtual machines running on damaged hardware to known "good" machines) can limit the impact of detected uncorrectable errors and the downstream effects of such errors. However, other factors must be considered during the migration of "live" machines, including the occurrence of additional memory errors that were not accounted for at the beginning of the migration process.
[0030] Typical mitigation techniques only occur after a host or VM running on it experiences a memory error. The techniques described herein rely on a scanner to proactively scan system memory to detect errors before the host or VM encounters them. By doing so, the mitigation techniques described herein can be implemented before the host or VM uses the "bad" memory containing the error. This prevents host and VM corruption and other issues.
[0031] Aspects of the disclosed technology include "live" migration of a running VM from one physical host to another. In some examples, the migration can occur in a series of steps, including migrating memory pages in order of criticality. In some examples, the most relevant or critical portions of the memory can be migrated. In some examples, simulation of memory errors can be performed to exclude certain memory segments or memory pages that are determined to be "poisoned." For example, poisoned memory pages can include memory pages having virtual memory locations corresponding to damaged memory elements on the host, such as physical memory locations with flipped bits or damaged memory components. Aspects of the disclosed technology allow for the preservation of certain types of memory errors (including migration of these errors) to enable a consistent view to the end user after a live migration event. In addition, the detection, identification, and handling of memory errors in a virtual environment (e.g., at the hypervisor abstraction level) can be used to improve live migration (e.g., tracking and isolating poisoned pages) so that they are not copied and transmitted to the target host as part of the natural live migration process. Other aspects may include notifying the target host of the poisoned page or corrupted memory location so that computations (eg, checksum computations) at the target host do not include the poisoned page or corrupted memory location.
[0032] Aspects of the disclosed technology include migration of one or more VMs. In some examples, virtual machines can be migrated in an order based on importance, current usage, or the number of critical errors associated with a particular VM. In some examples, upon detection of one or more uncorrectable memory errors, all VMs running on a particular host containing the one or more uncorrectable memory errors can be migrated to healthy physical hosts.
[0033] Aspects of the disclosed technology include abstracting from specific underlying microarchitecture platforms and generalizing the architecture to allow for a "universal" abstraction of virtual machines across multiple host platforms or architectures.
[0034] By migrating a VM from one host to another, aspects of the disclosed technology can contain certain types of memory errors to maintain data integrity, stability, scalability, and robustness of the virtual machine environment.
[0035] One aspect of the disclosed technology includes a cloud computing infrastructure that allows scanners to proactively detect memory errors (including uncorrectable memory errors), as well as locate and contain memory errors so that they do not impact other parts of the system, such as client VM workloads. For example, the disclosed technology includes configuring a host scanner (including associated memory elements) to enable error signaling to be recovered at the operating system (OS), enhancing and enabling the OS's recovery path when a memory error on a memory page is detected. One example of the disclosed technology includes a central processing unit (CPU) capability that can signal contextual information associated with a memory error to the operating system (OS) (e.g., address, severity, whether the error is signaled in isolation so that it is recoverable, etc.). For example, such a mechanism can include Intel's x86 Machine Check architecture, in which the CPU reports hardware errors to the OS. A Machine Check Exception (MCE) handler in the OS kernel (e.g., provided via Linux) can then use an application programming interface (API) (e.g., POSIX) to signal the presence of an error to the virtual machine manager, along with contextual information about the error (e.g., location, error type, whether it is unrecoverable, status of neighboring memory locations, etc.). The virtual machine manager can then use this error message as part of initiating the live migration process.
[0036] For example, one aspect of the disclosed technology includes a cloud computing system or architecture in which a mechanism is provided for a virtual machine manager or hypervisor to include the ability to be alerted by a host of memory errors, particularly uncorrectable memory errors. Upon being alerted, the hypervisor processes the memory error information it receives from the host to identify VMs that may be accessing (or may eventually access) a damaged memory element, which can be identified from the memory error information included in the alert. Upon identifying the affected VMs, the hypervisor can initiate a process to failover the VMs running on the affected host so that the host can eventually be repaired.
[0037] It will be appreciated that a cloud computing system or architecture implemented according to the aforementioned mechanisms can contain and allow for graceful recovery from uncorrectable memory errors. Specifically, by identifying the affected memory, the hypervisor can limit or eliminate the intended use (e.g., reading or access) of such memory. Furthermore, the hypervisor can limit the impact to only the affected VMs. Furthermore, the hypervisor can appropriately initiate a failover of the affected VMs and then manage the movement of unaffected VMs supported by the damaged host to another host to allow the damaged host to be repaired. In this way, the impact of a customer or user being exposed to an uncorrectable memory error may be limited to the affected VM, whose virtual memory is linked to the damaged physical memory element or address, while unrelated VMs are unaware of the error and are not affected by it.
[0038] Example System
[0039] Figure 1 is an example system 100 according to aspects of the present disclosure. The system 100 includes one or more computing devices 110 (which may include computing devices 1101 to 110 k ), network 140 and one or more cloud computing systems 150 (which may include cloud computing systems 1501 to 150 m ). Computing devices 110 may include computing devices located at customer locations that utilize cloud computing services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and / or Software as a Service (SaaS). For example, if computing device 110 is located at a commercial enterprise, computing device 110 may use cloud system 150 as a service that provides software applications (e.g., accounting, word processing, inventory tracking, etc.) to computing device 110 for use in operating enterprise systems. As an alternative example, computing device 110 may rent infrastructure in the form of virtual machines on which to run software applications to support enterprise operations.
[0040] like Figure 1 As shown, each computing device 110 may include one or more processors 112, memory 116 for storing data (D) and instructions (I), a display 120, a communication interface 124, and an input system 128, which are shown interconnected via a network 130. Computing devices 110 may also be coupled or connected to storage 136, which may include local storage (e.g., on a storage area network (SAN)) or remote storage, which stores data accumulated as part of the client's operations. Computing devices 110 may include standalone computers (e.g., desktop or laptop computers) or servers associated with a client. A given client may also implement multiple computing devices as servers as part of its business. In the case of standalone computers, network 130 may include a data bus within the computer, etc.; in the case of servers, network 130 may include one or more local area networks, virtual private networks, wide area networks, or other types of networks described below in conjunction with network 140. Memory 116 stores information accessible by one or more processors 112, including instructions 132 and data 134 that may be executed or used by processors 112. The memory 116 can be of any type capable of storing information accessible by the processor, including computing device-readable media or other media that stores data that can be read with the aid of an electronic device, such as a hard drive, memory card, ROM, RAM, DVD or other optical disk, and other writable and read-only memory. The systems and methods may include different combinations of the foregoing, whereby different portions of instructions and data are stored on different types of media.
[0041] Instructions 132 can be any set of instructions that are directly executed by a processor (e.g., machine code) or indirectly executed (e.g., a script). For example, the instructions can be stored as computing device code on a computing device readable medium. In this regard, the terms "instructions" and "program" are used interchangeably herein. The instructions can be stored in an object code format for direct processing by the processor, or in any other computing device language, including scripts or collections of independent source code modules that are interpreted on demand or pre-compiled. The processes, functions, methods, and routines of the instructions are explained in more detail below.
[0042] The processor 112 may retrieve, store, or modify data 134 according to the instructions 132. As an example, the data 134 associated with the memory 116 may include data used to support services for one or more client devices, applications, etc. Such data may include data supporting hosted network-based applications, file sharing services, communication services, games, shared video or audio files, or any other network-based service.
[0043] The one or more processors 112 may be any conventional processor, such as a commercially available CPU. Alternatively, the one or more processors may be a dedicated device, such as an ASIC or other hardware-based processor. Figure 1 The processor, memory, and other elements of the computing device 110 are functionally shown as being in a single block, but one of ordinary skill in the art will appreciate that the processor, computing device, or memory may actually include multiple processors, computing devices, or memories, which may or may not be located or stored in the same physical enclosure. In one example, the one or more computing devices 110 may include one or more server computing devices having multiple computing devices, such as a load-balanced server farm, which exchanges information with different nodes of a network to receive, process, and send data to and from other computing devices as part of a client's business operations.
[0044] The computing device 110 may also include a display 120 (e.g., a monitor having a screen, touch screen, projector, television, or other device operable to display information) that provides a user interface that allows control of the computing device 110 and access to one or more VMs associated with user-space applications and / or data supported in the cloud system 150, such as on a host in the cloud system 150. Such control may include, for example, using the computing device to cause data to be uploaded to the cloud system 150 via an input system 128 for processing, to cause data to be accumulated on a storage device 136, or more generally, to manage various aspects of the customer's computing system. In some examples, the computing device 110 may also have access to an API that allows the computing device 110 to specify workloads or jobs to be run on VMs in the cloud as part of IaaS or SaaS. While the input system 128 may be used to upload data, such as a USB port, the computing device 110 may also include a mouse, keyboard, touch screen, or microphone that may be used to receive commands and / or data.
[0045] The network 140 may include various configurations and protocols, including short-range communication protocols such as Bluetooth TM 、Bluetooth TM The communication may be carried out over a network 140, such as a wireless network, a LAN, a LTE network, the Internet, the World Wide Web, an intranet, a virtual private network, a wide area network, a local area network, a private network using one or more company-specific communication protocols, Ethernet, WiFi, HTTP, and the like, and various combinations thereof. Such communication may be facilitated by any device capable of transmitting data to and from other computing devices, such as modems and wireless interfaces. The computing device is connected to the network 140 via a communication interface 124, which may include the necessary hardware, drivers, and software to support a given communication protocol.
[0046] The cloud computing system 150 may include one or more data centers that may be linked via a high-speed communications or computing network. A given data center within the system 150 may include a dedicated space within a building that houses a computing system and its associated components (e.g., a storage system and a communications system). Typically, a data center will include racks of communications equipment, servers / hosts, and disks. The servers / hosts and disks comprise the physical computing resources used to provide virtual computing resources (e.g., VMs). To the extent that a given cloud computing system includes more than one data center, these data centers may be located in different geographic locations relatively close to each other, selected to provide services in a timely and cost-effective manner, as well as to provide redundancy and maintain high availability. Similarly, different cloud computing systems are typically provided in different geographic locations.
[0047] like Figure 1As shown, computing system 150 can be shown to include host 152, storage 154, and infrastructure 160. Host 152, storage 154, and infrastructure 160 can comprise a data center within cloud computing system 150. Infrastructure 160 can include one or more hosts as well as switches, physical links (e.g., optical fibers), and other devices for interconnecting hosts within the data center with storage 154. Storage 154 can include disks or other storage devices that can be partitioned to provide physical or virtual storage to virtual machines running on processing devices within the data center. Storage 154 can be provided as a SAN within the data center that hosts the virtual machines supported by storage 154, or provided in a different data center that does not share a physical location with the virtual machines it supports. One or more hosts or other computer systems within a given data center can be configured to act as a supervisory agent or hypervisor when creating and managing virtual machines associated with one or more hosts within the given data center. Generally speaking, a host or computer system configured to function as a hypervisor will contain, for example, the instructions necessary to manage operations resulting from requests for services originating from, for example, computing device 110 to provide IaaS, PaaS, or SaaS to customers or users.
[0048] exist Figure 2 In the example shown, the distributed system 200 (eg, Figure 1 The cloud system 150 of the present invention (e.g., a system related to the cloud system 150) includes a collection 204 of hosts 210 (e.g., hardware resources 210) that support or execute a virtual computing environment 300. The virtual computing environment 300 includes a virtual machine manager (VMM) 320 and a virtual machine (VM) layer 340 that runs one or more virtual machines (VMs) 350a-n, which are configured to execute instances 362a, 362a-n of one or more software applications 360. Each host 210 may include one or more physical central processing units (pCPUs) 212 ("data processing hardware 212") and associated memory hardware 216. Although each hardware resource or host 210 is shown as having a single physical processor 212, any hardware resource 210 may include multiple physical processors 212. Host 210 also includes physical memory 216, which can be divided into virtual memory by host operating system (OS) 220 and allocated for use by VM 350 in VM layer 340 or even for use by VMM 320 or host OS 220. Physical memory 216 can include random access memory (RAM) and / or disk storage devices (including Figure 1 Storage device 154 is shown as being accessible via infrastructure 160).
[0049] The host operating system (OS) 220 may be executed on a given host 210, or may be configured to operate on a set of multiple hosts 210. For convenience, Figure 2 The host OS 220 is shown as being on machines 2101 to 210 m Furthermore, although the host OS 220 is shown as part of the virtual computing environment 300, each host 210 is equipped with its own OS 218. However, from the perspective of the virtual environment, the operating system on each machine appears and is managed as the collective operating system 220 to the VMM 320 and VM layer 340.
[0050] In some examples, the VMM 320 corresponds to a hypervisor 320 (e.g., a compute engine) that includes at least one of software, firmware, or hardware configured to create, instantiate / deploy, and execute VMs 350. A computer (e.g., data processing hardware 212) associated with the VMM 320 that executes one or more VMs 350 is generally referred to as a host 210 (as used above), while each VM 350 can be referred to as a client. Here, the VMM 320 or hypervisor is configured to provide each VM 350 with a corresponding guest operating system (OS) 354 (e.g., 354a-n) having a virtual operating platform and manage the execution of the corresponding guest OS 354 on the VM 350. As used herein, each VM 350 can be referred to as an "instance" or "VM instance." In some examples, multiple instances of various operating systems can share virtualized resources. For example, a first VM 350 of an operating system, The second VM 350 of the operating system and the OS A third VM 350 of an operating system may all run on a single physical x86 machine.
[0051] The VM layer 340 includes one or more virtual machines 350. The distributed system 200 enables users (via one or more computing devices 110) to start VMs 350 on demand, i.e., by sending commands or requests 170 (e.g., a request to start a VM) to the distributed system 200 (including the cloud system 150) via the network 140. Figure 1 For example, the command / request 170 may include an image or snapshot associated with the corresponding operating system 220, and the distributed system 200 may use the image or snapshot to create the root resource 210 for the corresponding VM 350. Here, the image or snapshot within the command / request 170 may include a boot loader, the corresponding operating system 220, and a root file system. In response to receiving the command / request 170, the distributed system 200 may instantiate the corresponding VM 350 and automatically start the VM 350 upon instantiation.
[0052] VM 350 emulates a real computer system (e.g., host 210) and operates based on the computer architecture and functionality of a real or imaginary computer system, which may involve specialized hardware, software, or a combination thereof. In some examples, distributed system 200 authorizes and authenticates user device 110 before launching one or more virtual machines 350. An instance 362 of a software application 360, or simply an instance, refers to a VM 350 hosted on (executing on) the data processing hardware 212 of distributed system 200.
[0053] The host OS 220 virtualizes the underlying host hardware and manages the concurrent execution of one or more VM instances 350. For example, the host OS 220 can manage VM instances 350a-n, and each VM instance 350 can include an emulated version of the underlying host hardware or a different computer architecture. The emulated version of the hardware associated with each VM instance 350, 350an is referred to as virtual hardware 352, 352an. The virtual hardware 352 can include one or more virtual central processing units (vCPUs) ("virtual processors") that emulate one or more physical processors 212 of the host 210. The virtual processors can be interchangeably referred to as "computing resources" associated with the VM instance 350. The computing resources can include the target computing resource level required to execute the corresponding individual service instance 362.
[0054] The virtual hardware 352 may further include virtual memory that communicates with the virtual processor and stores guest instructions (e.g., guest software) that can be executed by the virtual processor for performing operations. For example, the virtual processor may execute instructions from the virtual memory that cause the virtual processor to execute a corresponding individual service instance 362 of the software application 360. Here, the individual service instance 362 may be referred to as a guest instance that cannot determine whether it is executed by the virtual hardware 352 or by the physical data processing hardware 212. The host's microprocessor may include processor-level mechanisms to enable the virtual hardware 352 to efficiently execute the software instance 362 of the application 360 by allowing guest software instructions to execute directly on the host's microprocessor without code rewriting, recompilation, or instruction emulation. The virtual memory may be interchangeably referred to as a "memory resource" associated with the VM instance 350. The memory resource may include a target memory resource level required to execute the corresponding individual service instance 362.
[0055] The virtual hardware 352 may further include at least one virtual storage device that provides runtime capacity for services on the physical memory hardware 212. The at least one virtual storage device may be referred to as a storage resource associated with the VM instance 350. The storage resource may include a target storage resource level required to execute the corresponding individual service instance 362. The guest software executing on each VM instance 350 may further be assigned a network boundary (e.g., an assigned network address) through which the corresponding guest software can communicate with other services running on the internal network 160 ( Figure 1 ), external network 140 ( Figure 1 ) or other processes that can be reached by both. A network boundary can be referred to as a network resource associated with a VM instance 350.
[0056] The guest OS 354 executing on each VM 350 includes software that controls the execution of a corresponding individual service instance 362 (e.g., one or more of 362a-n of applications 360) by the VM instance 350. The guest OS 354, 354an executing on a VM instance 350, 350an can be the same as or different from other guest OSs 354 executing on other VM instances 350. In some embodiments, the VM instance 350 does not require a guest OS 354 to execute individual service instances 362. The host OS 220 can further include virtual memory reserved for the kernel 226 of the host OS 220. The kernel 226 can include kernel extensions and device drivers and can perform certain privileged operations that are prohibited to processes running in the user process space of the host OS 220. Examples of privileged operations include accessing different address spaces, accessing special function processor units in the host 210, such as a memory management unit, and the like. The communication process 224 running on the host OS 220 may provide a portion of the VM network communication functionality and may execute in either the user process space or the kernel process space associated with the kernel 226 .
[0057] According to aspects of the disclosed technology, unrecoverable memory errors (e.g., bit flips) occurring on a host 210 implementing an MCE can be managed at the hypervisor level to mitigate and / or prevent crashes of affected guest VMs and limit the impact of unrecoverable memory errors to the affected guest VMs. For example, the BIOS associated with a given host 210 is configured such that an MCE generated by a pCPU 212 on the host is sent to the kernel 226. The MCE includes contextual information about the error, including, for example, the physical memory address, the severity of the error, whether the error is isolated, and the component within the pCPU that signaled the error. The kernel 226 forwards the error to the hypervisor 320. The hypervisor 320 then processes this information to identify the virtual memory associated with the error and any affected memory pages and associated VMs. Because VMs typically do not share virtual memory, a given memory error can be isolated to a given VM. Consequently, there is little to no risk of the error propagating beyond the affected VM. The hypervisor 320 then isolates the corrupted memory page to prevent guest OSes from accessing it. Next, the hypervisor notifies the affected guest operating system of the error by simulating the error. Specifically, the hypervisor injects an interrupt (e.g., interrupt 80) into the guest OS, which notifies the guest OS of the error. In this way, for example, only the VM affected by the error is notified of the error, and only the VM or the application associated with the VM can be restarted.
[0058] Furthermore, after being notified of corrupted virtual memory addresses or memory pages containing such addresses, the affected VM can avoid reading or accessing these memory locations, which results in error containment. For example, each memory read or access to a corrupted memory element generates an MCE. One aspect of the disclosed technology mitigates and / or avoids multiple reads or accesses to a corrupted memory element after a corrupted memory element is detected at the host level and notified of the error to the VMM and / or guest OS.
[0059] In other examples, user applications may be running on multiple virtual machines, and a memory error associated with a single VM may affect multiple VMs (e.g., a machine learning training job). In such an example, the impact of the error may require that more than one VM be notified of the error. For example, if the hypervisor has distributed one or more given jobs among multiple VMs, the hypervisor may broadcast the error to all affected VMs. In this instance, the user may decide that shutting down and restarting the affected applications is a viable option. In contrast, in cases involving a single virtual machine, keeping the virtual machine active by, for example, providing it with new memory pages or restarting it may be a viable option.
[0060] According to aspects of the disclosed technology, a scanner (e.g., scanner 301) can be used to identify memory errors before an MCE is detected by the BIOS. For each identified memory error, the scanner can provide contextual information about the error to the kernel, including, for example, the physical memory address, the severity of the error, whether the error is isolated, etc. The kernel can isolate the memory page belonging to each individual VM for which an error is detected by the scanner. The detection of memory errors and the isolation of memory pages do not require any interaction with the hypervisor and / or guest OS / guest applications. Therefore, the detection of memory errors and the isolation of affected memory pages are transparent to the execution of the guest VM.
[0061] The scanner actively searches for memory errors to isolate the affected memory from use by the host or VMs executing on it. By actively searching for memory errors, the scanner attempts to detect memory errors on "free pages" (i.e., memory pages not used by the host device or VMs running on the host device). By detecting memory errors on "free pages," the scanner, working in conjunction with the kernel, can isolate the bad memory pages and prevent them from being used by the host system of VMs executing on the host. As a result, the host device and VMs are completely isolated from bad memory and do not trigger the MCE, which could cause a freeze or shutdown. Furthermore, even if the scanner detects a memory error that is not on a free page, future accesses by the virtual machine will not cause the actual hardware MCE to signal, thus preventing susceptibility to various CPU bugs that could cause a recoverable MCE to signal an unrecoverable state. In other words, all future accesses to pages where the scanner detects errors are guaranteed to be recoverable; or, if the scanner detects errors on free pages, these future accesses are completely blocked.
[0062] Example procedure or method
[0063] Figure 3AAn example of a process flow or method 370 according to various aspects of the disclosed technology is shown. A host 372 includes a BIOS, a CPU, and a kernel (as part of its OS). The host is configured to detect uncorrectable memory errors and issue a machine check exception (MCE) in response to such detection. In addition, functionality is provided for classifying detected uncorrectable memory errors. For example, the classification may include where the error was found, whether the error is recoverable, and what type of recovery is allowed or required. For example, some hardware architectures forward context information that signals to the software that recovery is unavailable and therefore the kernel needs to enter emergency mode. A typical example of this occurring is when the execution context is corrupted (e.g., an error occurs during the execution of certain instructions by the CPU). When an uncorrectable memory error is detected in the host 372, the BIOS sends an MCE to the CPU, line 376.
[0064] The CPU then forwards the MCE message (depicted as #MC) to the kernel of host 372, line 378. #MC and MCE or MCE message may include the same context information or the same type of context information. A handler in the kernel (e.g., MCE or #MC handler) receives the MCE message (#MC) including the context information regarding the uncorrectable memory event and signals the MCE signal handler in hypervisor 386 (line 382). Signaling may occur via a bus error signal (e.g., SIGBUS). Hypervisor 386 decodes the MCE message and maps it to the virtual memory space associated with the VM supported by the affected host, line 388. In doing so, hypervisor 386 determines the virtual memory and memory page associated with the damaged memory element. In addition, hypervisor 386 emulates the MCE event, line 388. That is, hypervisor 386 translates the context information associated with the physical memory error into context information associated with the virtual memory location. Additionally, the hypervisor 386 may instantiate the processes necessary to migrate the VM on the affected host 372 to another host 373 , line 390 .
[0065] As described above, aspects of the disclosed technology include having the host kernel's MCE handler signal all relevant MCE details to the virtual machine manager or hypervisor. Utilizing the hypervisor, the MCE SIGBUS handler logs the memory error event in, for example, a VmEvents table. The event table may include fields for recording the following details: general VM metadata (e.g., VM id, project id); MCE details: DIMM, rank, bank, MCA registers from all relevant banks). Optionally, neighbor information may also be recorded, such as which other VMs are on the host, on the same socket, etc. Neighbor information may be important when analyzing potential security attacks (e.g., Row Hammer attacks). In such an example, the disclosed technology can notify the guest user space of all affected VMs and result in a more graceful failover to another host.
[0066] On the host, memory error containment and memory error recovery, as well as IIO stalls and screams, are enabled in the BIOS. Error signaling via a specific new MSI / NMI handler is added to the host kernel, and its behavior is simply a panic to the host. The host kernel is configured to know which address space the MCE error belongs to and whether the process is a VM.
[0067] Figure 3B An example of a process flow or method 870 according to aspects of the disclosed technology is shown. A host 872 includes a scanner 801, a CPU 802, a kernel (as part of its OS) 803, and a memory 816. The scanner 801 is configured to detect memory errors within the memory 816 of the host device. When the scanner 801 encounters a memory error, as shown by line 876, the CPU 802 can generate an MCE and send it to the kernel 803, as shown by line 878. The kernel 803 can then determine to which address space (e.g., memory page) the error notification belongs. The host device can then block ("poison") the affected memory page from being accessed by the host 872 and the VMs executing thereon, as shown by line 888.
[0068] The scanner 801 can be a software scanner executing on the host 872 and / or a separate hardware component within the host 872 or otherwise in communication with the host 872. The scanner can actively scan the entire system memory of the host 872 for errors. By doing so, the scanner can identify memory problems before the memory is used by the VM or the host itself.
[0069] For example, scanner 801 may scan the entire memory of the host device every X minutes. In this case, X is the upper limit of the number of errors that can be discovered. That is, any error will be identified by the scanner no more than X minutes after it occurs. By using the scanner to proactively detect memory errors, any memory pages with errors can be prevented from being used, so that the host device and VMs executing on the host do not rely on the erroneous memory pages. Furthermore, any virtual machines executing on the affected host can be migrated to a new host to allow the erroneous memory to be repaired. Each error detected by the scanner can also be provided to user space for future review by the user or other administrators of the host device and / or VMs executing on the host device.
[0070] The scanner 801 can perform read-only or read-write scanning to avoid changing the contents of the scanned memory. By performing read-only scanning, the scanner 801 can avoid interfering with the memory contents, which belong to any software executed on the host 872, including the operating system / kernel 803, applications, virtual machines, etc.
[0071] To minimize the processing overhead and memory bandwidth & cache contention introduced by the scanner 301, the memory copy can be offloaded to an integrated DMA engine, such as Crystal Beach DMA. Additionally, the scanner can ensure that non-uniform memory accesses are local socket-local chunked reads and return as early as possible once a memory error is detected in the current block, without having to complete the read of the remaining bytes of the block. In addition, the scanner can use non-temporal instructions to eliminate any cache pollution. In this regard, non-temporal instructions on x86 are a special set of instructions that provide the CPU cache hierarchy with a "non-temporal" property, i.e., "do not cause memory contents to be stored in the cache." In this way, the scanner can run continuously in the background without any or minimal impact on the workload performance of the host device and the VMs executing on the host device.
[0072] The CPU 802 may be configured to send an MCE to the kernel 803 of the host device 872. Upon receipt, the kernel 803 may review information contained in the MCE, such as the physical location of the memory where the scanner 801 encountered the memory error. The kernel 803 may poison the affected memory so that it cannot be accessed by a VM executing on the host device 872 or by the host device 872 itself.
[0073] This proactive scanning contrasts with previous error detection methods that rely on CPU-generated MCEs after a host or a VM executing on the host encounters a memory error. Therefore, before mitigation techniques are implemented, at least the host or one or more VMs executing on the host are affected by the error.
[0074] In the event that a memory error occurs on a free page, no additional errors related to the memory error are detected because the free page is not an option for use by the host device or a VM executing on the host device. That is, when the scanner encounters an error at these memory locations, the unused memory locations are not being used by the system. Therefore, these memory locations do not continue to trigger MCEs or other errors.
[0075] Even if a memory error occurs on a memory page used by a VM, a page fault is provided to the VM. The page fault will be detected (or otherwise provided) to the kernel 803 of the host device 872. The kernel can provide a "Sigbus" signal with an MCE error code to the hypervisor of the VM where the page fault occurred. The hypervisor can then send a simulated MCE to the guest vCPU, which can handle the simulated MCE as needed. Therefore, only the VM using the memory where the memory error occurred may be affected by the memory error. Other VMs and the host device may not be affected by the memory error.
[0076] As described above, aspects of the disclosed technology include having the host kernel's MCE handler signal all relevant MCE details to the virtual machine manager or hypervisor. Using the hypervisor, the MCE SIGBUS handler logs the memory error event in, for example, a VmEvents table. The event table can include fields for recording the following details: general VM metadata (e.g., VM id, project id); MCE details: DIMM, rank, bank, MCA registers from all relevant banks). Optionally, neighbor information can also be recorded, such as which other VMs are on the host, on the same socket, etc. Neighbor information can be important when analyzing potential security attacks (e.g., Row Hammer attacks). In such an example, the disclosed technology can notify the guest user space of all affected VMs and result in a more graceful failover to another host.
[0077] On the host, memory error containment and memory error recovery are enabled in the BIOS, along with IIO stalling and screaming. Error signaling via a specific new MSI / NMI handler is added to the host kernel, which behaves as a simple panic to the host. The host kernel is configured to know which address space the MCE error belongs to and whether the process is a VM.
[0078] Figure 3C374 according to aspects of the disclosed technology. Host 372 and host 373 may include various components, including BIOS, CPU, and OS / kernel. In addition, host 372 and host 373 may include volatile memory and non-volatile memory that may be divided into multiple segments. Host 372 and host 373 may be similar to distributed system 200 or host 210 as described above.
[0079] VMM / hypervisor 386 may run on the host. As described above, VMM / hypervisor 386 may control, coordinate, or otherwise enable the creation and operation of one or more VMs, such as VM 391. A to 391 N Although only two are shown for simplicity, it should be understood that more than two virtual machines (e.g., 100 or even 1000) may be instantiated or run on host 372. Each VM may correspond to a portion of volatile memory or other memory on host 372. Host 372 and host 373 need not reside in the same data center to be part of a given cloud system (e.g., see Figure 1 In some examples, migration may occur between hosts in different data centers within a cloud environment. In this case, VMM / hypervisor 386 may include different VMM / hypervisor components located in different physical locations or different data centers. Furthermore, in some examples, VMM / hypervisor 386 may include different components within a given data center, depending on how the underlying hosts are managed. Furthermore, the VMM / hypervisor may be functionally distributed across multiple hosts or machines.
[0080] In some examples, as indicated by the "!" in the memory, it may be known that certain sectors or portions of memory within the host 372 contain unrecoverable errors. Figure 3C As explained above, these unrecoverable errors can affect the operation of the virtual machine and the client applications or instances it supports. As an example, VM391 A A virtual machine may be running on a specific segment of memory that contains an unrecoverable error related to an MCE. Other virtual machines may be using physical hardware that does not contain the error, including volatile memory. The physical memory and other physical components used by a given VM are managed by the VMM / hypervisor 386. For example, even though a given VM may be using space on the host's physical memory, the actual physical memory addresses are typically unknown because the VMM typically maps them to virtual memory addresses within the virtual machine environment.
[0081] Various memory segments may correspond to one or more memory pages, which are Figure 3C372 is shown as part of host 372. In some examples, one or more storage pages may be a memory dump of one or more segments of volatile memory of host 372. The storage pages may be stored in any suitable memory, such as low-level cache memory, non-volatile memory, or volatile memory. Certain storage pages may be marked as corresponding to having unrecoverable errors (e.g., MCE). In some examples, a page may be marked as or contain information identifying the page as "poisoned" or containing poisoned memory. In some examples, a storage page may contain only "guest memory" or memory corresponding to a particular VM instance, such as Figure 2 Example 362a cited in .
[0082] Host 373 may be similar to host 372, and hypervisor 386 may control, coordinate, or otherwise enable the operation of one or more VMs, such as a particular VM 392 on host 373. A to 392 N In some examples, the number of VMs on host 373 can be the same as the number of VMs on host 372 .
[0083] Memory migration module 371 may include remote procedure calls, APIs, networking functions, and other "low-level" memory operations, such as those that occur below or at the OS level, to enable a virtual machine to be transferred or migrated from one host to another. Memory migration module 371 may be distributed across one or more physical machines, such as host 372 and host 373. Memory migration module 371 may also run on a network that connects host 372 and host 373 or other hosts or allows data to be transferred between host 372 and host 373 or other hosts.
[0084] The memory migration module 371 is also capable of generating checksums, reading from bounce buffers, and being aware of MCE errors in memory and storage pages. The memory migration module 371 may use RPC, software, or other APIs suitable for performing the dynamic migration functionality of the disclosed technology. The memory migration module 371 may include one or more modules that perform the functionality of the migration process (as described herein) and may be implemented as an instruction set running on one or more processing devices.
[0085] In some examples, the memory migration module 371 can be "generic" and include modules that are abstracted and compatible across different types of hardware and physical hosts (e.g., those containing different models of processors) and understand specific memory or other error codes generated by a particular physical machine.
[0086] Figure 4 A method or process 400 is shown in accordance with aspects of the disclosed technology.
[0087] Method 400 may include proactively detecting and forwarding MCEs associated with uncorrectable memory errors to a virtual machine manager or hypervisor by a scanner. The MCE information is decoded by the virtual machine manager or hypervisor and mapped to the affected memory pages, and therefore to the affected virtual machines. The virtual machine manager or hypervisor may then initiate a process to migrate the VM to another device. Further details regarding these operations are described herein.
[0088] As shown in block 401 , a scanner scans the host's memory for errors.
[0089] At block 403 , the scanner detects memory errors in the host's memory.
[0090] At block 405 , the scanner generates an error signal upon detecting a memory error.
[0091] At block 407 , the scanner transmits an error signal to one or more processors of the host.
[0092] Figures 5A-5D Various aspects of live memory migration from a "source VM" to a "target VM" are shown. Figures 5A-5D As shown, aspects of the migration can be described with respect to time, such as "before copy" and "after copy." Additionally, during a live memory migration, the operating states of the source and target VMs can be described. In some examples, and as in Figures 5A-5D As used in the present invention, the "time arrow" moves sequentially from left to right into the future, indicating that an example of an action block can be performed. However, one skilled in the art will recognize that the order of the processes can be exchanged or reversed, and that certain processes can be repeated.
[0093] As in Figures 5A-5D As used herein, a "source VM" may be a virtual machine from which data or information is migrated, and a "target VM" may be a virtual machine to which data or information is migrated. In some examples, the "source VM" may be associated with or run on a particular physical machine, such as host 210. In some examples, method 500 may begin once a particular error, such as the MCE described above, occurs on a physical machine associated with the source VM. "Source" may refer to a source VM or a host corresponding to the source VM, and "target" may refer to a target VM or a target machine corresponding to the target VM.
[0094] Those skilled in the art will recognize that Figures 5A-5D The specific implementation of the described method may vary and involve one or more software modules, APIs, RPCs, and use one or more types of data structures, logs, binary structures, and hardware to perform the method.
[0095] Figure 5A An example method 500 is shown. Figure 5A 5. In summary, the method 500 may consist of operations that may be conceptualized within a "pre-copy" phase and a "post-copy" phase. The method 500 may be described with reference to Figures 5B-5D Any combination of the described processes, including method 520 , method 530 , and method 540 .
[0096] During the pre-copy phase, "guest memory" may be copied from the source VM 510 to the target VM 515. The guest memory may include memory created in a guest user space or a guest user application. In some examples, guest memory may also refer to underlying physical memory that corresponds to specific virtual memory belonging to a specific guest user space or virtual machine instance. During the pre-copy phase, the source VM 510 runs on the associated source physical machine. During this phase, one or more processors copy the guest memory to the target. For example, the memory contents are copied to a network buffer and transmitted over the network (e.g., via an RPC protocol) via the client computer. Figure 1 The client memory is sent to the target VM 515 via the network 160), where a corresponding RPC receiver thread is provided on the target VM to receive the client memory and store the received client memory into the corresponding client physical address.
[0097] like Figure 5A As shown, during the pre-copy and post-copy periods, the source or target may enter a brownout period, during which the VM is not paused while the migration is in progress. During this phase, guest execution may slow down due to reasons such as dirty tracking or post-copy network page-ins.
[0098] Figure 5B Aspects of method 500 or method 520 related to the "pre-copy" phase are shown. One or more memory migration modules can be used, which can be sets of instructions for reading, writing, and tracking memory and storage pages. Method 520 can be performed while the source virtual machine is "live" or active, allowing the user to continue using the virtual machine while method 520 is in progress.
[0099] like Figure 5BAs shown, during the migration process from source to target, some storage pages may be modified due to user processes or other processing occurring on the source virtual machine. These differences can be tracked. Storage pages that have been modified during the transfer of guest memory can be referred to as "dirty pages." In some examples, only a subset of specific pages can be transferred during the pre-copy phase. In some cases, poison pages may include a subset of dirty pages, but such dirty pages will be skipped or not processed as part of regular dirty page processing on the migration target.
[0100] The guest memory on the source VM 510 can be read and written to the guest memory of the target VM 515. In some examples, the reading and writing processes can be performed using one or more remote procedure calls or RPCs. In some examples, the remote procedure calls can use pointers to specific memory contents to identify one or more portions of physical memory or virtual memory to be copied from the source VM 510 to the target VM 515.
[0101] In some examples, a bounce buffer can be used as part of the transfer. A bounce buffer is a type of memory that resides architecturally low enough in memory for the processor to copy data from and write data to. Pages can be allocated in the bounce buffer to organize memory. As part of incrementally updating a "dirty bitmap" and copying dirty pages, the memory migration module can repeatedly pass through the memory and dirty pages multiple times.
[0102] In some examples, "poisoned pages," or pages containing unrecoverable errors, can also be tracked and identified. In some examples, "poisoned" pages can be selectively excluded from the memory migration process and dirty pages. In some examples, when an MCE is discovered, the memory pages associated with the MCE can be marked as poisoned. The memory migration module can notify the memory bus that a particular page is "poisoned" and cause the memory bus to avoid copying that memory from the source to the target.
[0103] Method 520 may also include generating a checksum. After writing the client memory from the source to the target, a checksum may be generated. A checksum may be generated on the source memory page and the associated target memory page to ensure that the transfer of the memory page was error-free. In some examples, checksum generation or checksum checking may be skipped for poisoned pages.
[0104] Method 520 may include the following or similar process described as pseudo code:
[0105]
[0106]
[0107] In other words, the method 520 can implement tracking of dirty memory pages and, when a "blackout" process is not implemented, prepare a dirty memory page log and, for each dirty memory page log, send updates related to changes in the dirty memory page log from the source to the target, such as via a bitmap. In addition, as part of the method 520, checksums and tracking of changes can be performed by the memory migration module.
[0108] Figure 5C Aspects of method 500 or method 530 during a power outage are shown. During a power outage, the source is in a "suspended" state and users are unable to operate or use the source virtual machine. During method 530, poisoned pages can also be tracked and subtracted.
[0109] Method 530 may begin at the beginning of a power outage or when the source VM is paused. The memory migration module may traverse the source VM's memory to identify the most recent memory or the "last memory" before the power outage and send a dirty bitmap to the target.
[0110] During method 530, information related to the poisoned page or the poisoned page itself may be copied. Because poisoned pages are expected to be rare, in some examples, a structure other than a "bitmap" may be used to transmit the poisoned page or information related to the poisoned page to limit memory overhead. In some examples, the poisoned page may be sent only once at the beginning of the blackout period because changes to the poisoned page are expected to be minimal and the poisoned page itself is expected to be rare.
[0111] Method 530 may include the following or similar process described as pseudo code:
[0112] OnBlackout():
[0113] PauseGuest()
[0114] poisonedpages←ReadMemoryPoisonedLog()
[0115] dirtypages←ReadMemoryDirtyLog()
[0116] StartPostCopyOnTarget(dirtypages-poisonedpages)
[0117] In other words, in method 530 , one or more memory logs may be read, and by reading the memory logs, only dirty pages other than poison pages may be copied.
[0118] Figure 5DAspects of method 500 or method 540 are shown, which may involve a "post-copy" phase. At this phase, certain information has been transferred from the source to the target. At this phase, the virtual machine running on the source can now be run on the target. At this phase, the virtual machine running on the target may be different from the virtual machine running on the source because dirty and poisoned memory pages have not yet been transferred.
[0119] During the post-copy period or as part of method 540, the final dirty bitmap may be used to initialize "demand paging." The "demand paging" control may initialize the background fetcher module with the same dirty bitmap. As described above, this bitmap may already know or contain a list of poisoned pages that were subtracted.
[0120] During post-copy or as part of method 540, background fetches of memory pages that have not yet been fetched or migrated from the source may be accessed by the background fetcher module or the memory migration module.
[0121] In some examples, on the target, upon request of a particular memory page that has not yet been transferred from the target to the source, a remote memory access (RMA) or other remote procedure call may be used by the target to access memory pages that have not yet been migrated to the target.
[0122] After receiving the memory page at the target, a checksum can be generated for the obtained memory content when no MCE errors occur or are associated with the specific memory page. The checksum can be used to verify whether the memory migration process was performed correctly.
[0123] Method 540 may include the following process or similar processes, which may be described as pseudo code:
[0124]
[0125]
[0126] As described above, the disclosed technology provides techniques, systems, and apparatus for proactively detecting, containing, and recovering from uncorrectable memory errors in a distributed computing environment. One aspect of the disclosed technology includes a scanner on a host scanning the host's memory for errors. After the scanner detects an error, the scanner may generate an error notification. The scanner may send the error notification to one or more processors on the host to implement mitigation techniques.
[0127] Unless otherwise stated, the foregoing alternative examples are not mutually exclusive, but may be implemented in various combinations to achieve unique advantages. Since these and other variations and combinations of the above-described features can be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be considered as illustrative of, and not restrictive of, the subject matter defined by the claims. Furthermore, the provision of examples described herein and the use of terms such as "for example," "including," and the like should not be construed as limiting the subject matter of the claims to specific examples; rather, these examples are intended to illustrate only one of many possible embodiments. Furthermore, the same reference numerals in different figures may identify the same or similar elements.
Claims
1. A method for proactively detecting memory errors in a cloud computing environment, characterized in that: include: Scanning a memory of the host for errors by a scanner of the host; detecting, by the scanner, a memory error in the memory of the host; identifying one or more memory pages determined to be associated with the memory error as one or more poisoned memory pages based on a location in the memory where the memory error is detected by the scanner; generating, by one or more processors of the host, a machine check exception in response to the scanner detecting the memory error; providing, by the one or more processors, the machine check exception to a kernel executing on the host; as well as The kernel sends the machine check exception to a hypervisor of a virtual machine executing on the host; The hypervisor identifies the virtual memory and the virtual machine associated with the memory error and notifies an affected guest operating system of the memory error by simulating the memory error; Access by the host to the one or more poisoned memory pages is isolated.
2. The method according to claim 1, characterized in that The scanning is performed continuously by the scanner.
3. The method according to claim 1, characterized in that The scan is a read-only scan.
4. The method according to claim 1, wherein The memory error is an uncorrectable memory error.
5. The method according to claim 1, wherein The machine check exception includes an indication of a location in the memory where the memory error was detected by the scanner.
6. The method according to claim 1, characterized in that Further including: receiving a page fault associated with a read request made by a guest of the virtual machine executing on the host; as well as A SIGBUS signal is sent by the kernel to the hypervisor of the virtual machine.
7. The method according to claim 6, characterized in that Further including: generating, by the hypervisor, a machine check exception; as well as The machine check exception is sent to the client.
8. A cloud computing system, characterized in that: include: A host, capable of supporting one or more virtual machines; a scanner configured to scan a memory of the host for errors; as well as One or more processors configured to: receiving, from the scanner, an indication of a detection of a memory error in the memory of the host; identifying one or more memory pages determined to be associated with the memory error as one or more poisoned memory pages based on a location in the memory where the memory error is detected by the scanner; generating a machine check exception in response to receiving the indication of the memory error from the scanner; sending the machine check exception to a kernel of the host, the machine check exception including information associated with the memory error; as well as The kernel sends the machine check exception to a hypervisor of a virtual machine executing on the host; The hypervisor identifies the virtual memory and the virtual machine associated with the memory error and notifies an affected guest operating system of the memory error by simulating the memory error; Access by the host to the one or more poisoned memory pages is isolated.
9. The system according to claim 8, characterized in that The scanning is performed continuously.
10. The system according to claim 8, wherein: The scan is a read-only scan.
11. The system according to claim 8, wherein: The memory error is an uncorrectable memory error.
12. The system according to claim 8, wherein: The machine check exception includes an indication of a location in the memory where the memory error was detected by the scanner.
13. The system according to claim 8, wherein: The one or more processors are further configured to: receiving a page fault associated with a read request made by a guest of the virtual machine executing on the host; as well as A SIGBUS signal is sent to the hypervisor of the virtual machine.
14. The system according to claim 13, wherein: The one or more processors are further configured to: Generate a machine check exception; and The machine check exception is sent to the client.
15. A non-transitory computer-readable medium storing instructions, characterized in that: When executed by one or more processors, the instructions cause the one or more processors to: receiving, from the scanner, an indication that a memory error in a memory of the host was detected; identifying one or more memory pages determined to be associated with the memory error as one or more poisoned memory pages based on a location in the memory where the memory error is detected by the scanner; generating a machine check exception in response to receiving the indication of the memory error from the scanner; sending the machine check exception to a kernel of the host, the machine check exception including information associated with the memory error; as well as The kernel sends the machine check exception to a hypervisor of a virtual machine executing on the host; The hypervisor identifies the virtual memory and the virtual machine associated with the memory error and notifies an affected guest operating system of the memory error by simulating the memory error; Access by the host to the one or more poisoned memory pages is isolated.
16. The non-transitory computer-readable medium of claim 15, wherein: The machine check exception includes an indication of a location in the memory where the memory error was detected by the scanner.
Citation Information
Patent Citations
Processor error event handler
US20190034252A1
Systems and methods for memory failure prevention, management, and mitigation
US20200210272A1