Fault processing method and device, electronic equipment and computer readable storage medium
By obtaining and reverse mapping of the host physical address in memory failure information, determining the fault virtual address in the virtual machine simulator process and sending it to the user-state process, instructing the virtual machine to isolate the fault memory hardware, solving the problem of virtual machine memory failure isolation in the cloud computing environment, achieving the effect of reducing downtime and avoiding application errors.
Patent Information
- Application Number
- CN202411898050.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In a cloud computing environment, memory failures of virtual machines cannot be effectively isolated, resulting in errors in applications within the virtual machine or host downtime, and the existing technology cannot achieve memory failure isolation without placing all memory.
By obtaining the host physical address in the memory failure information, performing reverse mapping to determine the corresponding user state process, obtaining the fault virtual address in the virtual machine simulator process, and sending it to the user state process to instruct the virtual machine to isolate the failed memory hardware.
It effectively isolates memory failures without surviving all memory, reduces the downtime rate of virtual machines and hosts, and avoids errors in the virtual machine applications.
Smart Images

Figure CN120045274A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to technical fields such as cloud computing, virtual machines, and computer hardware failures. Specifically, embodiments of the present disclosure relate to a fault handling method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the field of cloud computing, virtualization technology is used to split large servers into multiple small virtual machines for multiple customers to use.
[0003] There may be VFIO (a device passthrough technology for the Linux operating system) devices in these small virtual machines. To support the DMA (Direct Memory Access) of the VFIO devices in the virtual machines, it is necessary to pin the HPA (Host Physical Address) memory corresponding to the GPA (Guest Physical Address) involved in the DMA transfer in the virtual machine to prevent this memory from being swapped by the host to the swap partition or moved due to memory consolidation. Summary of the Invention
[0004] The present disclosure provides a fault handling method and apparatus, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect of the present disclosure, a fault handling method is provided, the method comprising:
[0006] Obtaining memory fault information reported by a faulty memory hardware, the memory fault information including a fault host physical address of the faulty memory hardware on the host;
[0007] Determining, through reverse mapping according to the fault host physical address, a user-mode process corresponding to the fault host physical address;
[0008] When the user-mode process is a virtual machine emulator process, obtaining a fault host virtual address corresponding to the fault host physical address according to address information of a virtual memory area;
[0009] Sending the fault host virtual address to the user-mode process for the user-mode process to isolate the memory hardware of the faulty virtual machine running on the host according to the fault host virtual address.
[0010] According to a second aspect of the present disclosure, a fault handling apparatus is provided, the apparatus comprising:
[0011] An information collection module, configured to obtain memory fault information reported by a faulty memory hardware, where the memory fault information includes the physical address of the faulty memory hardware on the host machine in the host machine;
[0012] A reverse mapping module, configured to determine, through reverse mapping according to the physical address of the faulty host machine, the user-mode process corresponding to the physical address of the faulty host machine; in the case where the user-mode process is a virtual machine emulator process, obtain, according to the address information of the virtual memory area, the virtual address of the faulty host machine corresponding to the physical address of the faulty host machine;
[0013] A fault injection module, configured to send the virtual address of the faulty host machine to the user-mode process, so that the user-mode process isolates the memory hardware according to the virtual address of the faulty host machine for a faulty virtual machine running on the host machine.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the above-mentioned fault handling method.
[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above-mentioned fault handling method.
[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the above-mentioned fault handling method when executed by a processor.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0022] Figure 1 is a flowchart of a fault handling method provided by an embodiment of the present disclosure;
[0023] Figure 2Shows the overall architecture diagram of the host and virtual machines in the embodiments of the present disclosure;
[0024] Figure 3 Is a schematic flowchart of some steps of another fault handling method provided by the embodiments of the present disclosure;
[0025] Figure 4 Is a schematic flowchart of some steps of another fault handling method provided by the embodiments of the present disclosure;
[0026] Figure 5 Is a schematic flowchart of some steps of another fault handling method provided by the embodiments of the present disclosure;
[0027] Figure 6 Is a schematic flowchart of some steps of another fault handling method provided by the embodiments of the present disclosure;
[0028] Figure 7 Is a schematic flowchart of some steps of another fault handling method provided by the embodiments of the present disclosure;
[0029] Figure 8 Is a schematic structural diagram of a fault handling device provided by the embodiments of the present disclosure;
[0030] Figure 9 Is a block diagram of an electronic device for implementing the fault handling method of the embodiments of the present disclosure. Detailed implementation manners
[0031] The following makes an explanation of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0032] In some related technologies, QEMU (a simulation processor with source code distributed under the GPL license) is used as a virtual machine simulator.
[0033] When QEMU starts a virtual machine, it does not know which GPAs the virtual machine will use to participate in the DMA transfer of the virtual machine. Therefore, QEMU pins all the memory corresponding to the virtual machine, and the pinned memory cannot be moved. This leads to the situation that if a memory fault occurs in this memory, the host cannot isolate the faulty memory by copying the content of the faulty memory to a new memory page.
[0034] Therefore, once a memory fault occurs, one can only wait for the host to crash or the application in the virtual machine to malfunction.
[0035] The fault handling method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure aim to solve at least one of the above technical problems in the prior art.
[0036] Figure 1 The flowchart of the fault handling method provided by the embodiments of the present disclosure is shown. As Figure 1 shown, the fault handling method provided by the embodiments of the present disclosure may include step S110, step S120, step S130, and step S140.
[0037] In step S110, memory fault information reported by the faulty memory hardware is obtained. The memory fault information includes the fault host physical address of the faulty memory hardware in the host machine.
[0038] In step S120, the user-mode process corresponding to the fault host physical address is determined through reverse mapping based on the fault host physical address.
[0039] In step S130, when the user-mode process is a virtual machine emulator process, the fault host virtual address corresponding to the fault host physical address is obtained according to the address information of the virtual memory area.
[0040] In step S140, the fault host virtual address is sent to the user-mode process for the user-mode process to isolate the faulty memory hardware running on the host machine according to the fault host virtual address.
[0041] For example, a host machine refers to a physical computer or server running a hypervisor, which can manage and allocate physical resources such as a CPU (Central Processing Unit), memory, storage, and network connections for virtual machines to use. The host machine creates and runs virtual machines on its own, providing an isolated execution environment for each virtual machine.
[0042] Figure 2 The overall architecture diagram of the host machine and virtual machines in the embodiments of the present disclosure is shown. As Figure 2 shown, when a fault occurs in the memory hardware of the host machine, the memory hardware reports the fault event to the MCE (Machine Check Exception) driver in the host by sending memory fault information. The memory fault information will carry the HPA (Host Physical Address) of the faulty memory hardware in the host machine, that is, the fault HPA.
[0043] Other modules of the host machine can register callback functions with the MCE driver, and the MCE driver sequentially calls each callback function to handle the fault event.
[0044] The fault handling method provided by the embodiments of the present disclosure can be executed by a module of the host machine, or can be controlled by an electronic device such as a terminal device or a server to execute a module of the host machine. Specifically, a callback function implementing the fault handling method provided by the embodiments of the present disclosure can be registered with the MCE driver of the host machine, so that after the MCE driver receives the memory fault information reported by the memory hardware, the registered callback function is called, thereby running the fault handling method provided by the embodiments of the present disclosure on the host machine.
[0045] In some possible implementation manners, in step S110, as Figure 2 shown, after the MCE driver receives the memory fault information reported by the faulty memory hardware, it will call the callback function and transmit the memory fault information to the callback function.
[0046] In some possible implementation manners, the callback function can also be registered with the MCE to obtain
[0047] In some possible implementation manners, in step S120, as Figure 2 shown, the callback function can call the rmap (reverse mapping) function of the kernel of the host operating system to find the user-mode process corresponding to the faulty HPA.
[0048] Among them, the user mode is a running level in the operating system, mainly used to run user programs. The user-mode process corresponding to the faulty HPA is the process corresponding to the user program that uses the faulty HPA in the host machine.
[0049] In the kernels of some operating systems (such as linux), both processes and threads are regarded as Tasks, which are managed by a unified data structure task_struct. task_struct contains various attributes of the process, such as the process PID (Process ID, process control symbol), status, priority, scheduling information, etc., and is used to manage the life cycle, scheduling, memory management, file system operations, etc. of the process.
[0050] Among them, the comm field of task_struct is used to store the name of the process. The name of the user-mode process can be obtained through the comm field of task_struct of the user-mode process corresponding to the faulty HPA.
[0051] The user program corresponding to the user-mode process can be determined through the name of the user-mode process.
[0052] If the user-mode process corresponding to the faulty HPA is not a virtual machine simulator process, it means that the user program using the faulty HPA is other user programs except the virtual machine simulator.
[0053] If the user-mode process corresponding to the faulty HPA is a virtual machine emulator process, it indicates that the user program currently using the faulty HPA is a virtual machine emulator. That is to say, there is a virtual machine running on the host using the faulty HPA. If memory isolation is not performed on the faulty memory hardware, it may cause errors in the application programs within the virtual machine and even lead to the host crashing.
[0054] Therefore, when the user-mode process corresponding to the faulty HPA is a virtual machine emulator process, the HVA (Host Virtual Address) corresponding to the faulty HPA is calculated through the address information in the VMA (virtual memory area) area in the rmap, that is, the faulty HVA.
[0055] Among them, the host uses VMA to manage the virtual memory space. The virtual memory of each process is composed of multiple VMAs. Therefore, the faulty HVA corresponding to the faulty HPA can be obtained through the VMA.
[0056] In some possible implementation manners, QEMU (a simulated processor distributed with source code under the GPL license) is used as the virtual machine emulator. Therefore, if the user-mode process is a virtual machine emulator process, it is also a QEMU process.
[0057] In some possible implementation manners, in step S130, the faulty HVA is sent to the user-mode process (which is also a virtual machine emulator process) obtained through the rmap by the send_sig_mceerr method.
[0058] As Figure 2 shown, after the virtual machine emulator process obtains the faulty HVA, it searches for the GPA (Guest Physical Address) mapped by the faulty HVA through the faulty HVA, that is, the faulty GPA. The virtual MCE register of the faulty virtual machine corresponding to the virtual machine emulator process is filled with the faulty GPA, and then an MCE interrupt or a CMCI (Corrected Machine Check Interrupt) interrupt is injected into the faulty virtual machine to pass the faulty GPA into the faulty virtual machine.
[0059] After receiving the MCE interrupt or the CMC interrupt, the faulty virtual machine uses the Soft Offline or Hard Offline function of the memory to perform memory isolation on the memory corresponding to the faulty GPA.
[0060] In the fault handling method provided by the embodiments of the present disclosure, through the fault HPA of the memory hardware where the fault occurs, the virtual machine simulator process corresponding to the fault HPA and the corresponding fault HVA are obtained. By sending the fault HVA to the virtual machine simulator process, the memory fault is injected into the virtual machine, and the virtual machine isolates the memory hardware where the fault occurs, avoiding all the memory corresponding to the virtual machine from being pinned, preventing the virtual machine from accessing the memory hardware where the fault occurs, which may cause the application program in the virtual machine to go wrong and even cause the host machine to crash, and reducing the crash probability of the host machine.
[0061] The following specifically introduces the fault handling method provided by the embodiments of the present disclosure.
[0062] In some possible implementation manners, the memory hardware reports the fault event to the MCE driver in the host by sending memory fault information. In addition to carrying the HPA of the memory hardware where the fault occurs in the host machine, that is, the fault HPA, the memory fault information also carries the level of the memory fault that occurs.
[0063] In some possible implementation manners, the levels of memory faults may include CE, AO, and AR.
[0064] Among them, CE is Correctable Error, that is, a correctable error; AO is Action Optional, a fault that can take measures for processing; AR is Action Required, a fault that must take measures for processing.
[0065] In some possible implementation manners, a signal representing the level of the memory fault is also sent while sending the fault HVA to the virtual machine simulator process. Specifically, it may be the BUS_MCEERR_AO signal, representing that the memory fault is a fault that can take measures for processing; it may be the BUS_MCEERR_AR signal, representing that the memory fault is a fault that must take measures for processing; it may be the BUS_MCEERR_CE signal, representing that the memory fault is a correctable error.
[0066] After obtaining the fault HVA and the signal representing the level of the memory fault, the virtual machine simulator process searches for the fault GPA mapped by the fault HVA through the fault HVA, fills the virtual MCE register of the faulty virtual machine with the fault GPA and the memory fault level, and then injects an MCE interrupt or a CMCI interrupt into the faulty virtual machine to pass the fault GPA and the memory fault level into the faulty virtual machine.
[0067] After receiving the MCE interrupt or CMC interrupt, in the case of the centos series image of the operating system in the virtual machine, the faulty virtual machine uses the mcelog user-state daemon process in the virtual machine to judge the memory fault level.
[0068] In the case where the memory failure level is AO or AR, immediately call the Hard Offline function of the virtual machine to isolate the memory with the memory failure, and kill the virtual machine process that is using the faulty memory hardware.
[0069] In the case where the memory failure level is CE, count the received CE memory failures. If the number of CE memory failures reaches a preset threshold, call the Soft Offline function of the virtual machine to isolate the memory with the memory failure.
[0070] If there is a function for predicting memory UE (Uncorrectable Error) failures in the host machine, the risk of CE becoming UE can be predicted in the host machine, and then the fault GPA corresponding to the CE memory failure with the risk of becoming a UE failure is sent to the virtual machine. After receiving it, the virtual machine can set the preset threshold to 1, that is, isolate the memory with the memory failure when receiving 1 CE memory failure with the risk of becoming a UE failure.
[0071] After the faulty virtual machine receives an MCE interrupt or a CMC interrupt, in the case of the virtual machine's operating system (in the case of ubuntu series images), use the CEC kernel-mode driver in the virtual machine to judge the memory failure level.
[0072] In the case where the memory failure level is AO or AR, immediately call the Hard Offline function of the virtual machine to isolate the memory with the memory failure, and kill the virtual machine process that is using the faulty memory hardware.
[0073] In the case where the memory failure level is CE, count the received CE memory failures. If the number of CE memory failures reaches a preset threshold, call the Soft Offline function of the virtual machine to isolate the memory with the memory failure.
[0074] If there is a function for predicting memory UE failures in the host machine, the risk of CE becoming UE can be predicted in the host machine, and then the fault GPA corresponding to the CE memory failure with the risk of becoming a UE failure is sent to the virtual machine. After receiving it, the virtual machine can set the preset threshold to 1, that is, isolate the memory with the memory failure when receiving 1 CE memory failure with the risk of becoming a UE failure.
[0075] By sending the memory failure level to the virtual machine simulator process, the virtual machine can adopt different processing methods for memory failures of different memory failure levels, minimize the use of memory isolation, and reduce the impact of memory isolation on the operation of the virtual machine.
[0076] As described above, in some possible implementation manners, virtual machines running on a host may be subject to various operations such as cold restart, warm restart, deletion, creation, and live migration.
[0077] After the virtual machine is subject to various operations such as cold restart, warm restart, deletion, creation, and live migration, the isolation of the memory hardware performed by the virtual machine will also disappear. To prevent the virtual machine from still accessing the memory hardware with a memory fault after being subject to various operations such as cold restart, warm restart, deletion, creation, and live migration, which may cause the application programs in the virtual machine to malfunction or even cause the host to crash, it is necessary to save the obtained memory fault information so that the virtual machine can still isolate the memory hardware with a memory fault after cold restart, warm restart, deletion, creation, and live migration of the virtual machine.
[0078] The obtained memory fault information can be saved as historical fault information, and this historical fault information will not be deleted due to cold restart, warm restart, deletion, creation, and live migration of the virtual machine. After cold restart, warm restart, deletion, creation, and live migration of the virtual machine, the virtual machine isolates the memory hardware with a memory fault according to this historical fault information.
[0079] In some possible implementation manners, the physical address of the faulty host, the fault process identifier, and the correspondence between the physical address of the faulty host and the fault process identifier are saved as historical fault information;
[0080] Among them, the fault process identifier is a user-mode process, that is, the process control character of the virtual machine emulator process corresponding to the faulty HPA obtained through reverse mapping.
[0081] Specifically, when sending the faulty HVA to the virtual machine emulator process, the faulty HPA and the corresponding virtual machine emulator process are written into the fault information file as historical fault information.
[0082] Each line of this fault information file is a piece of historical fault information, including two fields, HPA and PID. Among them, the faulty HPA is written into the HPA field, and the fault process identifier is written into the PID field.
[0083] Each time the fault handling method corresponding to the embodiment of the present disclosure is executed, the obtained faulty HPA and the corresponding virtual machine emulator process need to be written into the fault information file as historical fault information.
[0084] In some possible implementation manners, by perceiving the operations on the virtual machine, according to the operations on the virtual machine, the historical fault information can be modified or the memory hardware with a fault can be isolated again.
[0085] Figure 3The flowchart shows an implementation method for modifying historical fault information or re-isolating faulty memory hardware according to operations on virtual machines, as Figure 2 shown, modifying historical fault information or re-isolating faulty memory hardware according to operations on virtual machines may include step S310, step S320, and step S330.
[0086] In step S310, obtain virtual machine operations for one or more virtual machines running on a host and operation identifiers corresponding to the virtual machine operations;
[0087] In step S320, perform a search in the historical fault information based on the virtual machine operations and operation identifiers;
[0088] In step S330, based on the search result, modify the historical fault information and / or isolate the memory hardware of the virtual machine on which the virtual machine operation acts through the operation identifier;
[0089] Among them, the operation identifier is the process control character of the virtual machine emulator process corresponding to the virtual machine on which the virtual machine operation acts.
[0090] In some possible implementation manners, in step S310, by creating a character device and opening the created character device in Libvirt, the perception of operations on virtual machines can be realized, and the operation identifier corresponding to the virtual machine operation can be obtained.
[0091] Among them, Libvirt is an open-source toolset for managing virtualization platforms, which provides a unified interface to manage virtual machines. Therefore, operations on virtual machines can be perceived through Libvirt.
[0092] The operation identifier is the PID of the virtual machine emulator process corresponding to the virtual machine on which the virtual machine operation acts.
[0093] For example, in the case of virtual machine hot restart, since the virtual machine emulator process corresponding to the virtual machine does not change before and after the hot restart, the operation identifier is the PID of the virtual machine emulator process corresponding to the virtual machine that is hot restarted.
[0094] For virtual machine deletion, the operation identifier is the PID of the virtual machine emulator process corresponding to the deleted virtual machine.
[0095] For creating a virtual machine, the operation identifier is the PID of the virtual machine emulator process corresponding to the created virtual machine.
[0096] Virtual machine migration is equivalent to deleting the virtual machine from the host, and virtual machine cold restart is equivalent to first deleting the virtual machine and then creating a new virtual machine.
[0097] In some possible implementation manners, in step S320, retrieval is performed based on the PID field in the fault information file according to the operation identifier.
[0098] In some possible implementation manners, in step S330, based on the retrieval result in the PID field, when the virtual machine simulator process corresponding to the faulty HPA changes, the historical fault information is modified. When a new virtual machine uses the faulty memory hardware, the memory hardware needs to be re-isolated.
[0099] The processing of the historical fault information required for different virtual machine operations will be specifically introduced below.
[0100] Figure 4 The flowchart shows an implementation manner of processing the historical fault information or isolating the memory hardware when the virtual machine operation is a warm restart, as Figure 4 shown, which may include step 410 and step S420.
[0101] In step S410, when the virtual machine operation is a warm restart, retrieval is performed in the historical fault information based on the operation identifier;
[0102] In step S420, when the fault process identifier consistent with the operation identifier is retrieved in the historical fault information, the virtual address of the faulty host corresponding to the physical address of the faulty host is obtained, and the virtual address of the faulty host is sent to the virtual machine simulator process corresponding to the operation identifier.
[0103] In some possible implementation manners, in step S410, when the virtual machine operation is a warm restart, since the virtual machine simulator process corresponding to the virtual machine does not change before and after the warm restart, and the memory corresponding to the virtual machine does not change either, the historical fault information in the fault information file does not need to change.
[0104] However, after the virtual machine is warm restarted, the memory isolation performed by the virtual machine will also disappear. Therefore, the faulty memory hardware needs to be re-isolated.
[0105] Retrieval is performed in the historical fault information in the fault information file based on the operation identifier to obtain whether the virtual machine isolated the faulty memory hardware before the warm restart, that is, the memory corresponding to the virtual machine before the warm restart includes the faulty memory hardware.
[0106] In some possible implementation manners, in step S420, if a fault process identifier consistent with the operation identifier is retrieved from the historical fault information, it indicates whether the virtual machine isolated the faulty memory hardware before the warm restart. That is to say, the memory corresponding to the virtual machine before the warm restart includes the faulty memory hardware. Therefore, it is necessary to isolate the faulty memory hardware again.
[0107] Since the fault HPA is saved in the historical fault information, only the method described above needs to be used to obtain the fault HVA corresponding to the fault HPA, and the fault HVA is sent to the virtual machine simulator process corresponding to the operation identifier. The memory isolation is performed using the method described above. That is, the virtual machine simulator process corresponding to the operation identifier finds the fault GPA mapped by the fault HVA through the fault HVA, fills the virtual MCE register of the faulty virtual machine corresponding to the virtual machine simulator process with the fault GPA, and then injects an MCE interrupt or a CMCI interrupt into the faulty virtual machine to pass the fault GPA into the faulty virtual machine. After receiving the MCE interrupt or the CMC interrupt, the faulty virtual machine uses the Soft Offline function of the memory to isolate the memory corresponding to the fault GPA.
[0108] Figure 5 The flowchart shows an implementation manner of processing the historical fault information or isolating the memory hardware in the case where the virtual machine operation is to delete the virtual machine. As Figure 5 shown, it may include step 510 and step S520.
[0109] In step S510, in the case where the virtual machine operation is to delete the virtual machine, a search is performed in the historical fault information based on the operation identifier;
[0110] In step S520, in the case where a fault process identifier consistent with the operation identifier is retrieved from the historical fault information, the fault process identifier consistent with the operation identifier in the historical fault information is modified to an invalid value.
[0111] In some possible implementation manners, in step S510 and step S520, in the case where the virtual machine operation is to delete the virtual machine, since the virtual machine is deleted, the virtual machine simulator process corresponding to the virtual machine will naturally disappear. Therefore, it is necessary to delete the fault process identifier consistent with the operation identifier in the fault information file to prevent other processes from being assigned the operation identifier and being wrongly recorded as historical fault information.
[0112] However, since the memory hardware failure still exists, the faulty HPA in the historical fault information cannot be deleted to prevent other processes from using the faulty HPA. Therefore, it is only necessary to modify the faulty process identifier consistent with the operation identifier in the historical fault information to an invalid value, such as -1, to retain the faulty HPA on the basis of removing the corresponding relationship between the PID of the virtual machine simulator process of the deleted virtual machine and the faulty HPA.
[0113] Figure 5 A flowchart of an implementation method of processing historical fault information or isolating memory hardware when the virtual machine operation is to create a virtual machine is shown. Figure 6 As shown, it may include step S610, step S620, and step S630.
[0114] In step S610, when the virtual machine operation is to create a virtual machine, a physical address of a faulty host machine corresponding to a faulty process identifier having an invalid value in the historical fault information is obtained;
[0115] In step S620, the faulty host machine physical address retrieved and obtained is used to determine the faulty process control symbol of the user-mode process corresponding to the faulty host machine physical address retrieved and obtained through reverse mapping;
[0116] In step S630, when the fault process control symbol is consistent with the operation identifier, the fault process identifier corresponding to the physical address of the fault host machine retrieved from the historical fault information is modified into the operation identifier.
[0117] In some possible implementations, in step S610, when the virtual machine operation is to create a virtual machine, all rows whose PID fields are invalid values are searched in the fault information file to obtain the fault HPA of the HPA fields of these rows to obtain all unused faulty memory hardware.
[0118] In some possible implementations, in step S620, the acquired faulty HPA is reverse mapped using the above method to obtain a user-mode process corresponding to the faulty HPA, that is, a process of an application using the faulty memory hardware corresponding to the faulty HPA.
[0119] In some possible implementation manners, in step S630, if the user-mode process is consistent with the operation identifier, it indicates that the memory corresponding to the created virtual machine includes the faulty memory hardware. It is necessary to use the method described above to obtain the faulty HVA corresponding to the faulty HPA, send the faulty HVA to the virtual machine simulator process corresponding to the operation identifier, and perform memory isolation using the method described above. That is, the virtual machine simulator process corresponding to the operation identifier looks up the faulty GPA mapped by the faulty HVA through the faulty HVA, fills the virtual MCE register of the faulty virtual machine corresponding to the virtual machine simulator process with the faulty GPA, and then injects an MCE interrupt or a CMCI interrupt into the faulty virtual machine to pass the faulty GPA into the faulty virtual machine. After receiving the MCE interrupt or the CMC interrupt, the faulty virtual machine uses the Soft Offline function of the memory to isolate the memory corresponding to the faulty GPA.
[0120] Meanwhile, since a virtual machine has re-pinned the faulty memory hardware, it is necessary to update the historical fault information, that is, write the PID of the virtual machine simulator process corresponding to the created virtual machine to the PID in the row where the faulty HPA is located.
[0121] In some possible implementation manners, virtual machine migration is equivalent to deleting the virtual machine from the host. Therefore, virtual machine migration only needs to be processed according to Figure 5 the method described above.
[0122] Virtual machine cold restart is equivalent to first deleting the virtual machine and then creating a new virtual machine. Therefore, virtual machine cold restart only needs to be processed according to Figure 6 and Figure 5 the methods described above. Among them, the PID of the virtual machine simulator process corresponding to the virtual machine after cold restart is the PID of the virtual machine simulator process corresponding to the created virtual machine.
[0123] In some possible implementation manners, if the host restarts, the PIDs of the virtual machine simulator processes corresponding to all virtual machines will change. Therefore, it is necessary to modify all historical faults in the fault information file and perform memory isolation again.
[0124] Figure 7 shows a flowchart of an implementation manner of modifying the historical fault information and performing memory isolation again after the host restarts, as Figure 7 shown, which may include step S710, step S720, and step S730.
[0125] In step S710, it is detected that the host restarts. Before the virtual machines running on the host restart, the fault process identifier in the historical fault information is modified to an invalid value;
[0126] In step S720, after the virtual machine running on the host is restarted, according to the physical address of the faulty host in the historical fault information, through reverse mapping, the user-mode process corresponding to the physical address of the faulty host is determined.
[0127] In step S730, when the user-mode process is a virtual machine emulator process, the fault process identifier corresponding to the physical address of the faulty host in the historical fault information is modified to the process control character of the user-mode process.
[0128] In some possible implementation manners, in step S710, when the host is restarted but the virtual machine running on the host is not restarted, the fault process identifiers in all historical fault information are modified to invalid values, such as -1, to facilitate rewriting the virtual machine emulator process of the virtual machine into the historical fault information after the virtual machine is restarted.
[0129] Specifically, the modification can be made in / etc / rc.local to implement modifying the fault process identifiers in all historical fault information to invalid values.
[0130] In some possible implementation manners, in steps S720 and S730, after the virtual machine is restarted, since after the virtual machine is restarted, it is equivalent to creating a new virtual machine on the host, therefore, the restarted virtual machine needs to be regarded as a newly created virtual machine and processed using the above method.
[0131] Therefore, steps S720 and S730 are essentially repetitions of steps S620 and S630, and will not be elaborated here.
[0132] In the case where the memory hardware fault is cleared, such as when the faulty memory module of the host is replaced, all historical fault information will become invalid, so it is necessary to delete the fault information corresponding to the physical address of the faulty host of the memory hardware on the host to avoid affecting the operation of the host.
[0133] Based on the same principle as the method shown in Figure 1 shown, Figure 8 FIG. shows a schematic structural diagram of a fault processing apparatus provided by an embodiment of the present disclosure. As Figure 8 shown, the fault processing apparatus 80 may include:
[0134] An information collection module 810, configured to obtain memory fault information reported by the faulty memory hardware, where the memory fault information includes the physical address of the faulty host of the faulty memory hardware on the host;
[0135] The reverse mapping module 820 is used to determine the user-mode process corresponding to the physical address of the faulty host by reverse mapping according to the physical address of the faulty host; when the user-mode process is a virtual machine simulator process, the virtual address of the faulty host corresponding to the physical address of the faulty host is obtained according to the address information of the virtual memory area.
[0136] The fault injection module 830 is used to send the virtual address of the faulty host to the user-mode process for the user-mode process to isolate the memory hardware of the faulty virtual machine running on the host according to the virtual address of the faulty host.
[0137] In the fault handling device provided by the embodiments of the present disclosure, through the fault HPA of the faulty memory hardware, the virtual machine simulator process corresponding to the fault HPA and the corresponding fault HVA are obtained. By sending the fault HVA to the virtual machine simulator process, the memory fault is injected into the virtual machine, and the virtual machine isolates the faulty memory hardware, avoiding all the memory corresponding to the virtual machine from being pinned, preventing the virtual machine from accessing the faulty memory hardware, resulting in errors in the application programs in the virtual machine, or even causing the host to crash, and reducing the crash probability of the host.
[0138] It can be understood that the above-mentioned modules of the fault handling device in the embodiments of the present disclosure have the functions of implementing the corresponding steps of the fault handling method in the embodiments shown in Figure 1 This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above-mentioned modules can be implemented separately or integrated by multiple modules. For the function descriptions of the above-mentioned modules of the fault handling device, reference can specifically be made to the corresponding descriptions of the fault handling method in the embodiments shown in Figure 1 and will not be elaborated here.
[0139] In some possible implementation manners, the fault handling device further includes a fault information saving module, which is used to: save the physical address of the faulty host, the fault process identifier, and the correspondence between the physical address of the faulty host and the fault process identifier as historical fault information; where the fault process identifier is the process control character of the user-mode process.
[0140] In some possible implementation manners, the fault handling device further includes a fault information processing module, and the fault information processing module includes: an information acquisition unit, configured to acquire virtual machine operations for one or more virtual machines running on a host and operation identifiers corresponding to the virtual machine operations; a retrieval unit, configured to retrieve in historical fault information based on the virtual machine operations and the operation identifiers; an information modification unit, configured to modify the historical fault information based on the retrieval result and / or indicate, through the operation identifier, the virtual machine isolated memory hardware on which the virtual machine operation acts; wherein, the operation identifier is the process control identifier of the virtual machine simulator process corresponding to the virtual machine on which the virtual machine operation acts.
[0141] In some possible implementation manners, the retrieval unit is configured to: when the virtual machine operation is a hot restart, retrieve in the historical fault information based on the operation identifier; the information modification unit is configured to: when a fault process identifier consistent with the operation identifier is retrieved in the historical fault information, acquire the virtual address of the faulty host corresponding to the physical address of the faulty host, and send the virtual address of the faulty host to the virtual machine simulator process corresponding to the operation identifier.
[0142] In some possible implementation manners, the retrieval unit is configured to: when the virtual machine operation is to delete a virtual machine, retrieve in the historical fault information based on the operation identifier; the information modification unit is configured to: when a fault process identifier consistent with the operation identifier is retrieved in the historical fault information, modify the fault process identifier consistent with the operation identifier in the historical fault information to an invalid value.
[0143] In some possible implementation manners, the retrieval unit is configured to: when the virtual machine operation is to create a virtual machine, acquire the physical address of the faulty host corresponding to the fault process identifier with an invalid value in the historical fault information; the information modification unit is configured to: according to the physical address of the faulty host obtained by retrieval, determine, through reverse mapping, the fault process control identifier of the user-mode process corresponding to the physical address of the faulty host obtained by retrieval; when the fault process control identifier is consistent with the operation identifier, modify the fault process identifier corresponding to the physical address of the faulty host obtained by retrieval in the historical fault information to the operation identifier.
[0144] In some possible implementation manners, the fault handling device further includes a host restart module, configured to: when it is detected that the host restarts, modify the fault process identifier in the historical fault information to an invalid value; according to the physical address of the faulty host in the historical fault information, determine, through reverse mapping, the user-mode process corresponding to the physical address of the faulty host; when the user-mode process is a virtual machine simulator process, modify the fault process identifier corresponding to the physical address of the faulty host in the historical fault information to the process control identifier of the user-mode process.
[0145] In some possible implementations, the fault handling device further includes a fault clearing module, which is configured to: delete the historical fault information corresponding to the physical address of the faulty host of the memory hardware on the host when the memory hardware fault is cleared.
[0146] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0147] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0148] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0149] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the fault handling method provided in the embodiment of the present disclosure.
[0150] Compared with the prior art, the electronic device obtains the virtual machine simulator process corresponding to the fault HPA and the corresponding fault HVA of the faulty memory hardware through the fault HPA of the faulty memory hardware, and injects the memory fault into the virtual machine by sending the fault HVA to the virtual machine simulator process. The virtual machine isolates the faulty memory hardware, avoiding all the memory corresponding to the virtual machine from being pinned, preventing the virtual machine from accessing the faulty memory hardware, which may cause the application program in the virtual machine to malfunction and even cause the host to crash, thereby reducing the crashing probability of the host.
[0151] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the fault handling method provided in the embodiment of the present disclosure.
[0152] Compared with the prior art, the readable storage medium obtains the virtual machine simulator process corresponding to the fault HPA and the corresponding fault HVA of the faulty memory hardware through the fault HPA of the faulty memory hardware, and injects the memory fault into the virtual machine by sending the fault HVA to the virtual machine simulator process. The virtual machine isolates the faulty memory hardware, avoiding all the memory corresponding to the virtual machine from being pinned, preventing the virtual machine from accessing the faulty memory hardware, which may cause the application program in the virtual machine to malfunction and even cause the host to crash, thereby reducing the crashing probability of the host.
[0153] The computer program product includes a computer program which, when executed by a processor, implements the fault handling method provided by the embodiments of the present disclosure.
[0154] Compared with the prior art, the computer program product obtains the virtual machine emulator process corresponding to the fault HPA and the corresponding fault HVA through the fault HPA of the faulty memory hardware. By sending the fault HVA to the virtual machine emulator process, the memory fault is injected into the virtual machine, and the virtual machine isolates the faulty memory hardware, avoiding all the memory corresponding to the virtual machine from being pinned, preventing the virtual machine from accessing the faulty memory hardware, which may cause the application program in the virtual machine to go wrong and even cause the host machine to crash, thus reducing the crashing probability of the host machine.
[0155] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0156] As Figure 9 shown, the device 900 includes a computing unit 901 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0157] Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0158] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the fault handling method. For example, in some embodiments, the fault handling method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the fault handling method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the fault handling method by any other suitable means (e.g., by means of firmware).
[0159] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0160] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0162] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0163] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0164] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0165] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is imposed herein.
[0166] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A fault handling method, comprising: Obtaining memory fault information reported by the failed memory hardware, wherein the memory fault information includes a failed host machine physical address of the failed memory hardware in the host machine; Determine the user state process corresponding to the physical address of the faulty host machine through reverse mapping according to the physical address of the faulty host machine; In the case where the user state process is a virtual machine simulator process, obtaining a virtual address of the faulty host machine corresponding to the physical address of the faulty host machine according to address information of the virtual memory area; The faulty host machine virtual address is sent to the user state process, so that the user state process instructs the faulty virtual machine running on the host machine to isolate the memory hardware according to the faulty host machine virtual address.
2. The method according to claim 1, further comprising: The physical address of the faulty host machine, the faulty process identifier, and the corresponding relationship between the physical address of the faulty host machine and the faulty process identifier are saved as historical fault information; The fault process identifier is the process control symbol of the user state process.
3. The method according to claim 2, further comprising: Obtaining virtual machine operations for one or more virtual machines running on the host machine and operation identifiers corresponding to the virtual machine operations; Based on the virtual machine operation and the operation identifier, searching in the historical fault information; Based on the search result, modify the historical fault information and / or isolate the memory hardware of the virtual machine indicated by the operation identifier as the effect of the virtual machine operation; The operation identifier is a process control symbol of a virtual machine simulator process corresponding to the virtual machine on which the virtual machine operation acts.
4. The method according to claim 3, wherein: The searching in the historical fault information based on the virtual machine operation and the operation identifier includes: In a case where the virtual machine operation is a hot restart, searching the historical fault information based on the operation identifier; The modifying the historical fault information based on the search result and / or isolating the memory hardware of the virtual machine indicated by the operation identifier as the effect of the virtual machine operation includes: When a fault process identifier consistent with the operation identifier is retrieved from the historical fault information, a fault host machine virtual address corresponding to the fault host machine physical address is obtained, and the fault host machine virtual address is sent to the virtual machine simulator process corresponding to the operation identifier.
5. The method according to claim 3, wherein: The searching in the historical fault information based on the virtual machine operation and the operation identifier includes: In a case where the virtual machine operation is to delete the virtual machine, searching the historical fault information based on the operation identifier; The modifying the historical fault information based on the search result and / or isolating the memory hardware of the virtual machine indicated by the operation identifier as the effect of the virtual machine operation includes: When a fault process identifier that is consistent with the operation identifier is retrieved from the historical fault information, the fault process identifier that is consistent with the operation identifier in the historical fault information is modified to an invalid value.
6. The method according to claim 5, wherein: The searching in the historical fault information based on the virtual machine operation and the operation identifier includes: In a case where the virtual machine operation is to create a virtual machine, obtaining a physical address of a faulty host machine corresponding to a faulty process identifier having an invalid value in the historical fault information; The modifying the historical fault information based on the search result and / or isolating the memory hardware of the virtual machine indicated by the operation identifier as the effect of the virtual machine operation includes: Determine the fault process control symbol of the user state process corresponding to the retrieved fault host machine physical address through reverse mapping according to the retrieved fault host machine physical address; When the fault process control symbol is consistent with the operation identifier, the fault process identifier corresponding to the physical address of the fault host machine retrieved from the historical fault information is modified to the operation identifier.
7. The method according to claim 2, further comprising: Upon detecting that the host machine is restarted, before the virtual machine running on the host machine is restarted, modifying the fault process identifier in the historical fault information to an invalid value; After the virtual machine running on the host machine is restarted, the user state process corresponding to the physical address of the faulty host machine is determined by reverse mapping according to the physical address of the faulty host machine in the historical fault information; In the case where the user state process is a virtual machine simulator process, the fault process identifier corresponding to the fault host machine physical address in the historical fault information is modified to the process control symbol of the user state process.
8. The method according to claim 2, further comprising: When the memory hardware fault is cleared, the fault information corresponding to the faulty host machine physical address of the memory hardware in the host machine is deleted.
9. A fault handling device, comprising: An information collection module, used to obtain memory fault information reported by the failed memory hardware, wherein the memory fault information includes a failed host machine physical address of the failed memory hardware in the host machine; A reverse mapping module is used to determine the user state process corresponding to the physical address of the faulty host machine through reverse mapping according to the physical address of the faulty host machine; when the user state process is a virtual machine simulator process, obtain the virtual address of the faulty host machine corresponding to the physical address of the faulty host machine according to the address information of the virtual memory area; The fault injection module is used to send the fault host machine virtual address to the user state process, so that the user state process can instruct the faulty virtual machine running on the host machine to isolate the memory hardware according to the faulty host machine virtual address.
10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Memory fault information determination method and device
CN114860432A
Memory fault processing method, device, system and equipment and storage medium
CN116126581A
Virtual machine memory fault processing method and apparatus, and electronic device
CN117472632A
Virtual machine fault information storage method and device, storage medium and electronic equipment
CN118567779A