Virtual machine management method under fault scene of fault-tolerant server management platform

By directly using the host's KVM virtualization function when the fault-tolerant server management platform fails, decoupling the virtual machine and management platform, adjusting the configuration and starting the virtual machine, the business interruption caused by the fault-tolerant server management platform is solved, and rapid recovery is achieved.

CN120371455APending Publication Date: 2025-07-25SHANGHAI HI TECH CONTROL SYST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510437678.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When the existing fault-tolerant server management platform fails, the virtual machine cannot operate normally, causing business interruption. The existing solutions have problems such as high learning costs, long time or data loss.

Method used

By accessing the host, use the KVM management tool to determine the virtual machine status, and decouple the virtual machine from the fault-tolerant server management platform in the unrun state, adjust the configuration file, and directly use the host's KVM virtualization function to start the virtual machine to ensure its normal operation.

Benefits of technology

It realizes rapid business recovery in the event of fault-tolerant server management platform failure, reduces the risk of downtime, improves business continuity and reliability, and avoids in-depth repairs and data loss to the underlying management platform software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371455A_ABST
    Figure CN120371455A_ABST
Patent Text Reader

Abstract

The invention provides a virtual machine management method under a fault scene of a fault-tolerant server management platform. The virtual machine management method comprises the following steps: accessing a host machine running a virtual machine; whether the virtual machine is in a running state or not is determined through the KVM management tool, and when the virtual machine is not in the running state, the virtual machine and the fault-tolerant server management platform are decoupled; the virtual machine is started through the KVM management tool, and whether the virtual machine can work normally or not is confirmed. According to the method, under the condition that the fault-tolerant server management platform fails due to faults, a management mechanism of the fault-tolerant server management platform is bypassed, the target virtual machine is managed by directly utilizing the KVM virtualization function of the host machine, and normal operation of the target virtual machine is recovered, so that quick recovery after key business interruption is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of server fault tolerance, and more specifically, relates to a virtual machine management method in a fault scenario of a fault-tolerant server management platform. Background Art

[0002] A fault-tolerant server is a server cluster with automatic fault detection and switching functions. Its main function is that when a server hardware, software, or network failure occurs, the system can automatically detect and switch to a standby device. In recent years, more and more enterprises have chosen to use fault-tolerant servers to ensure the continuity and availability of their own services.

[0003] Existing fault-tolerant servers achieve automatic fault switching based on corresponding management platform software. However, when the fault-tolerant server management platform itself fails, the original automatic fault recovery mechanism may fail, resulting in all managed virtual machines being unable to run properly and causing service interruption.

[0004] Currently, for the problem of service interruption caused by the failure of the fault-tolerant server management platform, the solutions mainly include the following two:

[0005] Repair the underlying layer of the management platform software: This method locates and repairs the problem from the underlying layer of the management platform software. However, different versions of the management platform software and different fault causes require different operations, which not only have a high learning cost but also usually take 1 to 2 days, during which the service will continue to be interrupted;

[0006] Reinstall the management platform software and the virtual machine system: After reinstalling the management platform software and the virtual machine system, although the failure of the fault-tolerant server management platform can be eliminated and the service can be restored, this method will cause the loss of user data, and the service will also continue to be interrupted during the reinstallation. Summary of the Invention

[0007] In view of this, the present invention provides a virtual machine management method in a fault scenario of a fault-tolerant server management platform. The virtual machine management method includes:

[0008] Access the host machine running the virtual machine;

[0009] Determine whether the virtual machine is in a running state through a KVM management tool, and decouple the virtual machine from the fault-tolerant server management platform when the virtual machine is not in a running state;

[0010] Start the virtual machine through the KVM management tool and confirm whether the virtual machine can work properly.

[0011] Optionally, the access methods of the host machine include the following three:

[0012] Access the host through the remote management interface of the host;

[0013] Access the host by means of SSH remote login;

[0014] Access the host through the physical console.

[0015] Optionally, if the fault-tolerant server management platform is the Hydraulic fault-tolerant server management platform, decoupling the virtual machine from the fault-tolerant server management platform includes:

[0016] Enter the / var / opt / ft / ax / directory and modify the configuration file of the virtual machine. The modification operation includes:

[0017] Delete <metadata>Information;

[0018] Modify the type of the virtual machine from "pc-i440fx-2.X" to "pc";

[0019] Modify "qemu-unity" to "qemu-kvm";

[0020] Delete "cache='directsync'".

[0021] Optionally, if the version of the fault-tolerant server management platform is above 7.8, the modification operation further includes:

[0022] Delete the genid information.

[0023] Optionally, if the version of the fault-tolerant server management platform is 7.6, the modification operation further includes:

[0024] Modify the CPUmodel and configure all policies to be selected as disable.

[0025] Optionally, starting the virtual machine through the KVM management tool includes:

[0026] If the virtual machine fails to start, perform a predetermined operation, and the predetermined operation is to forcibly restart the virtual machine or restart the virtual machine in the case of destroying the current instance;

[0027] If the virtual machine still fails to start after performing the predetermined operation, check the system resources of the host to ensure that the virtual machine has the running conditions, and perform the predetermined operation again.

[0028] Optionally, the criteria for the virtual machine to work properly include:

[0029] The virtual machine is continuously in a running state;

[0030] The virtual machine can enter the user environment;

[0031] The connection between the virtual machine and the external network is normal.

[0032] The beneficial effects of the present invention are as follows:

[0033] The virtual machine management method in the fault scenario of the fault-tolerant server management platform of the present invention bypasses the management mechanism of the fault-tolerant server management platform and directly uses the KVM virtualization function of the host to manage the target virtual machine in the case of the failure of the fault-tolerant server management platform due to a fault, and restores the normal operation of the target virtual machine, so as to achieve a rapid recovery after the interruption of key services.

[0034] Other features and advantages of the present invention will be described in detail in the following detailed implementation section. Description of the Drawings

[0035] The present invention can be better understood by referring to the descriptions made in conjunction with the drawings in the following text, where the same or similar reference numerals are used in all the drawings to represent the same or similar components.

[0036] Figure 1 The implementation flowchart of the virtual machine management method in the fault scenario of the fault-tolerant server management platform according to an embodiment of the present invention is shown. Detailed Implementation Modes

[0037] In order to enable those skilled in the art to more fully understand the technical solution of the present invention, the exemplary implementation modes of the present invention will be described more comprehensively and in detail in the following text in conjunction with the drawings. Obviously, one or more of the implementation modes of the present invention described below are merely one or more of the specific ways to implement the technical solution of the present invention, and are not exhaustive. It should be understood that other ways belonging to a general inventive concept can be used to implement the technical solution of the present invention, and should not be limited by the exemplary implementation modes described. Based on one or more implementation modes of the present invention, all other implementation modes obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0038] Embodiment: When a fault occurs in the fault-tolerant server management platform, the virtual machines managed by it may be in an unknown state and unable to automatically perform fault switching. In this case, through the underlying management function of KVM, the administrator can manually intervene and directly start or restart the virtual machines on the host, enabling them to resume operation without relying on the fault-tolerant server management platform. The solution for manually starting the virtual machines depends on the KVM management tool libvirt on the host, which allows the administrator to directly control the running state of the virtual machines, including starting, shutting down, restarting, and status monitoring, etc.

[0039] As a Linux kernel-level virtualization technology, KVM relies on the KVM management tool libvirt for virtual machine management. The configuration, status management, and running control of the virtual machines can all be independently operated through the underlying command-line tools. Therefore, even if the fault-tolerant server management platform fails, the administrator can still access the KVM components through the host and perform operations such as virtual machine status checking, configuration adjustment, and manual startup to restore the business functions.

[0040] Figure 1 The implementation flowchart of the virtual machine management method in the fault scenario of the fault-tolerant server management platform according to an embodiment of the present invention is shown. Refer to Figure 1 , the virtual machine management method in the fault scenario of the fault-tolerant server management platform according to the embodiments of the present invention includes the following steps:

[0041] Step S100, access the host machine running the virtual machine;

[0042] Step S200, determine whether the virtual machine is in a running state through the KVM management tool, and decouple the virtual machine from the fault-tolerant server management platform when the virtual machine is not in a running state;

[0043] Step S300, start the virtual machine through the KVM management tool, and confirm whether the virtual machine can work properly.

[0044] Further, in step S100 of the embodiments of the present invention, the access methods to the host machine include the following three types:

[0045] Access the host machine through the remote management interface of the host machine;

[0046] Access the host machine through SSH remote login;

[0047] Access the host machine through the physical console.

[0048] Specifically, in the embodiments of the present invention, when the fault-tolerant server management platform is unavailable, it is necessary to directly access the host machine running the virtual machine to obtain the control authority over the virtual machine. Common access methods include:

[0049] Remote management interface: Servers are usually equipped with management interfaces based on IPMI (such as iDRAC, iLO, BMC). Even if the operating system or network is unavailable, the underlying console can still be accessed through this method;

[0050] SSH remote login: If the server network is still available, the host machine can be connected through a remote terminal to perform management operations;

[0051] Physical console: In a data center environment, administrators can directly use the server terminal for operations to ensure access to the system even in extreme cases.

[0052] Still further, in the embodiments of the present invention, the fault-tolerant server management platform is the Hyd fault-tolerant server management platform.

[0053] Specifically, in the embodiments of the present invention, after entering the host machine, it is necessary to check the current running state of the target virtual machine to determine subsequent recovery operations. Generally, the virtual machine may be in the following three states:

[0054] Running: The virtual machine is in a normal running state and no further operation is required.

[0055] Paused: It may be forcibly paused due to system failures or abnormal management of the Fault Tolerance Server Management Platform. In this case, you can try to manually resume operation.

[0056] Shutoff: It indicates that the virtual machine is not running. It may be abnormally shut down or fail to complete automatic startup, and manual startup is required. If the virtual machine is in a paused state, the administrator can directly try to resume operation; if the virtual machine is already shut down, the startup operation needs to be executed and it is necessary to check whether it can enter the system normally.

[0057] Furthermore, in step S200 of the embodiment of the present invention, decoupling the virtual machine from the Fault Tolerance Server Management Platform includes:

[0058] Enter the / var / opt / ft / ax / directory and modify the configuration file of the virtual machine. The modification operations include:

[0059] Delete <metadata>Information;

[0060] Modify the type of the virtual machine from "pc-i440fx-2.X" to "pc";

[0061] Modify "qemu-unity" to "qemu-kvm";

[0062] Delete "cache='directsync'".

[0063] Furthermore, in the embodiment of the present invention, if the version of the Hydra fault-tolerant server management platform is above 7.8, the modification operation on the virtual machine configuration file further includes:

[0064] Delete the genid information.

[0065] Furthermore, in the embodiment of the present invention, if the version of the Hydra fault-tolerant server management platform is 7.6, the modification operation on the virtual machine configuration file further includes:

[0066] Modify the CPUmodel and configure all policies to be disabled.

[0067] Specifically, in the embodiment of the present invention, in some cases, even if the virtual machine is manually attempted to be started, it may still not run properly due to the relevant dependency configurations of the Hydra fault-tolerant server management platform. Therefore, it is necessary to adjust the virtual machine configuration file to ensure that it can be started independently of the Hydra fault-tolerant server management platform. The adjustment of the virtual machine configuration file includes the following items:

[0068] Delete <metadata>Information, avoiding virtual machines depending on the Hydraulic Fault Tolerant Server Management Platform;

[0069] Modify the type of the virtual machine from "pc-i440fx-2.X" to "pc";

[0070] Modify "qemu-unity" to "qemu-kvm" to ensure running with standard KVM;

[0071] Delete "cache='directsync'" to avoid restricted storage access;

[0072] If the version of the Hydraulic Fault Tolerant Server Management Platform is above 7.8, the genid information needs to be deleted, otherwise the creation of the domain will fail due to the verification of genid during the subsequent KVM startup process;

[0073] If the version of the Hydraulic Fault Tolerant Server Management Platform is 7.6, the CPUmodel needs to be modified, and all policy selections should be configured as disable. The purpose of this step is to decouple the network port from the Hydraulic Fault Tolerant Server Management Platform.

[0074] Furthermore, in step S300 of the embodiment of the present invention, starting the virtual machine through the KVM management tool includes:

[0075] If the virtual machine cannot be started, perform a predetermined operation, and the predetermined operation is to forcibly restart the virtual machine or restart the virtual machine in the case of destroying the current instance;

[0076] If the virtual machine still cannot be started after performing the predetermined operation, check the system resources of the host to ensure that the virtual machine has the running conditions, and perform the predetermined operation again.

[0077] Specifically, in the embodiment of the present invention, after completing the adjustment of the virtual machine configuration file, the administrator can manually start the virtual machine through the KVM management tool libvirt. Generally, the virtual machine should be able to resume normal operation, but if it still cannot be started, a forced restart needs to be performed, or the current instance needs to be destroyed and restarted to ensure that the virtual machine enters an available state. In addition, in some special cases, such as host resource contention, abnormal storage access, etc., it is necessary to check the resources such as CPU, memory, and disk to ensure that the virtual machine has sufficient running conditions.

[0078] Furthermore, in step S300 of the embodiment of the present invention, the criteria for the virtual machine to work properly include:

[0079] The virtual machine is continuously in a running state;

[0080] The virtual machine can enter the user environment;

[0081] The connection between the virtual machine and the external network is normal.

[0082] Specifically, in the embodiment of the present invention, after starting or restarting the virtual machine, it is necessary to confirm the status to ensure that the business system returns to normal. The main steps include:

[0083] Check the status of the virtual machine to ensure that it is continuously running and does not automatically stop or crash due to abnormalities;

[0084] Access the virtual machine console and observe the startup log to ensure that the system can successfully enter the user environment.

[0085] Test network connectivity to confirm whether the virtual machine is connected to the external network normally to eliminate business unavailability caused by network abnormalities.

[0086] The core of the virtual machine management method under the fault-tolerant server management platform failure scenario of the embodiment of the present invention is:

[0087] Directly access the physical host, bypass the fault-tolerant server management platform, and enter the server environment locally or remotely;

[0088] Check the current state of the virtual machine to determine whether it is running, paused, or powered off to decide the appropriate recovery action.

[0089] Adjust the virtual machine configuration file and unbind it from the fault-tolerant server management platform to ensure KVM compatibility and avoid startup failures caused by dependencies;

[0090] Manually start or restart the virtual machine, and after correcting the configuration, directly restore its running state through the KVM underlying mechanism;

[0091] Verify the business system recovery status to ensure that all business services can operate normally after the virtual machine is started.

[0092] The virtual machine management method in a fault-tolerant server management platform failure scenario according to an embodiment of the present invention has the following beneficial effects:

[0093] Improve business continuity and reliability: Compared with the traditional fault recovery solution that relies on the fault-tolerant server management platform itself, this method provides additional business recovery capabilities and greatly reduces the risk of downtime;

[0094] The fault-tolerant server management platform and the virtual machine system can be decoupled: when the management layer of the fault-tolerant server management platform fails, the business virtual machine can be forced to start directly from the physical machine layer to ensure that the business is not interrupted;

[0095] Strong compatibility: Maintenance personnel do not need to have an in-depth understanding of the underlying architecture of the fault-tolerant server management platform software. Moreover, the KVM startup method does not change the operation mode due to the failure of the fault-tolerant server management platform, does not depend on specific hardware or software versions, and is easy to deploy.

[0096] Fast repair time: The reasons for the failure of the fault-tolerant server management platform are often complex, taking a long time and having high technical requirements to repair. However, the KVM method of starting virtual machines is simple, fast, and standard unified. Although one or more embodiments of the present invention have been described above, those of ordinary skill in the art should be aware that the present invention can be implemented in any other form without departing from its gist and scope. Therefore, the embodiments described above are illustrative rather than restrictive, and many modifications and substitutions are obvious to those of ordinary skill in the art without departing from the spirit and scope of the present invention as defined by the appended claims.< / metadata> < / metadata> < / metadata>

Claims

1. A method for virtual machine management in a fault scenario of a fault-tolerant server management platform, characterized in that Including: Access the host machine running the virtual machine; Determine whether the virtual machine is in a running state through the KVM management tool, and decouple the virtual machine from the fault-tolerant server management platform when the virtual machine is not in a running state; Start the virtual machine through the KVM management tool and confirm whether the virtual machine can work properly.

2. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 1, characterized in that The access methods of the host machine include the following three types: Access the host machine through the remote management interface of the host machine; Access the host machine through SSH remote login; Access the host machine through the physical console.

3. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 1, wherein If the fault-tolerant server management platform is the Hydraulic fault-tolerant server management platform, decoupling the virtual machine from the fault-tolerant server management platform includes: Enter the / var / opt / ft / ax / directory and perform modification operations on the configuration file of the virtual machine. The modification operations include: Delete <metadata>Information;< / metadata> Modify the type of the virtual machine from "pc-i440fx-2.X" to "pc"; Modify "qemu-unity" to "qemu-kvm"; Delete "cache='directsync'".

4. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 3, characterized in that, If the version of the fault-tolerant server management platform is above 7.8, the modification operations further include: Delete the genid information.

5. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 4, characterized in that, If the version of the fault-tolerant server management platform is 7.6, the modification operations further include: Modify the CPU model and configure all policy images to be disabled.

6. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 1, characterized in that Starting the virtual machine through the KVM management tool includes: If the virtual machine cannot be started, perform a predetermined operation, which is to forcibly restart the virtual machine or restart the virtual machine in the case of destroying the current instance; If the virtual machine still cannot be started after performing the predetermined operation, check the system resources of the host machine to ensure that the virtual machine has the running conditions, and perform the predetermined operation again.

7. The virtual machine management method in the fault scenario of the fault-tolerant server management platform according to claim 1, characterized in that The criteria for the virtual machine to work properly include: The virtual machine is continuously in a running state; The virtual machine can enter the user environment; The connection between the virtual machine and the external network is normal.