Virtual machine high availability method and device, equipment and storage medium
By building an alarm monitoring system and a fault handling mechanism based on preset evaluation sequence, the problem of strong dependence of the high availability method of virtual machines on SaltStack is solved, high availability and security are improved, multi-scenario refinement processing is supported, and operation and maintenance efficiency is enhanced.
Patent Information
- Application Number
- CN202510223620.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
The existing high availability methods of virtual machines have strong dependence on SaltStack, which leads to the inability to handle node exceptions, and the operation is complex and does not support multi-scenario refinement processing, which has security and efficiency problems.
By using telegraf and preset collectors to build an alarm monitoring system, monitor the SNMP trap alarms, operating system resources and service status and network status of the computing node where the virtual machine is located, generate alarm information, and troubleshoot based on the preset evaluation sequence and business priority to avoid strong dependence on SaltStack.
It improves the efficiency and security of the highly available virtual machine methods, reduces manual operation errors, supports multi-scenario refinement processing, enhances comprehensive monitoring and management of the computing nodes where the virtual machine is located, and improves operation and maintenance efficiency and system reliability.
Smart Images

Figure CN120104255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing technology, and in particular to a method, device, equipment and storage medium for high availability of a virtual machine. Background Art
[0002] Although many cloud vendors have provided virtual machine high availability solutions, due to closed source reasons, no efficient and secure virtual machine high availability methods for different scenarios have been obtained from public solutions. In addition, the open source OpenStack Masakari (i.e., a high availability and disaster recovery solution for open cloud platforms) solution has a variety of security and efficiency issues: the operation is complex, involving a large number of manual operations, and manual operations are easy to miss; it does not support detailed processing of multiple scenarios, and there is no efficiency and security; there are few processing methods, and there are security risks that affect customers; the use of the disable computing service method to allow evacuation may cause disk read and write abnormalities; it is unable to automatically receive alarms, locate the cause of the fault, and is not easy to trace back and troubleshoot, and it is impossible to perform fine-grained processing for different storage types, and there is an incompatibility risk; the community is not active, and various security vulnerabilities and user experience have not been repaired and improved. There are virtual machine high availability methods in the existing technology, but this method is strongly dependent on SaltStack (i.e., a management tool). When the node is shut down, stuck or restarted, salt (i.e., SaltStack's configuration management and remote execution framework) cannot be used, which will cause all scenarios to be unable to be processed.
[0003] As can be seen from the above, how to avoid strong dependence on Salt and improve the efficiency and security of the virtual machine high availability method is an urgent problem to be solved. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a virtual machine high availability method, device, equipment and storage medium, which can avoid strong dependence on salt and improve the efficiency and security of the virtual machine high availability method. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a virtual machine high availability method, comprising:
[0006] An alarm monitoring system is constructed using telegraf and a preset collector. The SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored through the alarm monitoring system, and alarm information is generated based on the preset alarm type and monitoring results; the preset collector is a collector determined based on the Consul cluster;
[0007] Determine whether the target computing node corresponding to the alarm information is in a preset blacklist;
[0008] If the target computing node corresponding to the alarm information is not in the preset blacklist, the corresponding fault scenario is determined based on the preset evaluation order and the alarm information, so as to use the fault scenario, service priority and preset virtual machine evacuation switch to handle the fault; the fault scenario includes node shutdown, node stuck, node restart, network failure and hardware failure;
[0009] If the target computing node corresponding to the alarm information is in a preset blacklist, the computing node is determined as a node under maintenance, and triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information is prohibited.
[0010] Optionally, the alarm monitoring system is constructed by using telegraf and a preset collector, and the SNMP trap alarm of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored by the alarm monitoring system, and alarm information is generated based on the preset alarm type and the monitoring result, including:
[0011] Allocate the network cards used by the virtual machine for communication of the control network, the business network, and the storage network to different logical interfaces through the bridge technology, and after the allocation operation is completed, deploy the first Consul agent, the second Consul agent, and the third Consul agent corresponding to the control network, the business network, and the storage network based on the computing node to obtain the first Consul cluster, the second Consul cluster, and the third Consul cluster;
[0012] Building a preset collector based on the first Consul cluster, the second Consul cluster, and the third Consul cluster;
[0013] Using telegraf to monitor the SNMP trap alarm of the computing node where the virtual machine is located and the resources and services inside the operating system to obtain a first monitoring result;
[0014] Using a preset collector and based on a preset monitoring frequency, the network status of the computing node is monitored to obtain a second monitoring result;
[0015] Determine a target monitoring result that meets the preset alarm type based on the first monitoring result and the second monitoring result, and generate corresponding alarm information based on the target monitoring result;
[0016] The first Consul cluster, the second Consul cluster and the third Consul cluster communicate through a rumor protocol.
[0017] Optionally, the determining a corresponding fault scenario based on a preset evaluation order and the alarm information, so as to perform fault processing by using the fault scenario, service priority, and a preset virtual machine evacuation switch, includes:
[0018] If the alarm information indicates that the network states of the control network, the service network and the storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, the power supply state of the target computing node is checked through the intelligent platform management interface;
[0019] If the inspection result indicates that the power state of the target computing node is off, determining that the fault scenario corresponding to the alarm information is node shutdown, and performing a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority;
[0020] If the inspection result indicates that the power state of the target computing node is not off, performing a first network ping detection on each computing node based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the business network, and the storage network, respectively, to obtain a first network detection result;
[0021] Determine a first auxiliary node that meets a first preset packet loss rate by using the first network detection result, and perform a first node ping detection on the first auxiliary node and the target computing node based on a first preset detection round number to obtain a first node detection result;
[0022] If the detection result of the first node is that the packet loss rate corresponding to each network card on the target computing node meets the second preset packet loss rate and there is no recovery of the network status, it is determined that the fault scenario corresponding to the alarm information is a node stuck, and the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority.
[0023] Optionally, the performing the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority includes:
[0024] Determine whether the preset virtual machine evacuation switch is in an on state;
[0025] If the preset virtual machine evacuation switch is in the on state, determining whether the business resources meet the preset resource requirements;
[0026] If the service resources do not meet the preset resource requirements, the virtual machine evacuation operation on the target computing node is performed based on the service priority order from high to low.
[0027] Optionally, the determining a corresponding fault scenario based on a preset evaluation order and the alarm information, so as to perform fault processing by using the fault scenario, service priority, and a preset virtual machine evacuation switch, includes:
[0028] If the alarm information indicates that the network states corresponding to the control network, the business network and the storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, a second network ping test is performed on each computing node based on the first gateway, the second gateway and the third gateway corresponding to the control network, the business network and the storage network, respectively, to obtain a second network test result;
[0029] Determine a second auxiliary node that meets a second preset packet loss rate by using the second network detection result, and perform a second node ping detection on the second auxiliary node and the target computing node based on a second preset detection round number to obtain a second node detection result;
[0030] If the second node detection result indicates that there is a recovery of the network status, it is determined that the fault scenario corresponding to the alarm information is a node restart, and it is determined through the intelligent platform management interface whether the hardware status of the target computing node has a restart or a memory failure;
[0031] If the hardware status of the target computing node has a restart or memory failure, determining whether the software service status of the target computing node is in a running state;
[0032] If the software service state of the target computing node is in the running state, determining that the fault scenario corresponding to the alarm information is a node restart, and performing a hot migration operation of the virtual machine on the target computing node;
[0033] If the second node detection result indicates that there is no recovery of the network status, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the service priority.
[0034] Optionally, the determining a corresponding fault scenario based on a preset evaluation order and the alarm information, so as to perform fault processing by using the fault scenario, service priority, and a preset virtual machine evacuation switch, includes:
[0035] If the alarm information indicates that any network status of the control network, business network and storage network of the target computing node corresponding to the alarm information is abnormal within a preset time period, then the corresponding target fault network is obtained;
[0036] Performing a third network ping detection on each of the computing nodes based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the business network, and the storage network, respectively, to obtain a third network detection result;
[0037] Determine a third auxiliary node that meets a third preset packet loss rate by using the third network detection result, and perform a third node ping detection on the third auxiliary node and the target computing node based on a third preset detection round number to obtain a third node detection result;
[0038] Based on the detection result of the third node, the network status and hot migration support of the target fault network are checked, and if the check result indicates that the network status of the target fault network is abnormal and hot migration is not supported, it is determined whether the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition;
[0039] If the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition, shut down the physical machine corresponding to the target computing node, and then perform the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority;
[0040] If the inspection result indicates that the network status of the target faulty network is abnormal and supports hot migration, a hot migration operation of the virtual machine on the target computing node is performed.
[0041] Optionally, the determining a corresponding fault scenario based on a preset evaluation order and the alarm information, so as to perform fault processing by using the fault scenario, service priority, and a preset virtual machine evacuation switch, includes:
[0042] If the alarm information indicates that the physical machine on which the virtual machine on the corresponding target computing node depends has a fault of one of CPU, power supply, motherboard, memory and hard disk failure, then it is determined that the fault scenario corresponding to the alarm information is a hardware failure;
[0043] If the hardware fault is a memory fault, determining whether the memory fault is a UE error;
[0044] If the memory fault is a UE error, determining whether the software service state of the target computing node is a running state;
[0045] If the software service state of the target computing node is in the running state, determining that the fault scenario corresponding to the alarm information is a node restart, and performing a hot migration operation of the virtual machine on the target computing node;
[0046] If the hardware failure is a hard disk failure, a corresponding disk replacement operation is performed.
[0047] In a second aspect, the present application provides a virtual machine high availability device, comprising:
[0048] An alarm information generation module is used to build an alarm monitoring system using telegraf and a preset collector, monitor the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node through the alarm monitoring system, and generate alarm information based on the preset alarm type and monitoring results; the preset collector is a collector determined based on the Consul cluster;
[0049] A target computing node determination module is used to determine whether the target computing node corresponding to the alarm information is in a preset blacklist;
[0050] A fault handling module, for determining a corresponding fault scenario based on a preset evaluation order and the alarm information if the target computing node corresponding to the alarm information is not in a preset blacklist, so as to perform fault handling using the fault scenario, service priority, and a preset virtual machine evacuation switch; the fault scenario includes node shutdown, node freeze, node restart, network failure, and hardware failure;
[0051] The node under maintenance determination module is used to determine the computing node as a node under maintenance if the target computing node corresponding to the alarm information is in a preset blacklist, and prohibit triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information.
[0052] In a third aspect, the present application provides an electronic device, including:
[0053] Memory, used to store computer programs;
[0054] The processor is used to execute the computer program to implement the aforementioned virtual machine high availability method.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned virtual machine high availability method when executed by a processor.
[0056] This application uses telegraf and a preset collector to build an alarm monitoring system, through which the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored, and alarm information is generated based on the preset alarm type and the monitoring results; the preset collector is a collector determined based on the Consul cluster; it is determined whether the target computing node corresponding to the alarm information is in the preset blacklist; if the target computing node corresponding to the alarm information is not in the preset blacklist, the corresponding fault scenario is determined based on the preset evaluation order and the alarm information, so as to use the fault scenario, business priority and preset virtual machine evacuation switch for fault processing; the fault scenarios include node shutdown, node stuck, node restart, network failure and hardware failure; if the target computing node corresponding to the alarm information is in the preset blacklist, the computing node is determined as a node under maintenance, and the operation of triggering the corresponding fault scenario determined based on the alarm information is prohibited.
[0057] As can be seen from the above, this application builds an alarm monitoring system based on telegraf and in combination with a preset collector. The alarm monitoring system monitors the SNMP trap alarms of the computing nodes where the virtual machines are located, the resources and services inside the operating system, and the network status of the computing nodes to generate alarm information, and then determines whether the target computing node corresponding to the alarm information is in the preset blacklist. If it is not in the blacklist, the corresponding fault scenarios are determined based on the preset evaluation order and the alarm information, such as node shutdown, node stuck, and other fault scenarios. In this way, fault handling is performed according to business priority and the preset virtual machine evacuation switch to ensure that key businesses are restored first when resources are limited. Through a unified alarm monitoring system and alarm processing process, comprehensive monitoring and management of the computing nodes where the virtual machines are located is achieved, improving operation and maintenance efficiency and system reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0059] Figure 1 A flow chart of a virtual machine high availability method disclosed in this application;
[0060] Figure 2 A schematic diagram of a high availability device structure of a virtual machine disclosed in this application;
[0061] Figure 3This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] At present, the OpenStack Masakari solution has problems such as complex operation, lack of support for multi-scenario detailed processing, few processing methods and security risks, may cause disk read and write anomalies, difficult fault location, incompatibility with different storage types, and inactive community. The existing virtual machine high availability method is highly dependent on SaltStack and cannot handle node abnormalities. To this end, the present application provides a virtual machine high availability method, which handles faults according to business priorities and preset virtual machine evacuation switches to ensure that key businesses are restored first when resources are limited. Through a unified alarm monitoring system and alarm processing process, comprehensive monitoring and management of the computing nodes where the virtual machines are located are achieved, thereby improving operation and maintenance efficiency and system reliability.
[0064] See also Figure 1 As shown, an embodiment of the present invention discloses a method for high availability of a virtual machine, comprising:
[0065] Step S11, using telegraf and a preset collector to build an alarm monitoring system, through which the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored, and alarm information is generated based on the preset alarm type and the monitoring results; the preset collector is a collector determined based on the Consul cluster.
[0066] In this embodiment, the network card used by the virtual machine for communication of the control network, business network and storage network is allocated to different logical interfaces through network bridging technology, that is, each virtual machine can have multiple network interfaces, each interface is connected to a different network; based on the control network, business network and storage network on the computing node, a first Consul (i.e., a service network solution) agent, a second Consul agent and a third Consul agent are deployed respectively to build a first Consul cluster based on the control network, a second Consul cluster based on the business network and a third Consul cluster based on the storage network; the nodes of each cluster communicate through the rumor protocol; telegraf is used to monitor the SNMP (Simple Network Management Protocol) trap alarm of the computing node where the virtual machine is located and the resources and services inside the operating system, and the preset collector constructed based on the above-mentioned Consul cluster is used to monitor the network status of the computing node to obtain the monitoring results. The alarm information generated based on the monitoring results may include the alarm type, alarm level, occurrence time, and information about the relevant virtual machines and computing nodes.
[0067] Specifically, the alarm monitoring system is constructed by using telegraf and a preset collector, and the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored through the alarm monitoring system, and alarm information is generated based on the preset alarm type and the monitoring results, including: using the bridge technology to allocate the network card used by the virtual machine for communication with the control network, the business network, and the storage network to different logical interfaces, and after the allocation operation is completed, deploying the first Consul agent, the second Consul agent, and the third Consul agent corresponding to the control network, the business network, and the storage network based on the computing node to obtain the first Consul cluster, the second Consul cluster, and the third Consul cluster. l cluster; construct a preset collector based on the first Consul cluster, the second Consul cluster and the third Consul cluster; use telegraf to monitor the SNMP trap alarm of the computing node where the virtual machine is located and the resources and services inside the operating system to obtain a first monitoring result; use the preset collector and monitor the network status of the computing node based on a preset monitoring frequency to obtain a second monitoring result; determine a target monitoring result that meets the preset alarm type based on the first monitoring result and the second monitoring result, and generate corresponding alarm information based on the target monitoring result; wherein, the first Consul cluster, the second Consul cluster and the third Consul cluster communicate through a rumor protocol.
[0068] It is understandable that the network status of the computing node is monitored based on the preset monitoring frequency and using each Consul cluster in the preset collector. In a specific embodiment, the Consul cluster detects the network status of the control network, business network and storage network of each computing node once every 15 seconds. If any of the three networks of a computing node is judged to be unreachable or communication fails in 3 or more tests in 1 minute, that is, four tests, it means that there is a network problem with the computing node, and a corresponding alarm message is generated. The preset monitoring frequency can be adjusted according to the actual situation and is not specifically limited here. It should be noted that the preset alarm types include KubernetesNodeNotReady (abnormal node status), NovaServiceDown (computing service is unavailable or fails), PrometheusTargetDown (in the Prometheus monitoring system, one or more monitored targets are currently unreachable or unable to collect data), BmcTrap (alarm or event notification sent by the baseboard management controller of the server or device through SNMP), ConsulNetworkFailed (abnormal network indicators or network connection failure), ConsulAbnormalShutdown (the state of the Consul cluster or agent process suddenly shutting down or terminating under unexpected or abnormal circumstances), and SubpathStateWarning (path status alarm when the virtual machine uses centralized storage).
[0069] Step S12: Determine whether the target computing node corresponding to the alarm information is in a preset blacklist.
[0070] In this embodiment, after receiving the alarm information, the target computing node corresponding to the alarm information is determined based on the alarm information, and it is determined whether the target computing node is in the preset blacklist; the preset blacklist includes nodes that are known to have specific problems and are undergoing maintenance, such as shutdown maintenance. In addition, the historical alarm information and the behavior patterns of the computing nodes can be analyzed based on a preset time period, so that the computing nodes that frequently generate false alarms or do not need to be processed can be added to the preset blacklist using the analysis results; at the same time, for each computing node in the preset blacklist, if it is monitored that the computing node has returned to normal or maintenance is completed, it will be removed from the preset blacklist.
[0071] Step S13: If the target computing node corresponding to the alarm information is not in the preset blacklist, the corresponding fault scenario is determined based on the preset evaluation order and the alarm information, so as to perform fault processing using the fault scenario, business priority and preset virtual machine evacuation switch; the fault scenarios include node shutdown, node freezing, node restart, network failure and hardware failure.
[0072] In this embodiment, if the target computing node corresponding to the alarm information is not in the preset blacklist, the fault scenarios are evaluated in the order of node shutdown, node freeze, node restart, network failure, and hardware failure, and the fault is handled based on the fault scenario determined by the alarm information. If the fault scenario is node shutdown, the alarm information indicates that the network status of the control network, business network, and storage network corresponding to the target computing node corresponding to the alarm information are abnormal within a preset time period, and the inspection result obtained by checking the power status of the target computing node through the intelligent platform management interface indicates that the power is in the off state. At this time, the virtual machine evacuation operation on the target computing node is performed for the node shutdown fault scenario; if the inspection result indicates that the power is not in the off state, it is determined whether the fault scenario is a node freeze, based on the first gateway, the second network corresponding to the control network, the business network, and the storage network, respectively. The first auxiliary node and the target computing node perform a first network ping detection on each of the computing nodes, and screen out a first auxiliary node with a ping packet loss rate of 0 based on the first network detection result, so as to use the first auxiliary node and the target computing node to perform a first node ping detection. If the first node monitoring result indicates that the packet loss rates corresponding to the network cards on the target computing node are all 100% and there is no recovery of the network status, it is determined that the fault scenario is a node stuck, and the physical machine corresponding to the target computing node needs to be shut down, and the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority.
[0073] Specifically, the method of determining the corresponding fault scenario based on the preset evaluation order and the alarm information, and performing fault processing by using the fault scenario, business priority, and a preset virtual machine evacuation switch, includes: if the alarm information indicates that the network states corresponding to the control network, business network, and storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, then checking the power state of the target computing node through the intelligent platform management interface; if the inspection result indicates that the power state of the target computing node is off, then determining that the fault scenario corresponding to the alarm information is node shutdown, and performing a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the business priority; if the inspection result indicates that the power state of the target computing node is not off, then based on the control network, the business network, and the storage network corresponding to the target computing node, the power state of the target computing node is turned off, and then performing a virtual machine evacuation operation on the target computing node. The first gateway, the second gateway and the third gateway corresponding to the service network and the storage network respectively perform a first network ping test on each of the computing nodes to obtain a first network detection result; the first network detection result is used to determine the first auxiliary node that meets the first preset packet loss rate, and based on the first preset detection round number, the first node ping test is performed on the first auxiliary node and the target computing node to obtain the first node detection result; if the first node detection result is that the packet loss rate corresponding to each network card on the target computing node meets the second preset packet loss rate and there is no recovery of the network state, then it is determined that the fault scenario corresponding to the alarm information is a node stuck, and the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the service priority. It is worth mentioning that the first preset detection round number and the number of packets for ping detection default to two rounds of ping detection and ping50 packets, which can also be adjusted accordingly according to actual conditions.
[0074] It can be understood that the first auxiliary node with a ping packet loss rate of 0 is screened out based on the first network detection result. In this solution, the first auxiliary node is two nodes with a ping packet loss rate of 0 in the first gateway, the second gateway and the third gateway corresponding to the control network, the business network and the storage network respectively. The number of the first auxiliary nodes can be adjusted according to the actual situation. The corresponding counter can be set when performing ping detection, and the ping detection is stopped when the ping packet loss rate reaches a preset value. In addition, in addition to using salt to implement concurrent ping, concurrent ping can also be implemented based on Shell scripts or other scripts, and the ping detection can be monitored using preset network monitoring tools, such as Nagios (i.e., a network monitoring tool), so as to query the detection results in real time. It is worth mentioning that ping and consul can be implemented in a concurrent and topological manner respectively, and a timeout mechanism is set for both to ensure that the stuck scene can be quickly and accurately identified.
[0075] In this embodiment, after performing corresponding fault processing based on the fault scenario, first determine whether the preset virtual machine evacuation switch is in the on state; if the preset virtual machine evacuation switch is in the on state, determine whether the business resources meet the preset resource requirements, that is, by evaluating the current network resources, determine whether the business resources meet the preset resource requirements based on the evaluation results; if the business resources do not meet the preset resource requirements, perform the virtual machine evacuation operation on the target computing node based on the business priority from high to low, that is, start evacuating from the virtual machine with the highest business priority until the business resource requirements are met or all evacuable virtual machines have been considered. Specifically, the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the business priority includes: determining whether the preset virtual machine evacuation switch is in the on state; if the preset virtual machine evacuation switch is in the on state, determine whether the business resources meet the preset resource requirements; if the business resources do not meet the preset resource requirements, perform the virtual machine evacuation operation on the target computing node based on the business priority from high to low.
[0076] Furthermore, if the network states of the control network, business network and storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, and the fault scenario corresponding to the alarm information does not belong to node shutdown or node freezing, determine whether the fault scenario corresponding to the alarm information is a node restart. Similar to the previous two fault scenarios, network ping detection and node ping detection are also required, that is, based on the first gateway, the second gateway and the third gateway corresponding to the control network, the business network and the storage network, respectively, a second network ping detection is performed on each computing node, and the second network detection result is used to determine the second auxiliary node that meets the second preset packet loss rate, and based on the second preset number of detection rounds, a second node ping detection is performed on the second auxiliary node and the target computing node, so as to determine whether the network has recovered based on the second node detection result; if the second node detection result indicates that there is a recovery of the network state, it is determined that the fault scenario corresponding to the alarm information is a node restart, and further determined through the intelligent platform management interface. Step 1: Determine whether there is a restart or memory failure event in the hardware status of the target computing node; if there is a restart or memory failure event in the hardware status of the target computing node, determine whether the software service status of the target computing node has been restored to the running state; if the software service status of the target computing node has been restored to the running state, determine that the fault scenario corresponding to the alarm information is a node restart, and perform a hot migration operation of the virtual machine on the target computing node; if the second node detection result indicates that there is no recovery of the network status, indicating that the business has been in a faulty state, shut down the physical machine corresponding to the target computing node, and perform a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the business priority. It is worth mentioning that the second preset number of detection rounds and the number of packets for ping detection default to four rounds of ping detection and ping50 packets, which can also be adjusted accordingly according to actual conditions.
[0077] In a specific implementation, the status of nova compute service is checked through the command line tool or API (i.e., application programming interface) of OpenStack to determine whether the software service status of the target computing node has been restored to the running state; if the corresponding status of nova compute service is restored to up, it means that the software service status of the target computing node has been restored to the running state; if the software service status of the target computing node has not been restored to the running state, wait for 10 minutes. If the corresponding status of nova compute service is restored to up after 10 minutes, it is determined that the fault scenario corresponding to the alarm information is a node restart and supports virtual machine hot migration, and the virtual machine hot migration operation on the target computing node is performed. It should be noted that before hot migration, the nova computeservice of the target computing node needs to be disabled to avoid new virtual machines being scheduled to the target computing node during the migration process; when performing a virtual machine hot migration operation, the status of the virtual machine can be monitored to ensure the success of the hot migration. If the detection result of the second node indicates that there is no recovery of the network status, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority to quickly start the business on the new computing node.
[0078] Specifically, the method of determining the corresponding fault scenario based on the preset evaluation order and the alarm information, and performing fault processing using the fault scenario, service priority, and the preset virtual machine evacuation switch, includes: if the alarm information indicates that the network states corresponding to the control network, the service network, and the storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, then performing a second network ping test on each of the computing nodes based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the service network, and the storage network, respectively, to obtain a second network test result; determining a second auxiliary node that meets a second preset packet loss rate using the second network test result, and performing a second node ping test on the second auxiliary node and the target computing node based on a second preset number of test rounds to obtain a second node test result; if If the second node detection result indicates that there is recovery of the network status, it is determined that the fault scenario corresponding to the alarm information is node restart, and the intelligent platform management interface is used to determine whether the hardware status of the target computing node has a restart or a memory failure; if the hardware status of the target computing node has a restart or a memory failure, it is determined whether the software service status of the target computing node is in a running state; if the software service status of the target computing node is in a running state, it is determined that the fault scenario corresponding to the alarm information is node restart, and a hot migration operation of the virtual machine on the target computing node is performed; if the second node detection result indicates that there is no recovery of the network status, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority.
[0079] In this embodiment, if the network statuses of the control network, business network and storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, and the fault scenario corresponding to the alarm information does not belong to node shutdown, node freezing or node restart, it is determined whether the fault scenario corresponding to the alarm information is a network failure. Similar to the previous three fault scenarios, a third network ping detection is performed on each computing node based on the first gateway, the second gateway and the third gateway corresponding to the control network, the business network and the storage network respectively, and the third network detection result is used to determine the third auxiliary node that meets the third preset packet loss rate, and the third node ping detection is performed on the third auxiliary node and the target computing node based on the third preset detection round number, and then the network status and hot migration support of the target fault network are checked using the third node detection result. If the inspection result indicates that the network status of the target fault network is abnormal and does not support hot migration, it is determined whether the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition; if the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority; if the inspection result indicates that the network status of the target fault network is abnormal and supports hot migration, the virtual machine hot migration operation on the target computing node is performed. It is worth mentioning that the third preset number of detection rounds and the number of packets for ping detection are two rounds of ping detection and 50 ping packets by default, and can also be adjusted accordingly according to actual conditions.
[0080] In a specific implementation, if the target fault network is a service network, the network status of the service network is checked, that is, data_net_check is checked. If data_net_check is noPass, it indicates that the network status of the service network is abnormal. The network status and hot migration support of the service network are checked, that is, live_migration_support_check is checked. If live_migration_support_check is noPass, it indicates that the service network does not support hot migration. Therefore, if data_net_check is noPass and live_migration_support_check is noPass, data_net_force_evacuate is checked. If data_net_force_evacuate is True, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the service priority; if data_net_force_evacuate is True or not set, no operation is performed, and only a log alarm is required. If data_net_check is noPass and live_migration_support_check is Pass, a live migration operation of the virtual machine on the target computing node is performed.
[0081] Specifically, the determining of the corresponding fault scenario based on the preset evaluation order and the alarm information, and performing fault processing using the fault scenario, service priority, and the preset virtual machine evacuation switch, includes: if the alarm information indicates that there is an abnormality in any network status of the control network, the service network, and the storage network of the target computing node corresponding to the alarm information within a preset time period, obtaining the corresponding target fault network; performing a third network ping detection on each of the computing nodes based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the service network, and the storage network, respectively, to obtain a third network detection result; determining a third auxiliary node that meets a third preset packet loss rate using the third network detection result, and pinging the third auxiliary node with the target computing node based on a third preset number of detection rounds; A third node ping detection is performed on the target computing node to obtain a third node detection result; based on the third node detection result, the network status and hot migration support of the target fault network are checked; if the inspection result indicates that the network status of the target fault network is abnormal and does not support hot migration, whether the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition is determined; if the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority; if the inspection result indicates that the network status of the target fault network is abnormal and supports hot migration, the virtual machine hot migration operation is performed on the target computing node.
[0082] In this embodiment, if the fault scenario corresponding to the alarm information does not belong to node shutdown, node stuck, node restart or network failure, it is determined whether the fault scenario corresponding to the alarm information is a hardware fault, that is, whether the alarm information indicates that the fault of the physical machine on which the virtual machine on the corresponding target computing node depends is one of CPU, power supply, motherboard, memory and hard disk fault; if so, it is determined that the fault scenario corresponding to the alarm information is a hardware fault. Common hardware faults include memory faults, hard disk faults and motherboard faults. If the hardware fault is a memory fault, it is further determined that the memory fault is a UE (memory uncorrectable error) error or a CE (compilation error, i.e., correctable error) error; if the memory fault is a UE error, which is a fault scenario that usually causes a node restart, it is determined whether the software service status of the target computing node is a running state; if the software service status of the target computing node is a running state, it is determined that the fault scenario corresponding to the alarm information is a node restart, and a hot migration operation of the virtual machine on the target computing node is performed; if the memory fault is a CE error, the CE error is ignored. If the hardware failure is a hard disk failure, it will cause abnormal storage access by the virtual machine. If it is distributed storage, a broken disk may cause the iowait (the percentage of time waiting for I / O operations to complete) of disk access to increase, but it will not directly cause a virtual machine failure, so just replace the disk. If the hardware failure is a motherboard failure, it is divided into three situations according to the degree of failure: node shutdown, node stuck, and node restart. It can be handled according to the above-mentioned failure scenarios. If the hardware failure does not cause memory failure, hard disk failure, and motherboard failure, an alarm notification will be issued.
[0083] Specifically, the corresponding fault scenario is determined based on the preset evaluation order and the alarm information, so as to use the fault scenario, business priority and the preset virtual machine evacuation switch to perform fault processing, including: if the alarm information indicates that the fault of the physical machine on which the virtual machine on the corresponding target computing node depends is one of the CPU, power supply, motherboard, memory and hard disk failure, then the fault scenario corresponding to the alarm information is determined to be a hardware failure; if the hardware failure is a memory failure, whether the memory failure is a UE error is determined; if the memory failure is a UE error, whether the software service status of the target computing node is a running state is determined; if the software service status of the target computing node is a running state, then the fault scenario corresponding to the alarm information is determined to be a node restart, and a virtual machine hot migration operation is performed on the target computing node; if the hardware failure is a hard disk failure, a corresponding disk replacement operation is performed.
[0084] It is understandable that the concurrent evacuation of virtual machines based on the cloud computing platform, but too many concurrency numbers may cause great pressure on the computing nodes and storage, and ultimately lead to business recovery failure. Therefore, the maximum concurrency number needs to be set for different models of computing nodes and storage to ensure a balance between efficiency and success rate. The hot migration of virtual machines is performed one by one, and concurrency is not supported; when 98% of the Jinxiuhua process is executed, force complete will be executed; if one fails, the hot migration of other virtual machines will also stop; if the migration takes more than 20 minutes, it is considered a failure, and the hot migration of other virtual machines will also be stopped. In addition, the PRIORITY attribute of the virtual machine can be set for business priority, which supports settings from 1 to 10 levels. If not set, the default is 10. When sending a virtual machine evacuation command, evacuation requests are sent to virtual machines from priority 1 from high to low.
[0085] Step S14: If the target computing node corresponding to the alarm information is in the preset blacklist, the computing node is determined as a node under maintenance, and triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information is prohibited.
[0086] In this embodiment, if the target computing node corresponding to the alarm information is in the preset blacklist, the computing node is determined as a node under maintenance, that is, the node under maintenance may be generating alarm information, but the system will not perform the operation of determining the fault scenario based on the preset evaluation order and the alarm information. In other words, the system will not attempt to troubleshoot, isolate or repair the node under maintenance.
[0087] As can be seen from the above, this application builds an alarm monitoring system based on telegraf and in combination with a preset collector. The alarm monitoring system monitors the SNMP trap alarms of the computing nodes where the virtual machines are located, the resources and services inside the operating system, and the network status of the computing nodes to generate alarm information, and then determines whether the target computing node corresponding to the alarm information is in the preset blacklist. If it is not in the blacklist, the corresponding fault scenarios are determined based on the preset evaluation order and the alarm information, such as node shutdown, node stuck, and other fault scenarios. In this way, fault handling is performed according to business priority and the preset virtual machine evacuation switch to ensure that key businesses are restored first when resources are limited. Through a unified alarm monitoring system and alarm processing process, comprehensive monitoring and management of the computing nodes where the virtual machines are located is achieved, improving operation and maintenance efficiency and system reliability.
[0088] Accordingly, see Figure 2 As shown, the present application also provides a virtual machine high availability device, including:
[0089] The alarm information generation module 11 is used to build an alarm monitoring system using telegraf and a preset collector, monitor the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node through the alarm monitoring system, and generate alarm information based on the preset alarm type and the monitoring results; the preset collector is a collector determined based on the Consul cluster;
[0090] A target computing node determination module 12 is used to determine whether the target computing node corresponding to the alarm information is in a preset blacklist;
[0091] A fault handling module 13 is used to determine the corresponding fault scenario based on a preset evaluation order and the alarm information if the target computing node corresponding to the alarm information is not in a preset blacklist, so as to perform fault handling by using the fault scenario, service priority and a preset virtual machine evacuation switch; the fault scenario includes node shutdown, node stuck, node restart, network failure and hardware failure;
[0092] The node under maintenance determination module 14 is used to determine the computing node as a node under maintenance if the target computing node corresponding to the alarm information is in a preset blacklist, and prohibit triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information.
[0093] As can be seen from the above, this application builds an alarm monitoring system based on telegraf and in combination with a preset collector. The alarm monitoring system monitors the SNMP trap alarms of the computing nodes where the virtual machines are located, the resources and services inside the operating system, and the network status of the computing nodes to generate alarm information, and then determines whether the target computing node corresponding to the alarm information is in the preset blacklist. If it is not in the blacklist, the corresponding fault scenarios are determined based on the preset evaluation order and the alarm information, such as node shutdown, node stuck, and other fault scenarios. In this way, fault handling is performed according to business priority and the preset virtual machine evacuation switch to ensure that key businesses are restored first when resources are limited. Through a unified alarm monitoring system and alarm processing process, comprehensive monitoring and management of the computing nodes where the virtual machines are located is achieved, improving operation and maintenance efficiency and system reliability.
[0094] In some specific implementations, the alarm information generating module 11 may specifically include:
[0095] A network card allocation unit, configured to allocate the network cards used by the virtual machine for communication of the control network, the business network, and the storage network to different logical interfaces through a bridge technology, and after the allocation operation is completed, deploy a first Consul agent, a second Consul agent, and a third Consul agent corresponding to the control network, the business network, and the storage network based on the computing node to obtain a first Consul cluster, a second Consul cluster, and a third Consul cluster;
[0096] A preset collector construction unit, configured to construct a preset collector based on the first Consul cluster, the second Consul cluster, and the third Consul cluster;
[0097] A service monitoring unit, configured to monitor the SNMP trap alarms of the computing node where the virtual machine is located and the resources and services inside the operating system by using telegraf to obtain a first monitoring result;
[0098] A network status monitoring unit, configured to monitor the network status of the computing node using a preset collector and based on a preset monitoring frequency to obtain a second monitoring result;
[0099] A target monitoring result determination unit is used to determine a target monitoring result that meets the preset alarm type based on the first monitoring result and the second monitoring result, and generate corresponding alarm information based on the target monitoring result.
[0100] In some specific implementations, the fault processing module 12 may specifically include:
[0101] A power status checking unit, configured to check the power status of the target computing node through an intelligent platform management interface if the alarm information indicates that the network statuses of the control network, the service network, and the storage network of the target computing node corresponding to the alarm information within a preset time period are all abnormal;
[0102] A first virtual machine evacuation unit is used to determine that the fault scenario corresponding to the alarm information is node shutdown if the inspection result indicates that the power state of the target computing node is off, and perform a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority;
[0103] A first network detection unit, configured to perform a first network ping detection on each of the computing nodes based on a first gateway, a second gateway, and a third gateway respectively corresponding to the control network, the business network, and the storage network, to obtain a first network detection result if the inspection result indicates that the power state of the target computing node is not off;
[0104] A first node detection unit, configured to determine a first auxiliary node that meets a first preset packet loss rate by using the first network detection result, and perform a first node ping detection on the first auxiliary node and the target computing node based on a first preset detection round number to obtain a first node detection result;
[0105] A physical shutdown unit is used to determine that the fault scenario corresponding to the alarm information is a node freeze, and shut down the physical machine corresponding to the target computing node, and then perform a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the business priority if the first node detection result is that the packet loss rate corresponding to each network card on the target computing node meets the second preset packet loss rate and there is no recovery of the network status.
[0106] In some specific implementations, the fault processing module 12 may specifically include:
[0107] A virtual machine evacuation switch determination unit, used to determine whether the preset virtual machine evacuation switch is in an on state;
[0108] A business resource determination unit, configured to determine whether business resources meet preset resource requirements if the preset virtual machine evacuation switch is in an on state;
[0109] The second virtual machine evacuation unit is used to perform a virtual machine evacuation operation on the target computing node based on the order of business priorities from high to low if the business resources do not meet the preset resource requirements.
[0110] In some specific implementations, the fault processing module 12 may specifically include:
[0111] A second network detection unit is configured to perform a second network ping detection on each of the computing nodes based on the first gateway, the second gateway, and the third gateway respectively corresponding to the control network, the business network, and the storage network, to obtain a second network detection result if the alarm information indicates that the network states corresponding to the control network, the business network, and the storage network of the target computing node corresponding to the alarm information within a preset time period are all abnormal;
[0112] A second node detection unit, configured to determine a second auxiliary node that meets a second preset packet loss rate by using the second network detection result, and perform a second node ping detection on the second auxiliary node and the target computing node based on a second preset detection round number to obtain a second node detection result;
[0113] A hardware status judgment unit, configured to determine that the fault scenario corresponding to the alarm information is a node restart if the second node detection result indicates that the network status has been restored, and to judge whether the hardware status of the target computing node has a restart or a memory failure through an intelligent platform management interface;
[0114] A first software service status determination unit, configured to determine whether the software service status of the target computing node is in a running state if the hardware status of the target computing node has a restart or a memory failure;
[0115] A first virtual machine hot migration unit, configured to determine that the fault scenario corresponding to the alarm information is a node restart if the software service state of the target computing node is a running state, and perform a virtual machine hot migration operation on the target computing node;
[0116] The third virtual machine evacuation unit is used to shut down the physical machine corresponding to the target computing node if the second node detection result indicates that there is no recovery of the network status, and then perform the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the business priority.
[0117] In some specific implementations, the fault processing module 12 may specifically include:
[0118] A target fault network acquisition unit, configured to acquire a corresponding target fault network if the alarm information indicates that any network state of a control network, a business network, and a storage network of a target computing node corresponding to the alarm information within a preset time period is abnormal;
[0119] A third network detection unit, configured to perform a third network ping detection on each of the computing nodes based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the business network, and the storage network, respectively, to obtain a third network detection result;
[0120] A third node detection unit, configured to determine a third auxiliary node that meets a third preset packet loss rate by using the third network detection result, and perform a third node ping detection on the third auxiliary node and the target computing node based on a third preset detection round number to obtain a third node detection result;
[0121] a parameter value judgment unit, configured to check the network status and hot migration support of the target fault network based on the detection result of the third node, and if the inspection result indicates that the network status of the target fault network is abnormal and hot migration is not supported, determine whether the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition;
[0122] a fourth virtual machine evacuation unit, configured to shut down the physical machine corresponding to the target computing node if the parameter value corresponding to the virtual machine evacuation parameter satisfies the preset parameter condition, and then perform an evacuation operation of the virtual machine on the target computing node based on the preset virtual machine evacuation switch and the service priority;
[0123] The second virtual machine hot migration unit is used to perform a virtual machine hot migration operation on the target computing node if the inspection result indicates that the network status of the target fault network is abnormal and supports hot migration.
[0124] In some specific implementations, the fault processing module 12 may specifically include:
[0125] A fault scenario determination unit, configured to determine that the fault scenario corresponding to the alarm information is a hardware fault if the fault of the physical machine on which the virtual machine on the corresponding target computing node depends is one of a CPU, power supply, mainboard, memory, and hard disk fault;
[0126] a memory fault judgment unit, configured to judge whether the memory fault is a UE error if the hardware fault is a memory fault;
[0127] A second software service status judgment unit, configured to judge whether the software service status of the target computing node is a running status if the memory fault is a UE error;
[0128] a fifth virtual machine evacuation unit, configured to determine, if the software service state of the target computing node is in a running state, that the fault scenario corresponding to the alarm information is a node restart, and perform a virtual machine hot migration operation on the target computing node;
[0129] The hard disk failure determination unit is used to perform a corresponding disk replacement operation if the hardware failure is a hard disk failure.
[0130] Furthermore, the present application also discloses an electronic device. Figure 3 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input and output interface 25 and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the virtual machine high availability method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0131] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0132] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0133] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the virtual machine high availability method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0134] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disclosed virtual machine high availability method is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.
[0135] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0136] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0137] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0138] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0139] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A virtual machine high availability method, characterized in that: include: An alarm monitoring system is constructed using telegraf and a preset collector. The alarm monitoring system monitors the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node, and generates alarm information based on the preset alarm type and monitoring results. The preset collector is a collector determined based on the Consul cluster; Determine whether the target computing node corresponding to the alarm information is in a preset blacklist; If the target computing node corresponding to the alarm information is not in the preset blacklist, the corresponding fault scenario is determined based on the preset evaluation order and the alarm information, so as to use the fault scenario, service priority and preset virtual machine evacuation switch to handle the fault; the fault scenario includes node shutdown, node stuck, node restart, network failure and hardware failure; If the target computing node corresponding to the alarm information is in a preset blacklist, the computing node is determined as a node under maintenance, and triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information is prohibited.
2. The virtual machine high availability method according to claim 1, characterized in that: The alarm monitoring system is constructed by using telegraf and a preset collector, and the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node are monitored through the alarm monitoring system, and alarm information is generated based on the preset alarm type and the monitoring results, including: Allocate the network cards used by the virtual machine for communication of the control network, the business network, and the storage network to different logical interfaces through the bridge technology, and after the allocation operation is completed, deploy the first Consul agent, the second Consul agent, and the third Consul agent corresponding to the control network, the business network, and the storage network based on the computing node to obtain the first Consul cluster, the second Consul cluster, and the third Consul cluster; Building a preset collector based on the first Consul cluster, the second Consul cluster, and the third Consul cluster; Using telegraf to monitor the SNMP trap alarm of the computing node where the virtual machine is located and the resources and services inside the operating system to obtain a first monitoring result; Using a preset collector and based on a preset monitoring frequency, the network status of the computing node is monitored to obtain a second monitoring result; Determine a target monitoring result that meets the preset alarm type based on the first monitoring result and the second monitoring result, and generate corresponding alarm information based on the target monitoring result; The first Consul cluster, the second Consul cluster and the third Consul cluster communicate through a rumor protocol.
3. The virtual machine high availability method according to claim 2, characterized in that: The determining a corresponding fault scenario based on a preset evaluation sequence and the alarm information, so as to perform fault processing by using the fault scenario, service priority and a preset virtual machine evacuation switch, includes: If the alarm information indicates that the network states of the control network, the service network and the storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, the power supply state of the target computing node is checked through the intelligent platform management interface; If the inspection result indicates that the power state of the target computing node is off, determining that the fault scenario corresponding to the alarm information is node shutdown, and performing a virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority; If the inspection result indicates that the power state of the target computing node is not off, performing a first network ping detection on each computing node based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the business network, and the storage network, respectively, to obtain a first network detection result; Determine a first auxiliary node that meets a first preset packet loss rate by using the first network detection result, and perform a first node ping detection on the first auxiliary node and the target computing node based on a first preset detection round number to obtain a first node detection result; If the detection result of the first node is that the packet loss rate corresponding to each network card on the target computing node meets the second preset packet loss rate and there is no recovery of the network status, it is determined that the fault scenario corresponding to the alarm information is a node stuck, and the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the business priority.
4. The virtual machine high availability method according to claim 3, characterized in that: The performing the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority includes: Determine whether the preset virtual machine evacuation switch is in an on state; If the preset virtual machine evacuation switch is on, determining whether the business resources meet the preset resource requirements; If the business resources do not meet the preset resource requirements, the virtual machine evacuation operation on the target computing node is performed based on the business priority order from high to low.
5. The virtual machine high availability method according to claim 2, characterized in that: The determining a corresponding fault scenario based on a preset evaluation sequence and the alarm information, so as to perform fault processing by using the fault scenario, service priority and a preset virtual machine evacuation switch, includes: If the alarm information indicates that the network states corresponding to the control network, the business network and the storage network of the target computing node corresponding to the alarm information are all abnormal within a preset time period, a second network ping test is performed on each computing node based on the first gateway, the second gateway and the third gateway corresponding to the control network, the business network and the storage network, respectively, to obtain a second network test result; Determine a second auxiliary node that meets a second preset packet loss rate by using the second network detection result, and perform a second node ping detection on the second auxiliary node and the target computing node based on a second preset detection round number to obtain a second node detection result; If the second node detection result indicates that there is a recovery of the network status, it is determined that the fault scenario corresponding to the alarm information is a node restart, and it is determined through the intelligent platform management interface whether the hardware status of the target computing node has a restart or a memory failure; If the hardware status of the target computing node has a restart or memory failure, determining whether the software service status of the target computing node is in a running state; If the software service state of the target computing node is in the running state, determining that the fault scenario corresponding to the alarm information is a node restart, and performing a hot migration operation of the virtual machine on the target computing node; If the second node detection result indicates that there is no recovery of the network status, the physical machine corresponding to the target computing node is shut down, and then the virtual machine evacuation operation on the target computing node is performed based on the preset virtual machine evacuation switch and the service priority.
6. The virtual machine high availability method according to claim 2, characterized in that: The determining a corresponding fault scenario based on a preset evaluation sequence and the alarm information, so as to perform fault processing by using the fault scenario, service priority and a preset virtual machine evacuation switch, includes: If the alarm information indicates that any network status of the control network, business network and storage network of the target computing node corresponding to the alarm information is abnormal within a preset time period, then the corresponding target fault network is obtained; Performing a third network ping detection on each of the computing nodes based on the first gateway, the second gateway, and the third gateway corresponding to the control network, the business network, and the storage network, respectively, to obtain a third network detection result; Determine a third auxiliary node that meets a third preset packet loss rate by using the third network detection result, and perform a third node ping detection on the third auxiliary node and the target computing node based on a third preset detection round number to obtain a third node detection result; Based on the detection result of the third node, the network status and hot migration support of the target fault network are checked, and if the check result indicates that the network status of the target fault network is abnormal and hot migration is not supported, it is determined whether the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition; If the parameter value corresponding to the virtual machine evacuation parameter meets the preset parameter condition, shut down the physical machine corresponding to the target computing node, and then perform the virtual machine evacuation operation on the target computing node based on the preset virtual machine evacuation switch and the service priority; If the inspection result indicates that the network status of the target faulty network is abnormal and supports hot migration, a hot migration operation of the virtual machine on the target computing node is performed.
7. The virtual machine high availability method according to claim 2, characterized in that: The determining a corresponding fault scenario based on a preset evaluation sequence and the alarm information, so as to perform fault processing by using the fault scenario, service priority and a preset virtual machine evacuation switch, includes: If the alarm information indicates that the physical machine on which the virtual machine on the corresponding target computing node depends has a fault of one of CPU, power supply, motherboard, memory and hard disk failure, then it is determined that the fault scenario corresponding to the alarm information is a hardware failure; If the hardware fault is a memory fault, determining whether the memory fault is a UE error; If the memory fault is a UE error, determining whether the software service state of the target computing node is a running state; If the software service state of the target computing node is in the running state, determining that the fault scenario corresponding to the alarm information is a node restart, and performing a hot migration operation of the virtual machine on the target computing node; If the hardware failure is a hard disk failure, a corresponding disk replacement operation is performed.
8. A virtual machine high availability device, characterized in that: include: An alarm information generation module is used to build an alarm monitoring system using telegraf and a preset collector, monitor the SNMP trap alarms of the computing node where the virtual machine is located, the resources and services inside the operating system, and the network status of the computing node through the alarm monitoring system, and generate alarm information based on the preset alarm type and monitoring results; The preset collector is a collector determined based on the Consul cluster; A target computing node determination module is used to determine whether the target computing node corresponding to the alarm information is in a preset blacklist; A fault handling module, for determining a corresponding fault scenario based on a preset evaluation order and the alarm information if the target computing node corresponding to the alarm information is not in a preset blacklist, so as to perform fault handling using the fault scenario, service priority, and a preset virtual machine evacuation switch; the fault scenario includes node shutdown, node freeze, node restart, network failure, and hardware failure; The node under maintenance determination module is used to determine the computing node as a node under maintenance if the target computing node corresponding to the alarm information is in a preset blacklist, and prohibit triggering the operation of determining the corresponding fault scenario based on the preset evaluation order and the alarm information.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the virtual machine high availability method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the virtual machine high availability method according to any one of claims 1 to 7 is implemented.