A virtual machine management method and device, a cloud computing platform and a medium

CN117555711BActive Publication Date: 2026-09-11ZHONGKE FANGDE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311444718.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2026-09-11
Estimated Expiration
2043-11-01

AI Technical Summary

Technical Problem

[0003]相关技术中,可以采用通过在虚拟机内部部署代理程序,利用代理程序检测虚拟机是否正常,但是这种入侵方式存在一定的安全隐患,而且对故障判断的准确度不高,也无法个性化对虚拟机进行故障判断

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117555711B_ABST
    Figure CN117555711B_ABST
Patent Text Reader

Abstract

The present disclosure provides a virtual machine management method and device, a cloud computing platform and a medium. The method comprises: obtaining physical hardware running information and virtual machine running information of a computing node in at least one collection period; determining a host running state of the computing node in each collection period based on the physical hardware running information of the computing node in each collection period; and if it is determined that a target virtual machine is in a crash state in a first target collection period based on the virtual machine running information of the computing node in the first target collection period, managing the target virtual machine based on the host running state of the computing node in the first target collection period. The method can improve the accuracy of virtual machine fault judgment and the speed of fault elimination, avoid the adverse effects of false judgments on virtual machines, and can also manage virtual machines in a targeted manner to reduce the impact range on business continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing technology, and in particular to a virtual machine management method, apparatus, cloud computing platform, and medium. Background Technology

[0002] In related technologies, cloud computing platforms can provide users with processing, storage, networking and other computing resource services through the Infrastructure as a Service (IaaS) framework. The IaaS framework uses virtualization technology, which can eliminate the user's upfront hardware purchase costs and is widely used in industries such as finance, telecommunications and energy.

[0003] In related technologies, an agent program can be deployed inside the virtual machine to detect whether the virtual machine is normal. However, this intrusion method has certain security risks, and the accuracy of fault diagnosis is not high, nor can it be used to diagnose virtual machine faults in a personalized manner. Summary of the Invention

[0004] According to one aspect of this disclosure, a virtual machine management method is provided, comprising:

[0005] Obtain physical hardware operation information of the computing node and virtual machine operation information of the computing node in at least one acquisition cycle;

[0006] The host machine running status of the computing node in each acquisition cycle is determined based on the physical hardware operation information of the computing node in each acquisition cycle.

[0007] If it is determined that the target virtual machine is in a crash state during the first target acquisition period based on the virtual machine running information of the computing node during the first target acquisition period, the target virtual machine is managed based on the host machine running state of the computing node during the first target acquisition period.

[0008] According to another aspect of this disclosure, a virtual machine management apparatus is provided, comprising:

[0009] The acquisition module is used to acquire physical hardware operation information of the computing node and virtual machine operation information of the computing node in at least one acquisition cycle.

[0010] The determination module is used to determine the host machine running status of the computing node in each collection cycle based on the virtual machine running information of the computing node in each collection cycle;

[0011] The management module is used to manage the target virtual machine if it is determined from the virtual machine running information of the computing node in the first target collection period that the target virtual machine is in a crash state in the first target collection period, based on the host running state of the computing node in the first target collection period.

[0012] According to another aspect of this disclosure, a cloud computing platform is provided, comprising:

[0013] Processor; and,

[0014] Memory for stored programs;

[0015] The program includes instructions that, when executed by the processor, cause the processor to perform the method according to an exemplary embodiment of the present disclosure.

[0016] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method according to exemplary embodiments of this disclosure.

[0017] In one or more technical solutions provided in the exemplary embodiments of this disclosure, when a target virtual machine is in a crash state during a first target data collection period, the target virtual machine can be managed based on the host machine's running state during that first target data collection period. Therefore, when a virtual machine in an exemplary embodiment of this disclosure is in a crash state during a certain data collection period, the host machine's running state during the same data collection period can be referenced to manage the crashed target virtual machine, thereby improving the accuracy of virtual machine fault diagnosis and troubleshooting speed, and avoiding the adverse effects of incorrect diagnosis on the virtual machine. Therefore, the method in the exemplary embodiments of this disclosure has high availability, can improve the continuity of system operation, and ensure the reliability and robustness of virtual machine management, thereby achieving the goal of improving the quality of operation and maintenance services.

[0018] Furthermore, the exemplary embodiments of this disclosure determine the running status of each virtual machine in different collection cycles by using the virtual machine running information of the computing node in each collection cycle. Then, based on the running status of each virtual machine in different collection cycles, the target virtual machine in a crash state and the corresponding first target collection cycle are obtained. Then, referring to the host machine running status of the computing node in the first target collection cycle, the target virtual machine in a crash state is managed. It is evident that the method of the exemplary embodiments of this disclosure can specifically manage target virtual machines in a crash state without affecting other virtual machines running on the host machine. Therefore, when managing virtual machines, the exemplary embodiments of this disclosure can minimize the impact on business continuity. Attached Figure Description

[0019] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0020] Figure 1 The diagram shown is a schematic representation of a cloud computing platform architecture of an exemplary embodiment of this disclosure.

[0021] Figure 2 A schematic diagram of a virtual machine management framework according to an exemplary embodiment of this disclosure is shown;

[0022] Figure 3 A schematic diagram of a virtual machine management method flow according to an exemplary embodiment of this disclosure is shown;

[0023] Figure 4 A schematic diagram illustrating a virtual hardware fault warning and troubleshooting process according to an exemplary embodiment of this disclosure is shown.

[0024] Figure 5 This illustration shows a schematic diagram of host virtual hardware fault warning and troubleshooting in an exemplary embodiment of this disclosure;

[0025] Figure 6 This illustration shows a schematic diagram of a connectivity fault warning and troubleshooting process for an external hardware cluster, as exemplified in this disclosure.

[0026] Figure 7 A schematic diagram illustrating the service component fault warning and troubleshooting process of an exemplary embodiment of this disclosure is shown;

[0027] Figure 8 A schematic block diagram of the functional modules of a virtual machine management apparatus according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 9 A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown;

[0029] Figure 10 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0034] Before introducing the embodiments of this disclosure, the relevant terms involved in the embodiments of this disclosure are first defined as follows:

[0035] The host machine is a dedicated physical server with a virtualized environment deployed on it. Users have exclusive access to the resources of the entire physical server and are physically isolated from the servers of other tenants.

[0036] A virtual machine (VM) is a computer system emulator that uses software to simulate a complete computer system with full hardware system functionality, running in a completely isolated environment, and can provide the functions of a physical computer.

[0037] Hot migration, also known as dynamic migration or online migration, is a migration method that is seamless for users. Virtual machines do not need to be shut down, and business operations are not interrupted, but it is a relatively complex migration method.

[0038] Cold migration involves shutting down the virtual machine and migrating the data. Only the system disk data and data disk data need to be migrated, while memory data does not need to be migrated. This is done using a block migration method.

[0039] In related technologies, OpenStack is an open-source cloud computing platform designed to provide robust Infrastructure as a Service (IaaS) functionality. As the project matures, it has been widely adopted in industries such as finance, telecommunications, and energy. High availability methods for OpenStack mainly include the following two:

[0040] The first approach involves deploying an agent program within the virtual machine to detect whether the virtual machine is functioning correctly, thus addressing high availability from the virtual machine perspective. However, this solution requires intruding into the virtual machine and carries certain security risks. Furthermore, network fluctuations can easily lead to misjudgments.

[0041] The second method involves assessing the physical machine's functionality by monitoring the host machine's compute services and network. OpenStack's evacuate function is then used to migrate virtual machines from the faulty host to a healthy one, minimizing downtime and impacting business operations. However, this method lacks accuracy in fault diagnosis. For example, if the compute service fails, the virtual machines might remain unaffected. Forcibly shutting down and migrating them would cause service interruption. Furthermore, this method cannot provide personalized fault diagnosis for virtual machines.

[0042] To address the aforementioned issues, this exemplary embodiment provides a virtual machine management method that can reduce the probability of misjudgment of faults, decrease the impact of faults on business operations, and manage virtual machines in a crashed state by referring to the host machine's operating status, thereby achieving the purpose of targeted virtual machine management.

[0043] The method of the exemplary embodiments disclosed herein can be applied to Figure 1 The diagram shows a cloud computing platform architecture. Figure 1 As shown, the cloud computing platform architecture 100 of the exemplary embodiments of this disclosure may include a management node 101, a cloud service component 102, and multiple computing nodes 103.

[0044] like Figure 1 As shown, there can be multiple management nodes 101, each capable of managing multiple compute nodes. If a management node 101 malfunctions, other management nodes can take over the management of the compute nodes. Each compute node 103 can deploy multiple virtual machines, which can be desktop cloud virtual machines. The cloud service component 102 can provide basic resource services, such as: identity service, image service, compute service, networking service, block storage service, object storage service, telemetry service, and dashboard service.

[0045] Figure 2 A schematic diagram of a virtual machine management framework according to an exemplary embodiment of this disclosure is shown. Figure 2As shown, the virtual machine management framework 200 of the exemplary embodiments of this disclosure may include a management module 202 and a plurality of agent modules 201. The agent modules 201 may be deployed on the compute node, and the management module 202 may be deployed on the management node.

[0046] like Figure 2 As shown, the agent module 201 is responsible for collecting the hardware and software operation information of the computing nodes. The management module 202 consists of an evaluation module 2021 and a processing module 2022. The evaluation module 2021 evaluates the hardware and software operation information of the computing nodes collected by the agent module 201 to obtain fault warning categories. Then, the processing module 2022 combines the fault warning object and the fault warning category to perform differentiated management of the virtual machine, reducing the impact of computing node failures on business operations. Moreover, the exemplary embodiment of this disclosure uses a non-intrusive method for computing node fault judgment, which can eliminate security risks.

[0047] In the exemplary embodiments of this disclosure, the hardware and software operation information of the computing node may include physical hardware operation information and virtual machine operation information, and may also include service component operation information. Based on this, the agent module 201 of the exemplary embodiments of this disclosure can collect physical hardware operation information, virtual machine operation information and service component operation information, and then report them to the evaluation module 2021; the evaluation module 2021 can perform fault warnings for physical hardware, virtual machines and service components based on the above-mentioned physical hardware operation information, virtual machine operation information and service component operation information, and determine whether there is a risk to the virtual machine, service or physical hardware.

[0048] In practical applications, the agent module 201 of this exemplary embodiment can collect physical hardware operation information, virtual machine operation information, and service component operation information every preset time interval (e.g., 60 seconds) and then report them to the management module 202. The agent module 201 of this exemplary embodiment can also report all information collected within the preset cumulative time interval (e.g., 10 minutes) to the management module 202. To reduce the impact on business continuity, the evaluation module 2021 can integrate the fault warning results of physical hardware, virtual machines, and components from multiple collection cycles to determine whether risk handling by the processing module 2022 is necessary. This not only ensures the accuracy of fault detection but also reduces misjudgments caused by network jitter.

[0049] For example, the fault warning categories in this exemplary embodiment can be divided into potential fault risks and abnormal operation risks. The severity of the problem corresponding to potential fault risks is higher than that corresponding to abnormal operation risks. When the fault warning category is potential fault risk, it indicates that the virtual machines of the compute node need to be migrated, while when the fault warning category is abnormal operation risk, the virtual machines of the compute node can be managed according to the fault warning object.

[0050] In some cases, the virtualization clusters of the aforementioned cloud computing platforms may experience various single points of failure, preventing the continued normal collection, analysis, and processing of data. For example, ... Figure 2 As shown, if the agent module 201 fails, the management module 202 will not be able to obtain the agent module 201. Based on this, the exemplary embodiment of this disclosure can reduce the risk of single point of failure by introducing a redundant management node. For example, the redundant management node can monitor the communication link between the computing node and the management node. When the communication link between the computing node and the management node fails, it can automatically take over the computing node, thereby ensuring that the redundant management node can continue to process the information reported by the agent module deployed on the computing node.

[0051] The virtual machine management method of the exemplary embodiments of this disclosure can be executed by a management node, such as the management module of the management node. Figure 3 A schematic diagram of a virtual machine management method flow according to an exemplary embodiment of this disclosure is shown. Figure 3 As shown, the virtual machine management method of this exemplary embodiment may include:

[0052] Step 301: Obtain the physical hardware operation information of the computing node and the virtual machine operation information of the computing node in at least one acquisition cycle.

[0053] The exemplary embodiments disclosed herein can obtain physical hardware operation information and virtual machine operation information of a computing node in one acquisition cycle, and can also obtain physical hardware operation information and virtual machine operation information of a computing node in multiple acquisition cycles.

[0054] Step 302: Determine the host machine running status of the computing node in each acquisition cycle based on the physical hardware operation information of the computing node in each acquisition cycle.

[0055] In practical applications, the exemplary embodiments of this disclosure can analyze physical hardware operating information to determine the host operating state of the computing node in each acquisition cycle. Based on this, determining the host operating state of the computing node in each acquisition cycle based on the physical hardware operating information of the computing node in each acquisition cycle can include:

[0056] Based on the physical hardware operation information of the computing node in each acquisition cycle, the operating status of at least one physical hardware in each acquisition cycle is determined. If the operating status of at least one physical hardware in the first target acquisition cycle is in an overloaded operating state, the host machine operating status of the computing node in the first target acquisition cycle is confirmed to be in an overloaded operating state.

[0057] For example, the physical hardware in the exemplary embodiments of this disclosure may include the physical hardware of a physical machine. For instance, the physical hardware of a physical machine may include multiple physical hardware components such as a central processing unit (CPU), memory, disk, and network. Accordingly, the physical hardware operation information includes: physical machine physical hardware operation information, that is, the operation information of multiple physical hardware components such as the central processing unit (CPU), memory, disk, and network.

[0058] During each collection cycle, the agent module can use the dmesg tool to obtain basic information about the physical hardware of the physical machine, including the CPU, memory, hard disk, network, etc. It can also use tools such as mpstat, sar, vmstat, pmap, du, netstat to obtain physical hardware operating information such as memory usage, CPU utilization, storage utilization, network throughput, etc.

[0059] Meanwhile, when the exemplary embodiments of this disclosure obtain physical hardware operation information of a physical machine or physical hardware information of an external hardware cluster, whether it is physical hardware operation information of a physical machine or physical hardware information of an external hardware cluster, the target physical hardware that is overloaded can be determined by threshold judgment. Here, the target physical hardware can be one or multiple.

[0060] For example: when memory usage exceeds a memory usage threshold, the memory is in an overloaded state; when CPU utilization exceeds a CPU utilization threshold, the CPU is in an overloaded state; when disk utilization exceeds a disk utilization threshold, the disk is in an overloaded state; when network throughput exceeds a throughput threshold, the network is in an overloaded state. If at least one of the CPU, memory, disk, or network is in an overloaded state during the same data collection period, the host machine operating state of the computing node during that period can be determined as overloaded.

[0061] It can be seen that for each acquisition cycle, when at least one of the multiple physical hardware devices is in an overloaded state during a certain acquisition cycle, the acquisition cycle can be confirmed as the first target acquisition cycle, and the host machine of the computing node is in an overloaded state during the first target acquisition cycle. If none of the multiple physical hardware devices are in an overloaded state during a certain acquisition cycle, the host machine of the computing node is in a normal operating state during the acquisition cycle.

[0062] Step 303: If it is determined that the target virtual machine is in a crash state during the first target acquisition period based on the virtual machine running information of the computing node during the first target acquisition period, the target virtual machine is managed based on the host running status of the computing node during the first target acquisition period.

[0063] In practical applications, the virtual machine runtime information in the exemplary embodiments of this disclosure includes the virtual hardware runtime information of the virtual machine, which can be obtained through the libvirt process for each acquisition cycle. For example, the evaluation module can obtain the virtual machine runtime information through the vish command of the libvirt process, and then determine the running status of the virtual machine in each acquisition cycle based on the virtual hardware runtime information of the virtual machine in each acquisition cycle.

[0064] The virtual machine's running state can include: running, idle, paused, shut down, crashed, or dying. When the evaluation module detects a virtual machine in a crash state during a certain acquisition period, that acquisition period is the first target acquisition period, and the virtual machine in a crash state during that acquisition period is the target virtual machine. At this time, the target virtual machine in a crash state can be managed by referring to the host machine's running state on the compute node during that first target acquisition period.

[0065] As can be seen, when a virtual machine in the exemplary embodiment of this disclosure is in a crash state during a certain data collection period, the host machine's running state during the same data collection period can be referenced to manage the target virtual machine in the crash state, thereby improving the accuracy of virtual machine fault diagnosis and avoiding the adverse effects of incorrect diagnosis on the virtual machine. Therefore, the method of the exemplary embodiment of this disclosure has high availability, can improve the continuity of system operation, and ensure the reliability and robustness of virtual machine management, thereby achieving the goal of improving the quality of operation and maintenance services.

[0066] Furthermore, the exemplary embodiments of this disclosure determine the running status of each virtual machine in different collection cycles by using the virtual machine running information of the computing node in each collection cycle. Then, based on the running status of each virtual machine in different collection cycles, the target virtual machine in a crash state and the corresponding first target collection cycle are obtained. Then, the target virtual machine in a crash state is managed with reference to the host machine running status of the computing node in the first target collection cycle. It is evident that the method of the exemplary embodiments of this disclosure can specifically manage target virtual machines in a crash state without affecting other virtual machines running on the host machine. Therefore, when managing virtual machines, the exemplary embodiments of this disclosure can minimize the impact on business continuity.

[0067] In one alternative approach, an exemplary embodiment of this disclosure manages the target virtual machine based on the host machine running state of the computing node during the first target acquisition period, which may include:

[0068] If the host machine is in normal operating state during the first target acquisition cycle, it means that the crash of the target virtual machine is unrelated to the host machine's operating state and may be caused by the target virtual machine itself. Therefore, the exemplary embodiment of this disclosure can restart the target virtual machine on the computing node.

[0069] If the host machine is in an overloaded state during the first target acquisition period, it indicates that the crash of the target virtual machine may be related to the host machine's operating state. For example, the host machine may be experiencing abnormalities such as overload due to too many virtual machines running. Therefore, the target virtual machine can be migrated to the target compute node, and this migration can be a cold migration. It should be understood that the target compute node in the exemplary embodiments of this disclosure can be a redundant compute node or an idle compute node in a virtualization cluster, but is not limited to these.

[0070] For example, when the evaluation module detects that the target virtual machine of the compute node is in a crash state during the first target acquisition period, if the host machine of the compute node is in a normal operating state during the first target acquisition period, the fault warning category of the target virtual machine can be determined as abnormal operation risk; when the evaluation module detects that the target virtual machine of the compute node is in a crash state during the first target acquisition period, if the host machine of the compute node is in an overloaded operating state during the first target acquisition period, the fault warning category of the target virtual machine can be determined as potential fault risk. Once the evaluation module determines the fault warning type of the target virtual machine, it can submit the fault warning type of the target virtual machine to the processing module.

[0071] When the aforementioned processing module receives a fault warning of potential fault risk, it can migrate the target virtual machine to the target compute node. After migration, the virtual machine is redefined and then restarted. Thus, the method of this exemplary embodiment reduces service interruption time by automatically migrating a target virtual machine in a crashed state to the target compute node, thereby achieving earlier fault repair. Furthermore, migrating the target virtual machine to the target compute node enables dynamic resource allocation and load balancing, preventing resource starvation and over-allocation. This ensures that each virtual machine on the compute node receives an appropriate share of resources, resolving resource contention issues among multiple virtual machines deployed on the same host.

[0072] In one alternative embodiment, the exemplary embodiment of this disclosure may further include: determining the operating status of at least one virtual hardware in each acquisition cycle based on the virtual machine operating information of the computing node in each acquisition cycle; if the operating status of the same virtual hardware in at least one acquisition cycle included in the first target time period is an overloaded operating state, migrating the target virtual machine to the target computing node.

[0073] In practical applications, the exemplary embodiments of this disclosure can obtain multiple virtual hardware operation information such as virtual machine memory usage, central processing unit (CPU) utilization, disk usage, and network throughput through libvirt. The operating status of each virtual hardware in each acquisition cycle is determined based on the virtual hardware operation information of the computing node in each acquisition cycle, which can be referred to the relevant descriptions above and will not be repeated here.

[0074] For example, the first target time period of this exemplary embodiment may include one or more consecutive acquisition cycles. When the first target time period includes one acquisition cycle, as long as the operating state of a virtual hardware is detected to be overloaded during that acquisition cycle, the target virtual machine can be hot-migrated to the target compute node. When the first target time period includes multiple consecutive acquisition cycles, if the operating state of a virtual hardware is overloaded during any one or more acquisition cycles included in the first target time period, the target virtual machine can be hot-migrated to the target compute node.

[0075] To reduce misjudgments caused by network jitter or other factors, the exemplary embodiments of this disclosure can determine the operating status of a virtual hardware device within each collection cycle of a first target time period. If the virtual hardware is operating under overload conditions for most of the collection cycles within the first target time period, it can be determined that the target virtual machine corresponding to the virtual hardware is also operating under overload conditions. When a virtual hardware device is operating under overload conditions for most of the collection cycles within the first target time period, it is operating under overload conditions for most of the first target time period. Therefore, virtual hardware operating under overload conditions is highly likely to have adverse effects on the corresponding target virtual machine, and may even lead to service interruption. In this case, the target virtual machine can be hot-migrated to the target computing node.

[0076] In practical applications, when a virtual hardware device is in an overloaded operating state during a certain data acquisition period, it indicates that the virtual hardware is under relatively high load. However, this could be due to accidental factors or an abnormal operation of the target virtual machine hosting the virtual hardware. Therefore, when the target virtual hardware is in an overloaded operating state during a certain data acquisition period, this exemplary embodiment can use that data acquisition period as the first data acquisition period included in the first target time period, and statistically analyze the operating state of the target virtual hardware across multiple data acquisition periods within the first target time period. If the target virtual hardware is in an overloaded operating state for most of the data acquisition periods within the first target time period, the possibility of a high virtual hardware load due to accidental factors can be ruled out. In this case, it can be confirmed that the target virtual machine corresponding to the target virtual hardware is in an overloaded operating state, the fault warning type of the target virtual machine is defined as abnormal operation risk, and then the target virtual machine is hot-migrated to the target compute node. The migrated virtual machine is then redefined on the target compute node, without affecting normal business operations.

[0077] To facilitate understanding of the virtual machine fault warning and troubleshooting process in the exemplary embodiments of this disclosure, the following uses a target virtual machine as an example. Figure 4 Please provide a detailed explanation. For example... Figure 4 As shown, the virtual machine fault warning and troubleshooting method of the exemplary embodiments of this disclosure may include the following steps:

[0078] Step 401: Based on the target virtual machine running information of the compute node in each acquisition cycle, determine the running status of the target virtual machine in each acquisition cycle. The running status of the target virtual machine in each acquisition cycle can be the running status of any virtual machine on the compute node, and these running statuses can include crashed and non-crash states.

[0079] Assuming there is a first target acquisition cycle among multiple acquisition cycles, if the target virtual machine is detected to be in a crash state in the first target acquisition cycle, the virtual hardware running status of the target virtual machine in the first target acquisition cycle cannot be obtained. Therefore, step 402 is executed. If the target virtual machine is detected to be in a non-crash state in the first target acquisition cycle, the running information of multiple virtual hardware such as memory, CPU, disk, and network of the target virtual machine in the first target acquisition cycle can be obtained. Therefore, step 405 can be executed.

[0080] Step 402: Detect whether the host machine running status of the computing node in the first target acquisition cycle is in an overloaded state.

[0081] When the host machine is in an overloaded state during the first target acquisition period, it indicates that the target virtual machine is in a crash state during the first target acquisition period and is related to the host machine. Therefore, the fault warning type of the target virtual machine can be determined as a potential fault risk, and then step 403 is executed. When the host machine is in a normal state during the first target acquisition period, it indicates that the target virtual machine is in a crash state during the first target acquisition period and is unrelated to the host machine. It may be caused by the target virtual machine itself. Therefore, the fault warning category of the target virtual machine can be determined as an abnormal operation risk, and then step 404 is executed.

[0082] Step 403: Migrate the target virtual machine to the target compute node. After migrating the target virtual machine to the target compute node, you can redefine the target virtual machine and restart it.

[0083] Step 404: Restart the target virtual machine on the compute node. This means restarting the target virtual machine from its crashed state on the compute node.

[0084] Step 405: Based on the running information of multiple virtual hardware devices of the target virtual machine in the first target acquisition period, determine the running status of each virtual hardware device of the target virtual machine in the first target acquisition period.

[0085] In practical applications, when the target virtual machine is detected to be in a non-crash state during the first target acquisition period, the operating status of each virtual hardware in the first target acquisition period can be determined based on the operating information of the target virtual machine in the first target acquisition period, such as memory, CPU, disk, network, etc.

[0086] Step 406: If the target virtual hardware in the target virtual machine is in an overloaded state during the first target acquisition period, take the first target acquisition period as the first target time period and count the operating status of the target virtual hardware in each acquisition period included in the first target time period.

[0087] In practical applications, if one of the virtual hardware devices is in an overloaded operating state during the first target acquisition cycle, the virtual hardware device in the overloaded operating state can be defined as the target virtual hardware device. The start time of the first target acquisition cycle is taken as the start time of the first target time period, and the operating state of the target virtual hardware device in each acquisition cycle within the first target time period is statistically analyzed.

[0088] Step 407: When the target virtual hardware is in an overloaded state for most of the acquisition cycles in the first target time period, the target virtual machine is hot-migrated to the target computing node.

[0089] For example, the first target time period of this exemplary embodiment may include M consecutive acquisition cycles. When the target virtual machine operates in an overloaded state for N target acquisition cycles out of the M consecutive acquisition cycles, and M represents an integer greater than or equal to 2, and N represents an integer greater than M / 2 and less than or equal to M, it indicates that the target virtual machine operates in an overloaded state for most of the acquisition cycles included in the first target time period. In this case, the evaluation module can consider that the target virtual machine operates in an overloaded state for most of the first target time period.

[0090] If the target virtual machine continues to run under overload, it may crash, leading to business interruption. However, considering that the target virtual machine can still work normally when it is under overload, it can be migrated to the target compute node via hot migration and then redefined. The whole process does not affect the normal operation of the business.

[0091] For example, when M=3 and N=2, if the evaluation module determines that the target virtual machine is in an overloaded state in at least two of the three consecutive acquisition cycles, the fault warning type of the target virtual machine can be marked as abnormal operation risk, and the target virtual machine can be hot-migrated to the target computing node; if the evaluation module determines that the target virtual machine is in an overloaded state in one of the three acquisition cycles or is in a normal operating state in all of them, the target virtual machine can be considered to be in a normal operating state in the first target time period, and no fault or risk warning treatment is required for the target virtual machine.

[0092] As can be seen, the exemplary embodiments of this disclosure can perform differentiated management of target virtual machines based on the severity of the problem corresponding to the fault warning type of the target virtual machine, thereby eliminating faults, restoring the normal operation of the target virtual machine, minimizing business interruption, ensuring service continuity and availability, and thus improving system reliability. Furthermore, the exemplary embodiments of this disclosure execute [the following] for each virtual machine on the compute node. Figure 4The described method allows virtual machines that need to be restarted within a compute node to be restarted without affecting the operation of other virtual machines within the compute node, and allows virtual machines that need to be migrated within a compute node to be migrated without affecting the operation of other virtual machines within the compute node, thereby reducing the impact on other virtual machines within the compute node.

[0093] Furthermore, in the method of the exemplary embodiments of this disclosure, if the same virtual hardware is in an overloaded operating state during multiple acquisition cycles within the first target time period, the target virtual machine is hot-migrated to the target computing node. This process actually manages the target virtual machine by referring to the operating state of the same virtual hardware in multiple acquisition cycles, thus avoiding the risk of misjudging the operating state of the target virtual machine caused by network jitter.

[0094] In one alternative approach, considering that when a computing node experiences a hardware failure, all virtual machines running on that computing node are affected, and services are also affected to varying degrees, the method of the exemplary embodiments of this disclosure may further include a method for early warning and troubleshooting of the target physical hardware failure.

[0095] Figure 5 A schematic diagram illustrating the host virtual hardware fault warning and troubleshooting process of an exemplary embodiment of this disclosure is shown. Figure 5 As shown, the fault warning and troubleshooting method for the target physical hardware of this exemplary embodiment may include:

[0096] Step 501: If the target physical hardware is in an overloaded state during the second target acquisition period, which is included in the second target time period, obtain the identity information of the target physical hardware during the second target acquisition period.

[0097] The second target time period of this exemplary embodiment includes one or more consecutive acquisition cycles. When the target physical hardware of this exemplary embodiment is in an overloaded operating state during a certain acquisition cycle included in the second target time period, that acquisition cycle included in the second target time period is the second target acquisition cycle. Therefore, this exemplary embodiment can obtain the identity information of the target physical hardware during the second target acquisition cycle included in the second target time period. The definition of the second target acquisition cycle referred to below is hereby referred to.

[0098] Step 502: Based on the identity information acquisition method of the same target physical hardware in at least one second target acquisition cycle, manage the virtual machine corresponding to the target physical hardware.

[0099] This exemplary embodiment of the disclosure manages the virtual machine corresponding to the target virtual machine based on the identity information acquisition method of the target physical hardware in at least one second target acquisition cycle, including:

[0100] If the identity information of the same target physical hardware is obtained through the basic physical hardware information of the computing node in at least one second target acquisition cycle, then some virtual machines of the computing node can be hot-migrated to the target computing node; if the identity information of the same target physical hardware is obtained through the abnormal operation log of the host machine of the computing node in at least one second target acquisition cycle, then all virtual machines of the computing node can be migrated to the target computing node.

[0101] In practical applications, when the identity information of the target physical hardware in the second target acquisition cycle is obtained from the basic physical hardware information of the computing node, it indicates that the identity information of the target physical hardware in the second target acquisition cycle was obtained through the basic physical hardware information of the computing node. At this time, although the target physical hardware is operating under overload conditions in the second target acquisition cycle, it is not damaged and will not affect normal business operations.

[0102] If the target physical hardware's identity information cannot be obtained from the compute node's basic physical hardware information during the second target acquisition cycle, it indicates that the target physical hardware within the physical machine may be damaged. Therefore, the method for obtaining the target physical hardware's identity information during the second target acquisition cycle can be changed to obtain the target physical hardware's identity information during the second target acquisition cycle from the compute node's host machine abnormal operation log during the second target acquisition cycle. In this case, the method for obtaining the target physical hardware's identity information during the second target acquisition cycle is the compute node's host machine abnormal operation log. It can be seen that the method for obtaining the target physical hardware's identity information in each acquisition cycle of this exemplary embodiment indirectly reflects whether the target physical hardware is damaged.

[0103] Considering that the identity information acquisition method of the target physical hardware in each second target acquisition cycle can indirectly reflect whether the target physical hardware is damaged, the exemplary embodiment of this disclosure can combine the identity information acquisition method of the target physical hardware in the second target acquisition cycle included in the second target time period to manage the virtual machine corresponding to the target physical hardware in a differentiated manner, so as to reduce the impact on business continuity.

[0104] For example, if the identity information of the same target physical hardware is obtained from the basic physical hardware information of the compute node for at least one second target acquisition cycle within the second target time period, it indicates that the target physical hardware on the host machine is overloaded, or that some hardware has a certain risk, but is still usable. Therefore, the fault warning type of the target physical hardware can be determined as abnormal operation risk, and some virtual machines on the compute node can be hot-migrated to the target compute node to reduce the operating pressure on the host machine. At this time, the migrated virtual machines are redefined on the target compute node, and this process does not affect the normal operation of the business. In addition, since shared storage has already been configured, the system disk and data disk of the migrated virtual machines do not require further operation.

[0105] An exemplary embodiment of this disclosure may randomly select a portion of virtual machines from all running virtual machines deployed on a compute node and migrate them to the target compute node. The number of these virtual machines may be 1 / 4 to 1 / 2, for example, 1 / 3, of all virtual machines deployed on the compute node.

[0106] If the identity information of the same target physical hardware cannot be obtained from the basic physical hardware information of the compute node in at least one second target acquisition cycle within the second target time period, it indicates that the host hardware is damaged. In this case, the identity information of the target physical hardware in at least one second target acquisition cycle can be obtained from the abnormal operation log of the compute node's host. Therefore, the fault warning type of the target physical hardware can be determined as a potential fault risk, and all virtual machines on the compute node can be migrated to the target compute node to ensure that all virtual machines on the compute node can run normally on the target compute node. For example, all virtual machines on the compute node can be migrated to the target compute node, and their system disks and corresponding XML files can be used on the target compute node. Then, the virtual machines can be redefined and restarted.

[0107] For example, when the second target time period includes multiple consecutive collection cycles, and the multiple consecutive collection cycles include multiple second target collection cycles, for each second target collection cycle within the second target time period, the number of second target collection cycles corresponding to the same target virtual hardware with the same identity information acquisition method can be counted. If the number of second target collection cycles accounts for more than half of the total number of collection cycles included in the second target time period, the virtual machine corresponding to the target virtual machine can be managed based on the identity information acquisition method of at least one second target collection cycle included in the second target time period for the same target virtual hardware.

[0108] In other words, when the second target time period includes multiple consecutive acquisition cycles, the number of second target acquisition cycles included in the second target time period accounts for more than 50% of the total number of consecutive acquisition cycles. For example, when the second target time period includes M consecutive acquisition cycles, and the M consecutive acquisition cycles contain N second target acquisition cycles, N represents an integer of M / 2.

[0109] For example, when M=3 and N=2, if the evaluation module detects the same target physical hardware fault warning type as abnormal operation risk in at least two of the three consecutive acquisition cycles (i.e. the second target acquisition cycle), the evaluation module can upload the abnormal operation risk of the target physical hardware to the processing module. When the processing module receives the abnormal operation risk of the target physical hardware, it can hot-migrate some virtual machines of the computing node to the target computing node.

[0110] Similarly, if the evaluation module detects a potential fault risk in the same target physical hardware in at least two of the three consecutive acquisition cycles (the second target acquisition cycle), the evaluation module will upload the potential fault risk of the target physical hardware to the processing module. When the processing module receives the potential fault risk of the target physical hardware, it can migrate all virtual machines of the compute node to the target compute node.

[0111] For example, when the second target time period of this exemplary embodiment of the present disclosure includes multiple second target acquisition cycles, migrating all virtual machines of the computing node to the target computing node may include:

[0112] When the identity information is obtained through the abnormal operation logs of the host machine on the compute node, the number of abnormal log records of the target physical hardware in each second target collection period can be determined based on the abnormal operation logs of the host machine on the compute node in each second target collection period. If the number of abnormal log records of the same target physical hardware in multiple second target collection periods included in the second target time period increases with the extension of time, it indicates that the damage to the host hardware of the compute node has intensified. In this case, migrating all virtual machines of the compute node to the target compute node can ensure the stable and reliable operation of the business.

[0113] In one alternative embodiment, the physical hardware operation information of this disclosure may include not only the physical machine hardware but also connectivity information of the external hardware nodes in the external hardware cluster of the computing node. These external hardware clusters may be distributed storage systems, such as Ceph, or distributed file systems such as NFS. In this case, to ensure the connectivity of the external hardware cluster, the method of this disclosure may further include a connectivity fault warning and troubleshooting method for the external hardware cluster.

[0114] Figure 6A schematic diagram illustrating the connectivity fault warning and troubleshooting process of an external hardware cluster according to an exemplary embodiment of this disclosure is shown. Figure 6 As shown, the connectivity fault warning and troubleshooting method for an external hardware cluster in an exemplary embodiment of this disclosure may include:

[0115] Step 601: Based on the connectivity information of external hardware nodes in multiple acquisition cycles of the external hardware cluster, determine the number of external hardware nodes that the computing node lost connection in multiple acquisition cycles.

[0116] For example, the connectivity information of external hardware nodes for each collection period can be the monitoring node information details of the distributed storage cluster for each collection period. The specific acquisition method includes: first, obtaining the monitoring node information details of the distributed storage cluster; if the monitoring node information details of the distributed storage cluster cannot be obtained, it means that the entire distributed storage cluster is offline, that is, the external hardware node disconnection rate reaches 100%.

[0117] If you obtain the monitoring node information details of the distributed storage cluster, you can retrieve the distributed storage nodes included in the cluster from these details. For each distributed storage node, you can use the distributed cluster status check command and the ping command to determine whether the network connectivity between the compute node and the distributed storage node is valid. If the compute node and the distributed storage node are not connected, it indicates that the distributed storage node is offline.

[0118] Step 602: Determine whether the number of external hardware nodes that have gone offline during at least one acquisition cycle within the third target time period is greater than 0 and less than or equal to the preset maximum number of offline nodes. The third target time period may include one or more consecutive acquisition cycles.

[0119] For a given data collection period, let the maximum number of disconnections be n0. When the number of disconnections n = 0, it indicates that the connectivity between the external hardware cluster and the compute nodes is normal. When the number of disconnections n satisfies 0 < n ≤ n0, it indicates that although there are disconnected external hardware nodes between the external hardware cluster and the compute nodes, the overall connectivity is not particularly poor. Therefore, the fault warning category for the external hardware cluster in this data collection period can be determined as abnormal operation risk. When the fault warning category for the external hardware cluster in at least one data collection period within the third target time period is abnormal operation risk, step 603 can be executed. When the number of disconnections n satisfies n > n0, it indicates that the connectivity between the external hardware cluster and the compute nodes is relatively poor, potentially making it difficult to meet the virtual machine requirements of the compute nodes. Therefore, the fault warning category for the external hardware cluster in this data collection period can be determined as potential fault risk. When the fault warning category for the external hardware cluster in at least one data collection period within the third target time period is potential fault risk, step 604 can be executed.

[0120] The preset maximum number of disconnections in the exemplary embodiments of this disclosure can be set according to the actual situation or with reference to related technologies. For example, when the external hardware cluster is a Ceph distributed storage cluster, the total number N of Ceph monitoring nodes in the Ceph distributed storage cluster is actually the total number of distributed storage nodes, and the preset maximum number of disconnections n0 can be equal to N-(N+1) / 2.

[0121] As can be seen, when the number of disconnected distributed storage nodes exceeds the preset maximum number, meaning the number of disconnected distributed storage nodes accounts for more than half of the total number of distributed storage nodes, it can be considered that the distributed storage cluster has a relatively serious fault, which may have already affected the operation of all virtual machines on that compute node. Therefore, the fault warning category for the external hardware cluster in this collection period can be determined as potential fault risk. Conversely, when the number of disconnected distributed storage nodes is greater than 0 and less than or equal to the preset maximum number of disconnected distributed storage nodes n0, meaning the number of disconnected distributed storage nodes accounts for less than half of the total number of distributed storage nodes, it can be considered that the problem with the distributed storage cluster is not very serious and is still within the system's tolerance range. Therefore, the fault warning category for the external hardware cluster in this collection period can be determined as abnormal operation risk.

[0122] Step 603: Report an external hardware cluster failure message based on the identity information of the offline external hardware nodes that occurred during at least one acquisition period within the third target time period. This can be achieved by the evaluation module collecting the identity information of the offline external hardware nodes during at least one acquisition period within the third target time period, and then having the processing module report the external hardware failure message.

[0123] When the processing module reports external hardware failure messages, it can directly report the identity information of offline external hardware nodes for at least one acquisition cycle within the third target time period. Alternatively, it can first aggregate the identity information of offline external hardware nodes from multiple acquisition cycles within the third target time period, that is, merge the same offline external hardware node identity information from different acquisition cycles to ultimately form the offline external hardware node identity information of the computing node within the third target time period. Of course, it can also collect auxiliary information such as whether each offline external hardware node went offline in different acquisition cycles and the duration of the offline status. Finally, the offline external hardware node identity information and auxiliary offline information of the computing node within the third target time period are packaged into an external hardware failure message, which is then reported by the processing module.

[0124] In practical applications, the third target time period of the exemplary embodiments of this disclosure includes one or more collection cycles. If the fault warning category of the computing node in a certain collection cycle included in the third target time period is abnormal operation risk, the processing module can directly report the external hardware cluster fault message. However, in order to reduce misjudgment caused by network jitter or other random factors, when the third target time period includes multiple collection cycles, it can be determined whether the number of external hardware nodes disconnected in at least two collection cycles included in the third target time period is greater than 0 and less than or equal to the preset maximum number of disconnected nodes. If so, the fault warning category of the computing node can be determined to be abnormal operation risk, and the external hardware cluster fault message can be reported based on the identity information of the disconnected external hardware nodes.

[0125] Step 604: Migrate all virtual machines on the compute node to the target compute node. If the fault warning category of the compute node in a certain collection period included in the third target time period is a potential fault risk, all virtual machines on the compute node can be directly migrated to the target compute node. However, in order to reduce false judgments caused by network jitter or other random factors, when the third target time period includes multiple consecutive collection periods, it can be determined whether the number of external hardware nodes that have gone offline on the compute node in multiple consecutive collection periods is greater than the preset maximum number of offline nodes. If so, all virtual machines on the compute node are directly migrated to the target compute node.

[0126] To further ensure the reliability and accuracy of virtual machine management, all virtual machines on the compute node are migrated to the target compute node. This can include: if the number of external hardware nodes that fail during most of the data collection periods within the third target time period increases over time, it indicates that the number of external hardware nodes failing is increasing, suggesting a high probability of external hardware cluster failure. In this case, all virtual machines on the compute node can be migrated to the target compute node. For example, the system disk and the corresponding XML files of the virtual machines can be used to redefine and restart the virtual machines on a healthy node.

[0127] To prevent misjudgments caused by network jitter or other random factors, a third target time period can be defined as including M collection cycles. The system can then determine whether the number of external hardware nodes that have lost connection within the compute node during the N collection cycles of the third target time period exceeds a preset maximum number of lost connections, where N represents an integer greater than M / 2. This constraint method ensures that the ultimately determined fault warning type for the external hardware cluster is relatively accurate, thereby improving the reliability and accuracy of virtual machine management.

[0128] For example, when M=3 and N=2, the fault warning type of the external hardware cluster in each collection cycle can be collected during the third target time period. If the external hardware cluster has two or three collection cycles with abnormal operation risk in the fault warning type during the third target time period, the external hardware cluster fault message can be reported based on the identity information of the offline external hardware nodes in each collection cycle included in the third target time period. If the external hardware cluster has two or three collection cycles with potential fault risk in the fault warning type during the third target time period, and the number of offline external hardware nodes increases with the extension of time, all virtual machines of the computing node can be migrated to the target computing node.

[0129] In one alternative approach, considering that even when the compute node and its corresponding virtual machines are functioning normally, an anomaly in the service components of the compute node could still lead to high security vulnerabilities or risks. In this case, although the virtual machines corresponding to the compute node and their running business systems are unaffected, they face high security risks; therefore, the virtual machines corresponding to the compute node may fail at any time. Based on this, the exemplary embodiments of this disclosure may further include: a service component failure early warning and troubleshooting method.

[0130] Figure 7 A schematic diagram illustrating the service component fault warning and troubleshooting process of an exemplary embodiment of this disclosure is shown. For example... Figure 7 As shown, the service component fault warning and troubleshooting method of this exemplary embodiment may include:

[0131] Step 701: Obtain the service component running information of the compute node in multiple collection cycles. For example, in each collection cycle, the running status of L3 Agent (neutron-l3-agent), neutron-metadata-agent, neutron-dhcp-agent, neutron-openvswitch-agent, neutron-openvswitch-agent, neutron-netns-cleanup, and compute services can be determined by using systemctl list-units--type=service--state=running.

[0132] Step 702: Determine the running status of the target service in each collection cycle based on the service component running information of the computing node in each collection cycle.

[0133] Step 703: If the same target service is in a non-running state during at least one collection cycle included in the fourth target time period, hot-migrate all virtual machines on the compute node to the target compute node. The fourth target time period may include one or more consecutive collection cycles.

[0134] To verify whether the service is abnormal, when the service component operation information of the computing node in each collection cycle includes the operation information of multiple services in each collection cycle, the method of the exemplary embodiment of this disclosure may further include: if the same target service is in a non-operating state in a certain collection cycle, the target service can be restarted, and then the operation information of the restarted target service in the corresponding collection cycle can be updated. At this time, the operation state after restarting can be determined again based on the operation information of the restarted target service in the corresponding collection cycle.

[0135] If the target service remains in an inactive state, it can be determined that the target service is in an inactive state during the corresponding data collection period, and the fault warning category for the target service during the corresponding data collection period can be set as a potential fault risk. This process is essentially a way to verify the fault warning type of the target service by restarting it.

[0136] To reduce misjudgments caused by network jitter or other random factors, this exemplary embodiment can statistically analyze the fault warning categories of the same target service across multiple collection cycles within the fourth target time period. If the fault warning categories for the same target service across most collection cycles within the fourth target time period are potential fault risks, it indicates a significant risk to the virtual machine environment that cannot recover autonomously. Therefore, there is a security vulnerability on the compute node that could harm the virtual machines, and all virtual machines on the compute node can be migrated to the target compute node. This migration process can indiscriminately migrate the CPU, memory, and other components of the target virtual machines to the target compute node.

[0137] For example, the fourth target time period includes M consecutive collection cycles. The target service of the compute node is in a non-running state during N target collection cycles within the fourth target time period. All virtual machines on the compute node can be hot-migrated to the target compute node. N represents an integer greater than M / 2. For instance, when M=3 and N=2, if a service's fault warning category is detected as "potential fault risk" for the first time in the corresponding collection cycle, the fault warning category for that service can be detected in the subsequent two collection cycles. If the service's fault warning category is "potential fault risk" in at least one of the subsequent two collection cycles, all virtual machines on the compute node can be hot-migrated to the target compute node.

[0138] In one alternative embodiment, when at least two of the following fault warnings are received simultaneously: a fault warning for the target virtual machine, a fault warning for the target service, and a fault warning for the target physical hardware, the present exemplary embodiment may select one fault warning to manage the virtual machine of the compute node. Alternatively, it may randomly configure the fault warning priorities of at least two fault warnings and then manage the virtual machine of the compute node according to the randomly configured fault warning priorities of at least two fault warnings. Or, it may manage the virtual machine of the compute node according to a preset fault warning priority order.

[0139] For example, in the exemplary embodiments of this disclosure, the fault warning priority can be that the fault warning priority of the target virtual machine is higher than that of the target physical hardware, and the fault warning priority of the target physical hardware is higher than that of the target service. In this case, the service impact caused by directly migrating the virtual machine when the target service is detected to be abnormal but the virtual machine is normal can be avoided.

[0140] When fault warnings for the target virtual machine, the target service, and the target physical hardware are received simultaneously, the virtual machines on the compute node can be managed first based on the fault warning category of the target virtual machine, then based on the fault warning category of the target physical hardware, and finally based on the fault warning category of the target service.

[0141] In addition, considering that when managing virtual machines on a compute node based on the fault warning category of the target virtual machine, it is possible that the management of virtual machines on a compute node based on the fault warning category of the target physical hardware and the management of virtual machines on a compute node based on the fault warning category of the target service have been indirectly implemented, the exemplary embodiments of this disclosure may also decide whether to perform virtual machine management operations corresponding to lower priorities according to the actual situation.

[0142] In one or more technical solutions provided in the exemplary embodiments of this disclosure, when a target virtual machine is in a crash state during the first target data collection period, the target virtual machine can be managed based on the host machine's running state during the first target data collection period. Therefore, when a virtual machine in an exemplary embodiment of this disclosure is in a crash state during a certain data collection period, the host machine's running state during the same data collection period can be referenced to manage the crashed target virtual machine, thereby improving the accuracy of virtual machine fault diagnosis and troubleshooting speed, and avoiding the adverse effects of incorrect diagnosis on the virtual machine. Therefore, the method in the exemplary embodiments of this disclosure has high availability, can improve the continuity of system operation, and ensure the reliability and robustness of virtual machine management, thereby achieving the goal of improving the quality of operation and maintenance services.

[0143] Furthermore, the exemplary embodiments of this disclosure determine the running status of each virtual machine in different collection cycles by judging the virtual machine running information of the computing node in each collection cycle. Then, based on the running status of each virtual machine in different collection cycles, the target virtual machine in a crash state and the corresponding first target collection cycle are obtained. Then, the target virtual machine in a crash state is managed with reference to the host running status of the computing node in the first target collection cycle. It is evident that the method of the exemplary embodiments of this disclosure can specifically manage target virtual machines in a crash state without affecting other virtual machines running on the host machine. Therefore, when managing virtual machines, the exemplary embodiments of this disclosure can minimize the impact on business continuity.

[0144] In summary, the exemplary embodiments of this disclosure can collect three types of data through an agent module running on the computing node: the running information of service components, the running information of virtual machines (CPU, memory, network, storage, etc.) and the running information of physical hardware (CPU, memory, disk, network, external hardware node connectivity information, etc.) in each collection cycle. Then, the management module running on the management node can perform fault warnings on the three types of data respectively, and perform differentiated management of virtual machines based on the fault warning results.

[0145] The exemplary embodiments disclosed herein can differentiate virtual machine management methods based on the fault warning object and fault warning type, including operations such as partial migration of virtual machines on compute nodes, full migration of virtual machines on compute nodes, single virtual machine migration, and virtual machine restart.

[0146] Furthermore, the method of the exemplary embodiments of this disclosure can automatically realize the early warning and troubleshooting of physical hardware, virtual machines and service components of computing nodes, thereby improving the efficiency and productivity of operation and maintenance, reducing repetitive work, optimizing resource utilization, and saving time and costs. Moreover, the exemplary embodiments of this disclosure can comprehensively analyze similar data from multiple collection cycles (e.g., early warning of virtual machine overload, early warning of physical hardware failure and early warning of service component failure), or when a virtual machine is in a crash state in a certain collection cycle, refer to the host machine's operating status in the same collection cycle to achieve accurate analysis and prediction of different types of early warning of failure, thereby identifying and correcting potential problems in advance. Therefore, the method of the exemplary embodiments of this disclosure is conducive to improving the robustness of business systems, thereby improving the quality of operation and maintenance services.

[0147] Furthermore, when service components on compute nodes or certain virtual machines themselves fail or become unavailable, the processing module can automatically migrate virtual machines to available target compute nodes, ensuring service continuity and availability. Moreover, when a compute node fails, the target compute node can act as a backup node, quickly taking over the workload and reducing service interruption time, which is crucial for services and businesses with high availability requirements. Simultaneously, this troubleshooting method can analyze virtual machine resource usage based on virtual machine runtime information, rationally distributing the load across different compute nodes to prevent overload of any single compute node, thereby improving system performance.

[0148] Finally, when the method of the exemplary embodiments of this disclosure migrates the virtual machines of the current computing node to an available target computing node, it does not affect other virtual machines on the current computing node. This reduces system downtime, improves the efficiency of maintenance and upgrades of the computing node, and has better flexibility and operability, making it suitable for rapid recovery of system functions in the event of a catastrophic event.

[0149] It should be noted that when the first target time period, second target time period, third target time period, and fourth target time period included in the exemplary embodiments of this disclosure consist of multiple consecutive acquisition periods, the acquisition periods for which fault warnings occur are N. These N acquisition periods can be consecutive, partially consecutive, or all discontinuous. For example, M=3, N=2, and the M consecutive acquisition periods include the i-th acquisition period, the (i+1)-th acquisition period, and the (i+2)-th acquisition period, where i is an integer greater than or equal to 1. In this case, if the N acquisition periods are consecutive, the N acquisition periods may be the i-th acquisition period and the (i+1)-th acquisition period, or the (i+1)-th acquisition period and the (i+2)-th acquisition period, or the i-th acquisition period, the (i+1)-th acquisition period, and the (i+2)-th acquisition period; if the N acquisition periods are discontinuous, the N acquisition periods may be the i-th acquisition period and the (i+2)-th acquisition period.

[0150] The foregoing primarily describes the solutions provided by the embodiments of this disclosure from the perspective of the management node. It is understood that, in order to achieve the above functions, the management node includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0151] This embodiment of the disclosure can divide the management node into functional units according to the above method example. For example, it can divide each function into a separate functional module, or it can integrate two or more functions into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this embodiment of the disclosure is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.

[0152] By dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a virtual machine management device, which can be a management node or a chip applied to a management node. Figure 8 A schematic block diagram of the functional modules of a virtual machine management apparatus according to an exemplary embodiment of the present disclosure is shown. Figure 8 As shown, the virtual machine management device 800 includes:

[0153] The acquisition module 801 is used to acquire physical hardware operation information of the computing node and virtual machine operation information of the computing node in at least one acquisition cycle.

[0154] The determination module 802 is used to determine the host machine running status of the computing node in each collection cycle based on the virtual machine running information of the computing node in each collection cycle;

[0155] The management module 803 is used to manage the target virtual machine if it is determined from the virtual machine running information of the computing node in the first target collection period that the target virtual machine is in a crash state in the first collection period, based on the host running state of the computing node in the first target collection period.

[0156] As one possible implementation, the management module 803 is used to migrate the target virtual machine to the target computing node if the host machine running state of the computing node is overloaded during the first target acquisition period, and to restart the target virtual machine on the computing node if the host machine running state of the computing node is normal during the first target acquisition period.

[0157] As one possible implementation, the determining module 802 is further configured to determine the operating status of at least one virtual hardware device in each acquisition cycle based on the virtual machine running information of the computing node in each acquisition cycle;

[0158] The management module 803 is further configured to hot-migrate the target virtual machine to the target computing node if the same virtual hardware is in an overloaded operating state during at least one acquisition cycle in the first target time period, wherein the first target time period includes one or more consecutive acquisition cycles.

[0159] As one possible implementation, the physical hardware operation information includes: physical machine physical hardware operation information. The determining module 802 is used to determine the operation status of at least one physical hardware in each acquisition cycle based on the physical machine physical hardware operation information of the computing node in each acquisition cycle. If the operation status of the at least one physical hardware in the first target acquisition cycle is an overload operation state, the host machine operation status of the computing node in the first target acquisition cycle is confirmed to be an overload operation state.

[0160] As one possible implementation, the management module 803 is further configured to obtain the identity information of the target physical hardware in the second target acquisition cycle if the target physical hardware is in an overloaded state during the second target acquisition cycle included in the second target time period; and to manage the virtual machine corresponding to the target physical hardware based on the identity information acquisition method of the same target physical hardware in at least one second target acquisition cycle, wherein the second target time period includes one or more consecutive acquisition cycles.

[0161] As one possible implementation, the management module 803 is configured to, if the method of obtaining the identity information of the same target physical hardware in at least one second target acquisition cycle is the physical hardware basic information of the computing node, then hot-migrate some virtual machines of the computing node to the target computing node; if the method of obtaining the identity information of the same target physical hardware in at least one second target acquisition cycle is the host abnormal operation log of the computing node, then migrate all virtual machines of the computing node to the target computing node.

[0162] As one possible implementation, the second target time period includes multiple second target collection cycles. The management module 803 is used to determine the number of abnormal log records of the target physical hardware in each second target collection cycle based on the abnormal log records of the host machine of the computing node in the second target time period when the identity information is obtained through abnormal logs of the host machine of the computing node in the second target time period. If the number of abnormal log records of the same target physical hardware in multiple second target collection cycles in the second target time period increases with the extension of time, all virtual machines of the computing node are migrated to the target computing node.

[0163] As one possible implementation, the physical hardware operation information includes: connectivity information of the external hardware nodes included in the external hardware cluster of the computing node; the determining module 802 is further used to determine the number of external hardware nodes that the computing node has lost connection in multiple acquisition cycles based on the connectivity information of the external hardware nodes of the computing node in multiple acquisition cycles.

[0164] The management module 803 is further configured to migrate all virtual machines of the computing node to the target computing node if the number of external hardware nodes that have gone offline in at least one collection cycle included in the third target time period is greater than the preset maximum number of offline nodes; and to report an external hardware cluster fault message based on the identity information of the offline external hardware nodes of the computing node in at least one collection cycle included in the third target time period if the number of external hardware nodes that have gone offline in at least one collection cycle included in the third target time period is greater than 0 and less than or equal to the preset maximum number of offline nodes. The third target time period includes one or more consecutive collection cycles.

[0165] As one possible implementation, when the third target time period includes multiple consecutive acquisition cycles, the management module 803 is used to migrate all virtual machines of the computing node to the target computing node if the number of external hardware nodes that have gone offline during most of the acquisition cycles included in the third target time period increases with the extension of time.

[0166] As one possible implementation, the acquisition module 801 is also used to acquire the service component operation information of the computing node in multiple acquisition cycles;

[0167] The determining module 802 is further configured to determine the running status of the target service in each collection cycle based on the service component running information of the computing node in each collection cycle;

[0168] The management module 803 is further configured to hot migrate all virtual machines of the computing node to the target computing node if the same target service is in a non-running state in at least one collection cycle included in the fourth target time period, wherein the fourth target time period includes multiple consecutive collection cycles.

[0169] As one possible implementation, the service component operation information of the computing node in each collection cycle includes the operation information of multiple services in each collection cycle. The determining module 802 is further configured to restart the target service if the target service is in a non-running state in the collection cycle; and update the operation information of the restarted target service in the corresponding collection cycle.

[0170] Figure 9 A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown. Figure 9 As shown, the chip 900 includes one or more (including two) processors 901 and a communication interface 902. The communication interface 902 can support the management node in performing the data transmission and reception steps in the above method, and the processor 901 can support the management node in performing the data processing steps in the above method.

[0171] Optional, such as Figure 9 As shown, the chip 900 also includes a memory 903, which may include read-only memory and random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include non-volatile random access memory (NVRAM).

[0172] In some implementations, such as Figure 9 As shown, processor 901 executes corresponding operations by calling operation instructions stored in memory (which may be stored in the operating system). Processor 901 controls the processing operations of any terminal device; processor can also be called a central processing unit (CPU). Memory 903 may include read-only memory and random access memory, and provides instructions and data to processor 901. A portion of memory 903 may also include NVRAM. For example, in applications, memory, communication interfaces, and other components are coupled together via a bus system, which may include, in addition to a data bus, a power bus, a control bus, and a status signal bus, etc. However, for clarity, in... Figure 9 The general designated all buses as Bus System 904.

[0173] The methods disclosed in the embodiments of this disclosure can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0174] Exemplary embodiments of this disclosure also provide a cloud computing platform, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, causing the electronic device to perform a method according to an embodiment of this disclosure.

[0175] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.

[0176] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.

[0177] refer to Figure 10 The present invention describes a structural block diagram of an electronic device 1000 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0178] like Figure 10 As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of the device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0179] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, output unit 1007, storage unit 1008, and communication unit 1009. Input unit 1006 can be any type of device capable of inputting information to electronic device 1000. Input unit 1006 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 1007 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1008 may include, but is not limited to, disk and optical disk. Communication unit 1009 allows electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0180] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above. For example, in some embodiments, the methods of exemplary embodiments of this disclosure can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1000 via ROM 1002 and / or communication unit 1009. In some embodiments, the computing unit 1001 can be configured to perform the methods of exemplary embodiments of this disclosure by any other suitable means (e.g., by means of firmware).

[0181] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0182] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0183] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0184] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0185] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0186] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this disclosure are performed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0187] Although this disclosure has been described in conjunction with specific features and embodiments, it will be apparent that various modifications and combinations can be made therein without departing from the spirit and scope of this disclosure. Accordingly, this specification and drawings are merely exemplary illustrations of the disclosure as defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this disclosure. It is obvious that those skilled in the art can make various alterations and modifications to this disclosure without departing from its spirit and scope. Thus, this disclosure is also intended to include any such modifications and modifications that fall within the scope of the claims of this disclosure and their equivalents.

Claims

1. A virtual machine management method, characterized in that, include: Obtain physical hardware operation information of the computing node and virtual machine operation information of the computing node in at least one acquisition cycle; The host machine running status of the computing node in each acquisition cycle is determined based on the physical hardware operation information of the computing node in each acquisition cycle. If, based on the virtual machine running information of the computing node in the first target acquisition period, it is determined that the target virtual machine is in a crash state during the first target acquisition period, then based on the host machine running status of the computing node in the first target acquisition period, the target virtual machine is managed; specifically, If the host machine of the computing node is in an overloaded state during the first target acquisition period, the target virtual machine is migrated to the target computing node; if the host machine of the computing node is in a normal operating state during the first target acquisition period, the target virtual machine is restarted on the computing node. If the target physical hardware is in an overloaded state during the second target acquisition cycle within the second target time period, the identity information of the target physical hardware during the second target acquisition cycle is obtained. The second target time period includes one or more consecutive acquisition cycles. If the identity information of the same target physical hardware is obtained through the method of acquiring the basic physical hardware information of the computing node in at least one second target acquisition cycle, then some virtual machines of the computing node will be hot-migrated to the target computing node. If the identity information of the same target physical hardware is obtained through the abnormal operation log of the host machine of the computing node in at least one second target acquisition cycle, then all virtual machines of the computing node will be migrated to the target computing node.

2. The method according to claim 1, characterized in that, The method further includes: Based on the virtual machine running information of the computing node in each acquisition cycle, the running status of at least one virtual hardware device in each acquisition cycle is determined. If the same virtual hardware is in an overloaded state during at least one acquisition cycle in the first target time period, the target virtual machine will be hot-migrated to the target computing node. The first target time period includes one or more consecutive acquisition cycles.

3. The method according to claim 1, characterized in that, The physical hardware operation information includes: physical machine physical hardware operation information. Determining the host machine operation status of the computing node in each acquisition cycle based on the physical hardware operation information of the computing node in each acquisition cycle includes: Based on the physical hardware operation information of the computing node in each acquisition cycle, the operating status of at least one physical hardware in each acquisition cycle is determined. If the at least one physical hardware is in an overloaded operating state during the first target acquisition period, it is confirmed that the computing node is in an overloaded operating state during the host machine operation state during the first target acquisition period.

4. The method according to claim 1, characterized in that, The second target time period includes multiple second target acquisition cycles, and the migration of all virtual machines of the computing node to the target computing node includes: When the identity information is obtained through the abnormal operation log of the host machine of the computing node, the number of abnormal log records of the target physical hardware in each second target acquisition cycle is determined based on the abnormal operation log of the host machine of the computing node in each second target acquisition cycle during the second target time period. If the number of abnormal log records of the same target physical hardware increases over time in multiple second target acquisition cycles within the second target time period, all virtual machines of the computing node will be migrated to the target computing node.

5. The method according to any one of claims 1 to 4, characterized in that, The physical hardware operation information includes: connectivity information of the external hardware nodes in the external hardware cluster of the computing node; the method further includes: Based on the connectivity information of the external hardware nodes of the computing node in multiple acquisition cycles, determine the number of external hardware nodes that the computing node lost connection in multiple acquisition cycles. If the number of external hardware nodes that have gone offline in at least one collection cycle during the third target time period is greater than the preset maximum number of offline nodes, then all virtual machines of the computing node will be migrated to the target computing node. The third target time period includes one or more consecutive collection cycles. If the number of disconnected external hardware nodes of the computing node in at least one collection cycle included in the third target time period is greater than 0 and less than or equal to the preset maximum number of disconnected nodes, then the computing node reports an external hardware cluster fault message based on the identity information of the disconnected external hardware nodes in at least one collection cycle included in the third target time period.

6. The method according to claim 5, characterized in that, When the third target time period includes multiple consecutive collection cycles, migrating all virtual machines of the computing node to the target computing node includes: If the number of external hardware nodes that go offline during most of the acquisition cycles in the third target time period of the computing node increases over time, then all virtual machines of the computing node will be migrated to the target computing node.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain the service component operation information of the computing node in multiple collection cycles; The running status of the target service in each collection cycle is determined based on the service component running information of the computing node in each collection cycle; If the same target service is in a non-running state in at least one of the collection cycles included in the fourth target time period, all virtual machines of the computing node will be hot-migrated to the target computing node. The fourth target time period includes multiple consecutive collection cycles.

8. The method according to claim 7, characterized in that, The service component operation information of the computing node in each collection cycle includes the operation information of multiple services in each collection cycle, and the method further includes: If the target service is in a non-running state during the collection period, restart the target service; Update the running information of the target service after the restart in the corresponding collection period.

9. A virtual machine management device, characterized in that, include: The acquisition module is used to acquire physical hardware operation information of the computing node and virtual machine operation information of the computing node in at least one acquisition cycle. The determination module is used to determine the host machine running status of the computing node in each collection cycle based on the virtual machine running information of the computing node in each collection cycle; The management module is used to manage the target virtual machine based on the virtual machine running information of the computing node in the first target collection period, if it is determined that the target virtual machine is in a crash state in the first target collection period, and based on the host machine running status of the computing node in the first target collection period; specifically, If the host machine of the computing node is in an overloaded state during the first target acquisition period, the target virtual machine is migrated to the target computing node; if the host machine of the computing node is in a normal operating state during the first target acquisition period, the target virtual machine is restarted on the computing node. If the target physical hardware is in an overloaded state during the second target acquisition cycle within the second target time period, the identity information of the target physical hardware during the second target acquisition cycle is obtained. The second target time period includes one or more consecutive acquisition cycles. If the identity information of the same target physical hardware is obtained through the method of acquiring the basic physical hardware information of the computing node in at least one second target acquisition cycle, then a portion of the virtual machines of the computing node will be hot-migrated to the target computing node. If the identity information of the same target physical hardware is obtained through the abnormal operation log of the host machine of the computing node in at least one second target acquisition cycle, then all virtual machines of the computing node will be migrated to the target computing node.

10. A cloud computing platform, characterized in that, include: processor; as well as, Memory for stored programs; The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Fault tolerance method and system of virtual machines

    CN102455951A

  • Virtual machine high availability management method and system and storage medium

    CN111880906A