A Method, Device and Medium for Preventing Brain Split of Virtual Machines in a Cloud Platform
By obtaining computing node status information in the cloud computing platform, judging faults and creating snapshots, the split brain problem of virtual machines when computing node failure is solved, data corruption is avoided, and data availability and integrity are achieved.
Patent Information
- Application Number
- CN202210094309.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-01-26
AI Technical Summary
In the cloud computing platform, virtual machines may have brain splits when the computing node fails, resulting in file system corruption in the case of excessive writes and corruption in virtual machine data.
By obtaining the status information of the computing node, we judge whether there is a failure. If it is a failure, the backend storage is called to create a snapshot of the virtual machine's storage volume, retain the data of the original virtual machine, and send the snapshot to the new computing node, so that the new virtual machine can restore the data of the original virtual machine.
It avoids excessive writing of virtual machine data, protects the availability of virtual machine data, and prevents data corruption caused by split brain.
Smart Images

Figure CN114461341B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and in particular to a method, device, and medium for preventing brain split of virtual machines in a cloud platform. Background Art
[0002] Among the characteristics of a cloud computing platform, high availability of cloud hosts (virtual machines) is an effective way to ensure uninterrupted service of cloud host services and can improve service quality. When the host computer of the computing node where the virtual machine is located fails, the virtual machine can be restored on another computing node. However, the failed computing node does not delete the virtual machine therein or ensure that the virtual machine is shut down, resulting in the virtual machine data logic in the two computing nodes being the same, that is, the virtual machine experiences a brain split. When a virtual machine experiences a brain split, the computing nodes will perform "multiple writes" on the virtual machine data. "Multiple writes" refers to the situation where multiple computing nodes simultaneously read and write to the storage volume of a virtual machine, which will cause damage to the file system and damage to the virtual machine data.
[0003] In view of the above problems, designing a method for preventing brain split of virtual machines in a cloud platform to prevent damage to virtual machine data caused by "multiple writes" is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, and medium for preventing brain split of virtual machines in a cloud platform to prevent damage to virtual machine data caused by "multiple writes".
[0005] To solve the above technical problems, this application provides a method for preventing brain split of virtual machines in a cloud platform, which is applied to a cloud platform. The method includes:
[0006] Obtain the status information of the computing node;
[0007] Judge whether the computing node has failed according to the status information;
[0008] If so, call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine;
[0009] Send the snapshot to a new computing node for the new computing node to access the snapshot to restore the data of the original virtual machine for the new virtual machine.
[0010] Preferably, after sending the snapshot to the new computing node, it further includes:
[0011] Detect the data of the original virtual machine and the data of the new virtual machine respectively to obtain a detection result;
[0012] Delete the virtual machines that do not meet the preset requirements according to the detection result.
[0013] Preferably, the status information is network information and heartbeat data information;
[0014] Wherein, the heartbeat data information is a timestamp regularly written to a heartbeat data disk, and the heartbeat data disk is arranged in the backend storage.
[0015] Preferably, the judging whether a failure occurs according to the status information includes:
[0016] Judging whether the computing node fails according to the network information;
[0017] If it is determined according to the network information that the computing node does not fail, it is determined that the computing node does not fail;
[0018] If it is determined according to the network information that the computing node fails, a request for updating the heartbeat data information is sent to update the heartbeat data information;
[0019] Read the new heartbeat data information;
[0020] Judge whether the new heartbeat data information meets a preset condition;
[0021] If not, it is determined that the computing node fails;
[0022] If so, it is determined that the computing node does not fail.
[0023] Preferably, after deleting the virtual machines that do not meet the preset requirements according to the detection result, it further includes:
[0024] Output a fault message to the computing node to prompt the user that the computing node fails.
[0025] Preferably, the preset condition is that the time intervals of the timestamps are equal.
[0026] To solve the above technical problems, the present application further provides a method for preventing brain split of virtual machines in a cloud platform, which is applied to a computing node, and the method includes:
[0027] Receive an opening instruction sent by the cloud platform to start the virtual machine;
[0028] Receive a closing instruction sent by the cloud platform to close the virtual machine;
[0029] Wherein, the closing instruction is sent after the cloud platform judges that the original computing node fails according to the status information, the cloud platform calls the backend storage to create a snapshot of the storage volume of the virtual machine, and sends the snapshot to a new computing node.
[0030] To solve the above technical problems, the present application also provides a cloud platform virtual machine anti-split brain device, including:
[0031] An acquisition module, configured to acquire the status information of a computing node;
[0032] A judgment module, configured to judge whether the computing node has a fault according to the status information; if so, trigger a call module;
[0033] The call module is configured to call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine;
[0034] A sending module, configured to send the snapshot to a new computing node for the new computing node to access the snapshot to enable the new virtual machine to restore the data of the original virtual machine.
[0035] To solve the above technical problems, the present application also provides another cloud platform virtual machine anti-split brain device, including:
[0036] A memory, configured to store a computer program;
[0037] A processor, configured to implement the steps of the above-mentioned cloud platform virtual machine anti-split brain method when executing the computer program.
[0038] To solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned cloud platform virtual machine anti-split brain method are implemented.
[0039] The cloud platform virtual machine anti-split brain method provided by the present application obtains the status information of a computing node, judges whether the computing node has a fault according to the status information; if so, calls the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine, and sends the snapshot to a new computing node for the new computing node to access the snapshot to enable the new virtual machine to restore the data of the original virtual machine. It can be seen that in the above technical solution, when the computing node has a fault, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine of the new computing node, so that the data logic of the original virtual machine is divided into two parts, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thereby protecting the availability of the virtual machine data.
[0040] In addition, the present application also provides a cloud platform virtual machine anti-split brain device and a computer-readable storage medium, and the effect is the same as above. Description of the Drawings
[0041] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of a method for preventing brain split of virtual machines in a cloud platform provided by an embodiment of the present application;
[0043] Figure 2 It is a flowchart of another method for preventing brain split of virtual machines in a cloud platform provided by an embodiment of the present application;
[0044] Figure 3 It is a schematic structural diagram of a device for preventing brain split of virtual machines in a cloud platform provided by an embodiment of the present application;
[0045] Figure 4 It is a schematic structural diagram of another device for preventing brain split of virtual machines in a cloud platform provided by an embodiment of the present application. Detailed implementation manners
[0046] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0047] The core of the present application is to provide a method, device and medium for preventing brain split of virtual machines in a cloud platform to prevent data damage to virtual machines caused by "multiple writes".
[0048] To enable those skilled in the art to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0049] A virtual machine refers to a complete computer system with the functions of a complete hardware system simulated by software and running in a completely isolated environment. Among the features of a cloud computing platform that can manage multiple virtual machines, high availability of cloud hosts (virtual machines) is an effective way to ensure uninterrupted services for cloud host services and can improve service quality. High availability means that after the computing node where the virtual machine is located fails, through a certain judgment logic and operation, the virtual machine is restarted from another computing node to continue providing services. Specifically, when the computing node where the virtual machine is located fails, the virtual machine can be restored on another computing node. However, the failed computing node does not delete the virtual machine in it or ensure that the virtual machine is shut down, resulting in the virtual machine data logic in the two computing nodes being the same, that is, the situation of virtual machine split-brain occurs. When the virtual machine has a split-brain, the computing nodes will perform "multiple writes" on the virtual machine data. "Multiple writes" refers to the situation where multiple computing nodes simultaneously read and write to the storage volume of a virtual machine, which will cause damage to the file system and damage to the virtual machine data. Therefore, to solve the above problems, the present application provides a method for preventing virtual machine split-brain in a cloud platform.
[0050] Figure 1 The flowchart of a method for preventing virtual machine split-brain in a cloud platform provided by an embodiment of the present application. Applied to a cloud platform, as Figure 1 shown, the method includes:
[0051] S10: Obtain the status information of the computing node.
[0052] S11: Determine whether the computing node has failed according to the status information; if so, enter step S12.
[0053] S12: Call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine.
[0054] S13: Send the snapshot to the new computing node for the new computing node to access the snapshot to restore the data of the original virtual machine.
[0055] In a specific implementation, the cloud platform manages the operation of multiple computing nodes and their virtual machines; the cloud platform masters the current status of each computing node by obtaining the status information of each computing node. For example, network information, running time information, and some other information. In this embodiment, there is no limitation on the specific content of the status information, which depends on the specific implementation situation. During the operation of the virtual machine, the cloud platform determines whether the computing node has failed according to the status information. There is no limitation on the specific logic for determining a failure in this embodiment, which depends on the specific implementation situation.
[0056] When the cloud platform determines that a certain computing node has failed based on the status information, it calls the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node. Here, the backend storage refers to the server or storage system that actually stores data and is used to store the data of numerous computing nodes and their virtual machines; at the same time, it can perform operations on the storage volume of the virtual machine such as snapshot, replication, creation, and deletion. Specifically, the backend storage creates a snapshot of the storage volume of the original virtual machine of the failed computing node; a snapshot refers to an instantaneous copy of the virtual machine disk file (VMDK) at a certain point; when the system crashes or is abnormal, the disk file system and system storage can be maintained by restoring to the snapshot. Therefore, the data of the original virtual machine of the failed computing node can be retained. Then, the snapshot is sent to a new computing node, that is, a non-failed computing node, so that the new computing node can access the obtained snapshot and allocate the snapshot to the new virtual machine therein, ultimately realizing the restoration of the data of the original virtual machine.
[0057] It should be noted that although the failed computing node has been determined to be faulty, in fact, limited by the high-availability judgment logic, it cannot fully guarantee that the original virtual machine can no longer be read and written. Even with a fencing mechanism, that is, by means of the Intelligent Platform Management Interface (IPMI) or Secure Shell (SSH) and other means to shut down the original computing node to shut down the virtual machine, it is difficult to ensure that the original virtual machine can be shut down, because there may be a situation of complete network isolation or the physical machine of the computing node is in a false dead state and cannot be processed. After the backend storage takes a snapshot of the data of the source virtual machine and provides the snapshot to the new virtual machine, the data of the virtual machine can be logically divided into two parts, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thus protecting the data availability.
[0058] In this embodiment, by obtaining the status information of the computing node, it is determined whether the computing node has failed according to the status information; if so, the backend storage is called to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine, and the snapshot is sent to the new computing node for the new computing node to access the snapshot to restore the data of the original virtual machine to the new virtual machine. It can be seen that in the above technical solution, when the computing node fails, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine of the new computing node, so that the data of the original virtual machine is logically divided into two parts, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thus protecting the availability of the virtual machine data.
[0059] Figure 2 It is a flowchart of another method for preventing a cloud platform virtual machine from brain split provided by an embodiment of the present application. As Figure 2As shown, in order to accurately obtain the data of the virtual machine, after sending the snapshot to the new computing node, it further includes:
[0060] S14: Detect the data of the original virtual machine and the data of the new virtual machine respectively to obtain a detection result.
[0061] S15: Delete the virtual machine that does not meet the preset requirements according to the detection result.
[0062] It can be understood that after high availability is triggered and the virtual machine of the failing computing node runs on the new computing node, the virtual machine can run on the new computing node and provide services externally at the same time; since the data logics of the two virtual machines are in two copies, whether the original virtual machine runs and reads / writes on the failing computing node has no impact on the new virtual machine and will not cause data damage to it; however, in the actual implementation process, it is still necessary to determine which virtual machine's data to use. By actually detecting and judging the data of the original virtual machine and the new virtual machine, a detection result is obtained to decide which data to use. Because whether it is the original virtual machine or the new virtual machine, their network information is exactly the same, so there will be no situation where business data is written into the two virtual machines simultaneously; then, according to the detection result, it is judged which virtual machine to retain, and the virtual machine that does not meet the preset requirements, that is, the virtual machine that does not need to be retained, is deleted. The preset requirements are not limited in this embodiment and are determined according to the specific implementation situation.
[0063] In this embodiment, in order to accurately obtain the data of the virtual machine, after high availability is triggered, the data of the original virtual machine and the data of the new virtual machine are detected respectively to obtain a detection result, and the virtual machine that does not meet the preset requirements is deleted according to the detection result, and finally the accurate virtual machine data is retained.
[0064] Based on the above embodiment:
[0065] As a preferred embodiment, the status information is network information and heartbeat data information;
[0066] Among them, the heartbeat data information is the timestamp regularly written into the heartbeat data disk, and the heartbeat data disk is set in the backend storage.
[0067] It can be understood that the triggering method of high availability is usually determined by the network information of the computing nodes. Therefore, the status information includes network information. In order to improve the accuracy of the judgment logic for triggering high availability, it is also necessary to determine based on the heartbeat data information in the status information. The heartbeat data information is the timestamp regularly written to the heartbeat data disk to achieve regular recording. The specific time interval for timing is not limited in this embodiment and is determined according to the specific implementation situation. The heartbeat data disk is set in the backend storage and is used to store the heartbeat data information. In this embodiment, the specific process of judging whether the calculation is faulty according to the status information, that is, according to the network information and the heartbeat data information, is not limited and is determined according to the specific implementation situation.
[0068] In this embodiment, the status information is network information and heartbeat data information, which improves the accuracy of the judgment logic for triggering high availability.
[0069] Based on the above embodiments:
[0070] As a preferred embodiment, judging whether a failure has occurred according to the status information includes:
[0071] Judging whether the computing node has failed according to the network information;
[0072] If it is determined according to the network information that the computing node has not failed, it is determined that the computing node has not failed;
[0073] If it is determined according to the network information that the computing node has failed, a request to update the heartbeat data information is sent to update the heartbeat data information;
[0074] Read the new heartbeat data information;
[0075] Judge whether the new heartbeat data information meets the preset conditions;
[0076] If not, it is determined that the computing node has failed;
[0077] If so, it is determined that the computing node has not failed.
[0078] In the above embodiments, the status information includes network information and heartbeat data information; the specific steps for determining whether a computing node fails based on the status information are as follows: First, determine whether the computing node fails based on the network information. If it is determined that no failure has occurred, it is determined that the computing node has not failed. If it is determined based on the network information that the computing node has failed, a request to update the heartbeat data information is sent, and the cloud platform processes this request with the highest priority, updates the heartbeat information to the heartbeat data disk, and updates the heartbeat data; reads the new heartbeat data information and determines whether it meets the preset conditions. If it meets the conditions, it is determined that the computing node has not failed. If it does not meet the conditions, it is determined that the computing node has failed. In this embodiment, there is no limitation on the preset conditions, which are determined according to the specific implementation situation.
[0079] In this embodiment, based on the determination of failure according to the network information, the heartbeat data information is updated, the new heartbeat data information is read, and whether the computing node fails is re-determined by judging whether it meets the preset conditions, so as to avoid misjudgment of high availability triggered by untimely heartbeat writing caused by excessive input / output pressure.
[0080] As Figure 2 shown, after deleting virtual machines that do not meet the preset requirements according to the detection results, it further includes:
[0081] S16: Output the fault information to the computing node to prompt the user that the computing node has failed.
[0082] It can be understood that after the cloud platform determines that the computing node has failed through the status information of the computing node, it calls the backend storage to create a snapshot of the storage volume of the original virtual machine and send it to the new virtual machine of the new computing node; after detecting the data of the original virtual machine and the new virtual machine, the virtual machines that do not meet the preset requirements are deleted; at this time, the failed computing node no longer undertakes the task of running the virtual machine. In order to enable the user to determine the failed computing node, the fault information is output to the computing node, enabling the user to clarify the specific fault information, so as to maintain and debug the failed computing node and restore its ability to run virtual machines normally.
[0083] In this embodiment, by outputting the fault information to the task node, it is realized to prompt the user that the computing node has failed for maintenance and debugging.
[0084] Based on the above embodiments:
[0085] As a preferred embodiment, the preset condition is that the time intervals of the timestamps are equal.
[0086] In the above embodiments, there is no limitation on the preset conditions for the new heartbeat data information, which depends on the specific implementation situation. In this embodiment, as a preferred embodiment, the preset condition is that the time intervals of the timestamps of the new heartbeat data information are equal. Finally, the misjudgment caused by the high-availability logic trigger is prevented.
[0087] Based on the above embodiments, the present application further provides another method for preventing a cloud platform virtual machine from splitting its brain, which is applied to a computing node. The method includes:
[0088] Receiving an opening instruction sent by the cloud platform to start the virtual machine;
[0089] Receiving a closing instruction sent by the cloud platform to shut down the virtual machine;
[0090] Wherein, the instruction is sent after the cloud platform determines that the original computing node has failed according to the status information, the cloud platform calls the backend storage to create a snapshot of the storage volume of the virtual machine, and sends the snapshot to the new computing node.
[0091] It can be understood that the computing node is used to run the virtual machine, and the instruction for controlling the virtual machine running on the computing node is sent by the cloud platform. Specifically, receive the opening instruction sent by the cloud platform to start the virtual machine. When the cloud platform determines that the current computing node has failed according to the status information, the cloud platform calls the backend storage to create a snapshot of the storage volume of the virtual machine, sends the snapshot to the new computing node, and sends a closing instruction to shut down the original computing node where the failure occurred, thereby shutting down the virtual machine.
[0092] Wherein, the cloud platform obtains the status information of the computing node, and determines whether the computing node has failed according to the status information; if so, calls the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine, and sends the snapshot to the new computing node for the new virtual machine on the new computing node to access the snapshot to restore the data of the original virtual machine. It can be seen from the above technical solution that when the computing node fails, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine on the new computing node, so that the data logic of the original virtual machine is divided into two parts, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thereby protecting the availability of the virtual machine data.
[0093] In the above embodiments, the method for preventing a cloud platform virtual machine from splitting its brain is described in detail. The present application also provides corresponding embodiments of the device for preventing a cloud platform virtual machine from splitting its brain. It should be noted that the present application describes the embodiments of the device part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware structure.
[0094] Figure 3 This is a schematic structural diagram of a device for preventing a cloud platform virtual machine from splitting its brain provided by an embodiment of the present application. AsFigure 3 As shown in the figure, the cloud platform virtual machine anti-brain-split device includes:
[0095] An acquisition module 10, configured to acquire the status information of a computing node.
[0096] A judgment module 11, configured to judge whether the computing node fails according to the status information; if so, trigger a call to a call module 12.
[0097] A call module 12, configured to call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node, so as to retain the data of the original virtual machine.
[0098] A sending module 13, configured to send the snapshot to a new computing node, so that the new computing node can access the snapshot to restore the data of the original virtual machine for the new virtual machine.
[0099] The cloud platform virtual machine anti-brain-split device provided in this embodiment acquires the status information of the computing node, judges whether the computing node fails according to the status information; if so, calls the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node, so as to retain the data of the original virtual machine, and sends the snapshot to a new computing node, so that the new computing node can access the snapshot to restore the data of the original virtual machine for the new virtual machine. It can be seen from the above technical solution that when the computing node fails, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine of the new computing node, so that the data logic of the original virtual machine is divided into two copies, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thereby protecting the availability of the virtual machine data.
[0100] Figure 4 It is a schematic structural diagram of another cloud platform virtual machine anti-brain-split device provided in an embodiment of the present application. As Figure 4 shown, the cloud platform virtual machine anti-brain-split device includes:
[0101] A memory 20, configured to store a computer program.
[0102] A processor 21, configured to implement the steps of the cloud platform virtual machine anti-brain-split method mentioned in the above embodiment when executing the computer program.
[0103] The cloud platform virtual machine anti-brain-split device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a notebook computer, or a desktop computer, etc.
[0104] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0105] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the cloud platform virtual machine anti-brain-split method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the cloud platform virtual machine anti-brain-split method.
[0106] In some embodiments, the cloud platform virtual machine anti-brain-split device may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.
[0107] Those skilled in the art can understand that Figure 4 the structure shown in
[0108] The anti-brain-split device for cloud platform virtual machines provided in this embodiment includes a memory and a processor. The processor is used to implement the steps of the anti-brain-split method for cloud platform virtual machines mentioned in the above embodiment when executing a computer program. By obtaining the status information of the computing node, it is determined whether the computing node has failed according to the status information; if so, the backend storage is called to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine, and the snapshot is sent to a new computing node for the new virtual machine of the new computing node to access the snapshot to restore the data of the original virtual machine. It can be seen that in the above technical solution, when the computing node fails, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine of the new computing node, so that the data logic of the original virtual machine is divided into two copies, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thereby protecting the availability of virtual machine data.
[0109] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps recorded in the above method embodiments (which can be the method corresponding to the cloud platform side, the method corresponding to the computing node side, or the method corresponding to both the cloud platform side and the computing node side).
[0110] It can be understood that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0111] This embodiment provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps recorded in the above method embodiment. By obtaining the status information of the computing node, it is determined whether the computing node fails according to the status information; if so, the backend storage is called to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine, and the snapshot is sent to a new computing node for the new computing node to access the snapshot to restore the data of the original virtual machine to the new virtual machine. It can be seen that in the above technical solution, when the computing node fails, the snapshot executed by the backend storage on the original virtual machine is provided to the new virtual machine of the new computing node, so that the data logic of the original virtual machine is divided into two copies, and the two volumes can be read and written simultaneously, avoiding multiple writes to the same volume, thereby protecting the availability of the virtual machine data.
[0112] The above has introduced in detail a method, device and medium for preventing brain split of virtual machines in a cloud platform provided by the present application. The embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0113] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A method for preventing brain split of virtual machines in a cloud platform, characterized in that, applied to a cloud platform, the method includes: Obtain the status information of the computing node; the status information is network information and heartbeat data information; wherein, the heartbeat data information is a timestamp regularly written to the heartbeat data disk, and the heartbeat data disk is set in the backend storage; the backend storage is a server or a storage system for storing data of multiple computing nodes and their virtual machines; Judge whether the computing node has a fault according to the status information; If so, call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine; wherein, the original virtual machine and the new virtual machine in the new computing node are allowed to read and write the corresponding storage volume simultaneously; Send the snapshot to the new computing node for the new computing node to access the snapshot to restore the data of the original virtual machine by the new virtual machine.
2. The method for preventing brain split of virtual machines in a cloud platform according to claim 1, characterized in that, after sending the snapshot to the new computing node, it further includes: Detect the data of the original virtual machine and the data of the new virtual machine respectively to obtain a detection result; Delete the virtual machine that does not meet the preset requirements according to the detection result.
3. The method for preventing brain split of virtual machines in a cloud platform according to claim 1, characterized in that, judging whether there is a fault according to the status information includes: Judge whether the computing node has a fault according to the network information; If it is determined according to the network information that the computing node has no fault, it is determined that the computing node has no fault; If it is determined according to the network information that the computing node has a fault, send a request to update the heartbeat data information for updating the heartbeat data information; Read the new heartbeat data information; Judge whether the new heartbeat data information meets the preset conditions; If not, it is determined that the computing node has a fault; If so, it is determined that the computing node has no fault.
4. The method for preventing brain split of virtual machines in a cloud platform according to claim 2, characterized in that, after deleting the virtual machine that does not meet the preset requirements according to the detection result, it further includes: Output a fault message to the computing node to prompt the user that the computing node has a fault.
5. The method for preventing brain split of virtual machines in a cloud platform according to claim 3, characterized in that, the preset condition is that the time intervals of the timestamps are equal.
6. A method for preventing brain split of virtual machines in a cloud platform, characterized in that, applied to a computing node, the method includes: Receive an opening instruction sent by the cloud platform to start the virtual machine; Receive a closing instruction sent by the cloud platform to close the virtual machine; Among them, the shutdown instruction is sent after the cloud platform determines that the original computing node has failed according to the status information, the cloud platform calls the backend storage to create a snapshot of the storage volume of the virtual machine, and sends the snapshot to the new computing node; the status information is network information and heartbeat data information; the heartbeat data information is the timestamp regularly written to the heartbeat data disk, and the heartbeat data disk is set in the backend storage; the backend storage is a server or a storage system for storing data of multiple computing nodes and their virtual machines; the virtual machine and the new virtual machine in the new computing node are allowed to read and write the corresponding storage volume simultaneously.
7. A cloud platform virtual machine anti-split brain device, characterized in that, it includes: An acquisition module, configured to acquire the status information of the computing node; the status information is network information and heartbeat data information; among them, the heartbeat data information is the timestamp regularly written to the heartbeat data disk, and the heartbeat data disk is set in the backend storage; the backend storage is a server or a storage system for storing data of multiple computing nodes and their virtual machines; A judgment module, configured to judge whether the computing node has failed according to the status information; if so, trigger the call module; The call module is configured to call the backend storage to create a snapshot of the storage volume of the original virtual machine of the corresponding computing node to retain the data of the original virtual machine; among them, the original virtual machine and the new virtual machine in the new computing node are allowed to read and write the corresponding storage volume simultaneously; A sending module, configured to send the snapshot to the new computing node for the new computing node to access the snapshot to enable the new virtual machine to restore the data of the original virtual machine.
8. A cloud platform virtual machine anti-split brain device, characterized in that, it includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the cloud platform virtual machine anti-split brain method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the cloud platform virtual machine anti-split brain method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Computer fault-tolerant method and computer fault-tolerant system
CN104391764A
Virtual machine anti-cerebral fissure management method and main server
CN110825487A