Virtual machine evacuation method and device
By obtaining the list of candidate hosts and sorting based on weights, determining the target computing node, and creating a target virtual machine on the node, the evacuation failure caused by resource conflicts during virtual machine migration is solved, and the success rate and efficiency of virtual machine evacuation is improved.
Patent Information
- Application Number
- CN202412000514.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
During the virtual machine migration process, you may encounter resource conflict problems, resulting in failure of evacuation of virtual machines and reducing the success rate of high availability of virtual machines.
By obtaining the candidate host list and sorting it based on the weight of the computing node, the target computing node to be migrated is determined. Create a target virtual machine on the target compute node and migrate the data of the source virtual machine to enable evacuation of the virtual machine. If the primary node fails, the next compute node is automatically selected from the candidate list and try again until successful or all nodes have tried.
By introducing resource weighted sorting and multi-try mechanisms, the problem of virtual machine evacuation failure is solved, and the efficiency and success rate of the virtual machine evacuation process is significantly improved.
Smart Images

Figure CN120045388A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of cloud computing technology, and in particular, to a virtual machine evacuation method and device. Background Art
[0002] The high availability of virtual machines is an important feature of cloud computing services. Even when a data center fails, the user's virtual machines can be quickly migrated to another healthy data center, ensuring continuous operation of the business, reducing the risk of data loss, and at the same time improving the stability and reliability of the entire system.
[0003] During the virtual machine migration process, resource problems may be encountered, resulting in migration failure. For example, when two virtual machine requests arrive at different nova-scheduler services simultaneously, the nova-scheduler service is very likely to select the same candidate host to evacuate these two virtual machines. At this time, the two virtual machines will conflict when applying for resources on the candidate host. If the resources on the candidate host are not sufficient to create two virtual machines, one of the virtual machines will fail to be evacuated, thus reducing the success rate of virtual machine high availability.
[0004] In response to the above problem, no effective solution has been proposed yet. Summary of the Invention
[0005] The embodiments of the present application provide a virtual machine evacuation method and device to at least solve the problem in the related technology that when two virtual machine requests select the same candidate host and the two virtual machines conflict when applying for resources on the candidate host, resulting in the failure of virtual machine evacuation.
[0006] According to an embodiment of the present application, a virtual machine evacuation method is provided. The method includes: in response to receiving a virtual machine evacuation request, obtaining a candidate host list, where the virtual machine evacuation request is used to indicate that a source computing node in a cloud platform fails, and a source virtual machine is carried on the source computing node, and the candidate host list includes multiple candidate computing nodes, and the multiple candidate computing nodes in the candidate host list are sorted based on the weights of the computing nodes; determining a first target computing node to be migrated from the multiple candidate computing nodes based on the candidate host list, where the first target computing node is the candidate computing node with the highest weight in the candidate host list; creating at least one target virtual machine on the first target computing node, and migrating the data stored in the source virtual machine to the target virtual machine to evacuate the source virtual machine.
[0007] In an exemplary embodiment, obtaining a list of candidate hosts includes: obtaining a plurality of first computing nodes; screening the plurality of first computing nodes to obtain a plurality of second computing nodes; selecting a plurality of candidate computing nodes from the plurality of second computing nodes based on a preset number of lists; and generating a list of candidate hosts based on the plurality of candidate computing nodes.
[0008] In an exemplary embodiment, screening the plurality of first computing nodes to obtain a plurality of second computing nodes includes: screening the plurality of first computing nodes based on a filter to obtain a plurality of third computing nodes, where the filter is used to screen out computing nodes that do not meet preset requirements, and the preset requirements include at least one of the following: the computing node has the network interfaces and bandwidth required by the virtual machine, the computing node complies with access control, firewall, and encryption standards, the computing node provides the required storage type and capacity, the computing node runs the required operating system and software packages, the computing node's hardware supports the graphics processor architecture and central processor architecture corresponding to the virtual machine; determining the weights of the plurality of third computing nodes respectively based on a weigher, where the weigher is used to calculate the weights from at least one of the following aspects: resource utilization rate, load status, geographical location, power status, workload balance between nodes; sorting the plurality of third computing nodes based on a preset sorting strategy and the weights of the plurality of third computing nodes to obtain a plurality of second computing nodes, where the preset sorting strategy is used to represent sorting in descending order based on the weights.
[0009] In an exemplary embodiment, the virtual machine evacuation method further includes: in response to a failure of a source computing node, generating a virtual machine evacuation request based on a high-availability mechanism.
[0010] In an exemplary embodiment, in response to a failure of a source computing node, generating a virtual machine evacuation request based on a high-availability mechanism includes: monitoring the status of any computing node in the cloud platform based on a first service component to obtain a status monitoring result; in response to the status monitoring result indicating that the source computing node has failed, identifying the source virtual machine running on the source computing node based on the high-availability mechanism to obtain an identification result; and generating a virtual machine evacuation request corresponding to any source virtual machine based on the identification result.
[0011] In an exemplary embodiment, the virtual machine evacuation method further includes: in response to a failure to create at least one target virtual machine on a first target computing node, determining a second target computing node from the list of candidate hosts, where the second target computing node is the next computing node of the first target computing node;
[0012] Creating at least one target virtual machine on the second target computing node.
[0013] In an exemplary embodiment, the virtual machine evacuation method further includes: in response to the failure of creating at least one target virtual machine on multiple candidate computing nodes, stopping the evacuation of the source virtual machine.
[0014] According to another embodiment of the present application, there is provided a virtual machine evacuation device, including: an acquisition module, configured to acquire a candidate host list in response to receiving a virtual machine evacuation request, where the virtual machine evacuation request is used to indicate that a source computing node in a cloud platform fails, and a source virtual machine is hosted on the source computing node, the candidate host list includes multiple candidate computing nodes, and the multiple candidate computing nodes in the candidate host list are sorted based on the weights of the computing nodes; a determination module, configured to determine a first target computing node to be migrated from the multiple candidate computing nodes based on the candidate host list, where the first target computing node is the candidate computing node with the highest weight in the candidate host list; a creation module, configured to create at least one target virtual machine on the first target computing node, and migrate the data stored in the source virtual machine to the target virtual machine to evacuate the source virtual machine.
[0015] According to still another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0016] According to still another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0017] According to still another embodiment of the present application, there is also provided a computer program product, including a computer program, where the computer program implements the steps in any one of the above method embodiments when executed by a processor.
[0018] Through this application, when a virtual machine evacuation request is received, candidate computing nodes are screened based on the resource information of the computing nodes, and a list containing multiple candidate computing nodes is returned. After receiving the list of candidate computing nodes returned by the scheduling service, an attempt is made to create a virtual machine on the first computing node in the list of candidate computing nodes. When the creation of the virtual machine fails on the first computing node, the control service automatically selects the next computing node from the list of candidate computing nodes and continues to attempt to create the virtual machine until the virtual machine is successfully created on any computing node in the list of candidate computing nodes or all computing nodes in the list have been tried. When screening candidate computing nodes, the scheduling service calculates a resource weight value based on the resource status of each computing node and sorts the candidate computing nodes according to the resource weight value to ensure an optimized sorting of the list of candidate computing nodes. Therefore, by introducing resource weight sorting and a multi-attempt mechanism, the problem in the related art where when two virtual machine requests screen out the same candidate hosts and resource requests for the two virtual machines conflict on the candidate hosts, resulting in the failure of virtual machine evacuation, can be solved, achieving the effect of significantly improving the efficiency and success rate of the virtual machine evacuation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a hardware structure block diagram of a server device for a virtual machine evacuation method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of the virtual machine high-availability evacuation process;
[0021] Figure 3 is a flowchart of the virtual machine evacuation method according to an embodiment of the present application;
[0022] Figure 4 is a flowchart of the virtual machine high-availability evacuation process based on the OpenStack cloud platform according to an embodiment of the present application;
[0023] Figure 5 is a block diagram of the structure of a virtual machine evacuation device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] For ease of understanding, some explanations of concepts related to the embodiments of the present invention are exemplarily given for reference. As follows:
[0025] OpenStack: An open-source cloud computing management platform project based on OpenStack, aiming to provide open-source software for the construction and management of public and private clouds, and is a combination of a series of open-source software projects.
[0026] The High Availability (HA) module generally refers to a software or hardware mechanism designed to ensure the uninterrupted operation of applications or services in the event of failures in a server cluster or virtualization environment.
[0027] 2) Virtual Machine HA: The high-availability mechanism for virtual machine services. The virtual machine high-availability mechanism ensures that when one or more computing nodes fail, the virtual machines on the failed nodes can be effectively managed and restored, and the virtual machines running in the OpenStack environment can continue to run without affecting the user experience.
[0028] Nova-API: A key component in the OpenStack Compute project (Nova), responsible for handling all requests from users, including operations such as creating, deleting, starting, and stopping virtual machines, as well as querying the status of virtual machines and obtaining computing resource information. It is the entry point of the Nova service, providing a RESTful API interface that allows users to manage and control virtual machines through HTTP requests.
[0029] Virtual Machine Evacuation: In a cloud computing environment, when the physical host or underlying hardware hosting a virtual machine fails, such as a power outage or hardware damage, to avoid the impact of this failure on the services on the virtual machine, the cloud platform will automatically or manually migrate the virtual machine from the failed physical host to another healthy physical host. During the evacuation process, the cloud platform will attempt to maintain the integrity of the virtual machine's state, minimize the service interruption time, and ensure that the applications and services on the virtual machine can continue to run seamlessly on the new host.
[0030] Nova-Conductor: A core component in the OpenStack Nova Compute project, which is mainly responsible for handling long transactions that require database access or do not require real-time responses. Nova-Conductor receives requests from Nova-API, executes database operations such as querying or updating the status of virtual machines, and coordinates the work of multiple Nova-Compute service instances.
[0031] Nova-Scheduler: The scheduler in the OpenStack Nova Compute project, responsible for allocating resources among multiple computing nodes. When Nova-Conductor needs to create or migrate a virtual machine on a computing node, it will request Nova-Scheduler to select a suitable node.
[0032] In the following text, the embodiments of the present application will be described in detail with reference to the accompanying drawings and in combination with the embodiments.
[0033] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0034] The method embodiments provided in the embodiments of this application can be executed in a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is a hardware structure block diagram of a server device for a virtual machine evacuation method according to an embodiment of this application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 the figure is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than
[0035] shown in
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0037] Figure 2 is a flowchart of the high-availability evacuation process of virtual machines, as Figure 2 shown, in the high-availability process of community native logic, the nova-api service processes the virtual machine evacuation process triggered by the virtual machine HA module, the nova-condutor service calls nova-scheduler to obtain the host resources required for virtual machine evacuation, and the nova-scheduler service scheduler returns an available host for the creation of a virtual machine. However, when two virtual machine requests arrive at different nova-scheduler services simultaneously, the nova-scheduler service is very likely to select the same candidate host to evacuate these two virtual machines. At this time, resource applications for these two virtual machines on the candidate host will conflict. If the resources on the candidate host are not sufficient to create two virtual machines, it will cause the evacuation of one of the virtual machines to fail, thereby reducing the success rate of virtual machine high availability.
[0038] In this embodiment, a method for evacuating virtual machines is provided. Figure 3 is a flowchart of the method for evacuating virtual machines according to an embodiment of the present application, as Figure 3 shown, the process includes the following steps:
[0039] Step S302, in response to receiving a virtual machine evacuation request, obtain a list of candidate hosts, where the virtual machine evacuation request is used to indicate that a source computing node in a cloud platform fails, and there is a source virtual machine hosted on the source computing node. The list of candidate hosts includes multiple candidate computing nodes, and the multiple candidate computing nodes in the list of candidate hosts are sorted based on the weights of the computing nodes;
[0040] In the embodiment of the present invention, the virtual machine evacuation request can be understood as a request sent to the system when the monitoring or HA module of the OpenStack cloud platform detects that a computing node (physical server) fails. The content of the request is to migrate the virtual machines on the faulty node to healthy computing nodes to ensure the business continuity and high availability of the virtual machines. The evacuation request usually contains detailed information about the source virtual machine, such as ID, specifications, required resources, etc., which are not limited here.
[0041] The candidate host list can be understood as follows: when a virtual machine evacuation request is received, the scheduling service (such as nova-scheduler) will select a batch of nodes from all available computing nodes in the cloud platform that can host the evacuated virtual machines and return them as a candidate list, which is not restricted here.
[0042] The source computing node can be understood as the computing node that has failed, that is, the physical server that originally hosted the virtual machine to be evacuated. Exemplarily, the failure may include but is not limited to hardware failure, software crash, network interruption, etc., resulting in the virtual machine on the node being unable to run properly, which is not restricted here.
[0043] The source virtual machine can be understood as the virtual machine running on the source computing node. When the source computing node fails, these virtual machines need to be migrated to other healthy nodes to continue running their services. There may be multiple source virtual machines, and each virtual machine has its specific configuration and requirements, which is not restricted here.
[0044] The candidate computing node can be understood as the computing node in the candidate host list. The candidate computing nodes are selected through a series of filtering and weight calculations and are considered to have sufficient resources and capabilities to host the source virtual machines. The number of candidate computing nodes may be determined by configuration parameters, and when resource conflicts are obtained, the scheduling service will try other nodes in the order of the list, which is not restricted here.
[0045] The weight of a computing node can be understood as a value that evaluates the ability and priority of the computing node in hosting virtual machine evacuation requests. The weight may be calculated based on multiple factors, such as CPU usage, memory usage, disk space, network bandwidth, load, etc. A computing node with a high weight indicates that it is more suitable for hosting the evacuated virtual machines in terms of resources and performance, so it will be given priority in scheduling, which is not restricted here.
[0046] Responding to the receipt of a virtual machine evacuation request, obtaining the candidate host list can be understood as follows: when OpenStack detects that a computing node hosting a virtual machine fails, the system will automatically trigger an evacuation process, and the scheduling component (such as nova-scheduler) will, based on the current resource status of the cloud platform, select a series of computing nodes that can meet the resource requirements of the virtual machines, and these nodes are the candidate host list.
[0047] Sorting multiple candidate computing nodes in the candidate host list based on the weights of the computing nodes can be understood as follows: during the process of obtaining the candidate host list, the computing nodes are not simply listed, but sorted according to the weights of each candidate computing node. Exemplarily, the sorted candidate host list means that when the scheduling component attempts to recreate a virtual machine, it will first select the node with the highest weight in the list, and so on until the virtual machine is successfully created or all candidate nodes have been tried. There is no limitation here.
[0048] In an embodiment of the present invention, when the monitoring system of the OpenStack cloud platform detects that a certain computing node (source computing node) fails and can no longer provide a normal virtual machine running environment, the system automatically triggers a virtual machine evacuation request. After receiving the virtual machine evacuation request, the scheduling service of OpenStack (usually nova-scheduler) starts to work, filters out the nodes that can host the source virtual machine from all available computing nodes, and forms a candidate host list. The computing nodes in the candidate host list are not randomly arranged, but sorted based on the weights of the computing nodes. The sorting based on weights ensures more efficient utilization of resources. Nodes with higher weights are preferentially selected, which can better balance the load of the cloud platform and avoid resource waste and over-concentration.
[0049] Step S304: Determine a first target computing node to be migrated from multiple candidate computing nodes based on the candidate host list, where the first target computing node is the candidate computing node with the highest weight in the candidate host list;
[0050] In an embodiment of the present invention, the first target computing node can be understood as the computing node determined based on the weights of the computing nodes from the candidate host list for hosting and creating the evacuated virtual machine. Specifically, the first target computing node is the candidate computing node with the highest weight in the list, that is, the first target computing node is considered the best choice in terms of resource availability, performance, load balancing, etc. to meet the resource requirements of virtual machine evacuation.
[0051] Determining a first target computing node to be migrated from multiple candidate computing nodes based on the candidate host list can be understood as follows: based on the generated candidate host list, after receiving the virtual machine evacuation request, OpenStack selects the computing node with the highest weight from the candidate host list and determines it as the first target computing node to be migrated.
[0052] In an embodiment of the present invention, determining a first target computing node to be migrated from multiple candidate computing nodes based on the candidate host list ensures that OpenStack can allocate resources in an optimal manner during virtual machine evacuation, preferentially using the computing node with the best resource status, thereby improving the success rate of evacuation and the overall efficiency of the system.
[0053] Step S306: Create at least one target virtual machine on the first target computing node, and migrate the data stored in the source virtual machine to the target virtual machine, so as to evacuate the source virtual machine.
[0054] In the embodiment of the present invention, the target virtual machine can be understood as a new virtual machine created on the first target computing node, which is used to receive the services and data of the source virtual machine evacuated from the source computing node. Exemplarily, when the OpenStack cloud platform detects that the source computing node fails and can no longer support the operation of the source virtual machine, the system will select the first target computing node as the new hosting node according to the evacuation request and the sorting in the candidate host list to ensure business continuity and data security, which is not limited herein.
[0055] Creating at least one target virtual machine on the first target computing node and migrating the data stored in the source virtual machine to the target virtual machine to evacuate the source virtual machine can be understood as when the cloud platform detects that the source computing node (the node hosting the source virtual machine) fails and can no longer provide services, the system will select the first target computing node (the healthiest node with the highest weight) according to the previously generated candidate host list. On this node, a new virtual machine, that is, the target virtual machine, is created based on the configuration of the source virtual machine (such as CPU, memory, disk size). The creation of the target virtual machine is to replace the source virtual machine and continue to execute its business process to ensure the continuity and stability of the cloud service.
[0056] In the embodiment of the present invention, by creating a target virtual machine on the first target computing node and completing the data migration, the evacuation of the source virtual machine from the failed node to the healthy node is realized. The evacuation process ensures business continuity. Even in the case of the failure of the original node, the user's application programs and data can be quickly transferred to a new node, avoiding business interruption and data loss.
[0057] Through the above steps, when a virtual machine evacuation request is received, candidate computing nodes are screened based on the resource information of the computing nodes, and a list containing multiple candidate computing nodes is returned. After receiving the list of candidate computing nodes returned by the scheduling service, an attempt is made to create a virtual machine on the first computing node in the list of candidate computing nodes. When the creation of the virtual machine fails on the first computing node, the control service automatically selects the next computing node from the list of candidate computing nodes and continues to attempt to create the virtual machine until the virtual machine is successfully created on any computing node in the list of candidate computing nodes or all computing nodes in the list have been tried. When screening candidate computing nodes, the scheduling service calculates a resource weighted value according to the resource status of each computing node, and sorts the candidate computing nodes according to the resource weighted value to ensure an optimized sorting of the list of candidate computing nodes. Therefore, by introducing resource weighted sorting and a multi-attempt mechanism, the problem in the related art where two virtual machine requests screen out the same candidate hosts and resource conflicts occur when the two virtual machines apply for resources on the candidate hosts, resulting in the failure of virtual machine evacuation, can be solved, achieving the effect of significantly improving the efficiency and success rate of the virtual machine evacuation process.
[0058] According to another embodiment of the present application, in step S302, obtaining the list of candidate hosts includes the following execution steps:
[0059] Step S3021, obtain multiple first computing nodes;
[0060] Step S3022, screen the multiple first computing nodes to obtain multiple second computing nodes;
[0061] Step S3023, select multiple candidate computing nodes from the multiple second computing nodes based on a preset list number;
[0062] Step S3024, generate a list of candidate hosts based on the multiple candidate computing nodes.
[0063] In the embodiments of the present invention, the multiple first computing nodes can be understood as the set of all currently available, unscreened, and unsorted computing nodes in the OpenStack cloud platform. That is, the multiple first computing nodes can carry virtual machine resources, but have not undergone detailed resource evaluation and screening. When the virtual machine HA mechanism triggers the evacuation process, nova-scheduler will start screening from all these nodes to determine which nodes can meet the specific virtual machine resource requirements.
[0064] The second computing nodes can be understood as the set of computing nodes after preliminary resource screening of the multiple first computing nodes. Exemplarily, nova-scheduler will filter the first computing nodes to exclude those nodes with insufficient resources, abnormal status, or those that do not meet the scheduling policy, thereby obtaining a smaller but more evacuation requirement-compliant set of second computing nodes.
[0065] The number of preset lists can be understood as a parameter preset in the OpenStack configuration, indicating the number of candidate computing nodes that should be included in the final candidate host list. The setting of the number of preset lists can be adjusted according to the scale of the cloud platform, the load status, and the specific requirements of the customer for the HA mechanism to optimize the efficiency and success rate of the evacuation process. Exemplarily, through configuration, the number of candidate host lists can be set, and the number can be set to be greater than 1.
[0066] Obtaining multiple first computing nodes can be understood as that in the initial stage of the evacuation process, the nova-scheduler service obtains information about all currently available computing nodes that can participate in the evacuation scheduling from the computing node pool of the entire cloud platform.
[0067] Filtering multiple first computing nodes to obtain multiple second computing nodes can be understood as that after obtaining multiple first computing nodes, nova-scheduler performs detailed resource evaluation, health status check, and screening of other conditions that meet the evacuation requirements on the obtained first computing nodes. Exemplarily, it will check the CPU utilization rate, memory usage, disk space, network status, etc. of the nodes, and exclude those nodes with insufficient resources, abnormal status, or non-compliance with the evacuation policy. After this series of evaluations and screenings, the remaining set of computing nodes that meet the conditions is called the second computing nodes, which are not restricted here.
[0068] Selecting multiple candidate computing nodes from multiple second computing nodes based on the number of preset lists can be understood as that nova-scheduler selects a certain number of computing nodes from the set of second computing nodes according to the preset number of candidate host lists (usually set in the configuration file to optimize the evacuation efficiency and success rate), and these nodes are considered as candidate targets for evacuating virtual machines. Exemplarily, the selection process may consider factors such as node performance, load, resource sufficiency, and geographical location to generate an optimal set of candidate computing nodes, which are not restricted here.
[0069] Generating a candidate host list based on multiple candidate computing nodes can be understood as that nova-scheduler sorts the selected candidate computing nodes according to a certain weight or priority to form a candidate host list.
[0070] In an embodiment of the present invention, OpenStack can screen out a second computing node that meets the evacuation requirements from the entire computing node pool. Based on the number of preset lists, candidate computing nodes are selected, and finally a candidate host list is generated. By pre-generating the candidate host list, when an evacuation event occurs, the nova-condutor service can directly select a node from the list, reducing the decision-making time and accelerating the execution speed of the evacuation process, which is crucial for handling sudden failure events.
[0071] According to another embodiment of the present application, in step S3022, screening the multiple first computing nodes to obtain multiple second computing nodes includes the following execution steps:
[0072] Screen the multiple first computing nodes based on a filter to obtain multiple third computing nodes, where the filter is used to screen out computing nodes that do not meet the preset requirements, and the preset requirements include at least one of the following: the computing node has the network interfaces and bandwidth required by the virtual machine, the computing node complies with access control, firewall, and encryption standards, the computing node provides the required storage type and capacity, the computing node runs the required operating system and software packages, and the computing node hardware supports the image processor architecture and central processor architecture corresponding to the virtual machine;
[0073] Determine the weights of the multiple third computing nodes respectively based on a weigher, where the weigher is used to calculate the weights from at least one of the following aspects: resource utilization rate, load status, geographical location, power status, and workload balance between nodes;
[0074] Sort the multiple third computing nodes based on a preset sorting strategy and the weights of the multiple third computing nodes to obtain multiple second computing nodes, where the preset sorting strategy is used to represent sorting in descending order based on the weights.
[0075] In an embodiment of the present invention, the filters can be understood as a set of rules used in the nova-scheduler service to evaluate and screen whether a computing node meets the virtual machine creation requirements.
[0076] The third computing node can be understood as a set of computing nodes obtained after being screened by the filter. Exemplarily, the set of third computing nodes includes computing nodes that meet the virtual machine evacuation requirements in terms of resources, hardware, network, software, and security, and is the input for the subsequent weighing and sorting processes, which is not limited here.
[0077] Weighers can be understood as key components for evaluating the resource availability and performance of computing nodes, which determine the priorities of computing nodes. Exemplarily, a weigher assigns a weight value to each third computing node passing through the filter. The weight value is calculated based on the comprehensive performance and resource status of the node. The higher the weight value of a node, the higher its priority and the more suitable it is as the target for virtual machine evacuation, which is not limited herein.
[0078] Filtering multiple first computing nodes based on a filter to obtain multiple third computing nodes can be understood as the nova-scheduler service using a series of preset filtering rules to check all first computing nodes (i.e., all currently available computing nodes) to ensure that these nodes meet the specific requirements for virtual machine evacuation. The content checked by the filter may include the network capabilities, storage resources, operating system, hardware compatibility, etc. of the nodes. Any computing node that does not meet these preset requirements will be excluded. The result of the filtering is a smaller set containing all compliant nodes, i.e., the set of third computing nodes, which is not limited herein.
[0079] Determining the weights of multiple third computing nodes based on the weigher can be understood as calculating a comprehensive weight value for each node in the set of third computing nodes according to factors such as its resource status, load condition, geographical location, power status, and workload balance among nodes, which is not limited herein.
[0080] Sorting multiple third computing nodes based on a preset sorting strategy and the weights of the multiple third computing nodes to obtain multiple second computing nodes can be understood as, after determining the weights of each third computing node, sorting the set of third computing nodes according to a preset sorting strategy (e.g., sorting by weight value from high to low). The purpose of the sorting is to generate a list of nodes with decreasing priorities, and this list is the set of second computing nodes, which is not limited herein.
[0081] In the embodiments of the present invention, by filtering, evaluating, and sorting multiple first computing nodes, the best candidate node list is selected from all available computing nodes for virtual machine evacuation operations, ensuring the efficient utilization of cloud platform resources and the continuity of services.
[0082] According to another embodiment of the present application, the virtual machine evacuation method further includes: in response to a failure of the source computing node, generating a virtual machine evacuation request based on the high availability mechanism.
[0083] In the embodiments of the present invention, the high-availability mechanism, i.e., the HA mechanism, aims to ensure that critical services and applications (such as virtual machines) on the cloud platform can continue to run even in the event of unexpected situations such as hardware failures, software errors, or network outages, thereby maintaining business continuity and data integrity. Exemplarily, when a certain source computing node fails, the virtual machine HA mechanism of OpenStack will automatically trigger a series of response measures to protect the virtual machines running on the faulty node, which is not limited herein.
[0084] Generating a virtual machine evacuation request based on the high-availability mechanism in response to a failure of a source computing node can be understood as that when the monitoring system detects a failure of a certain source computing node (such as hardware failure, operating system crash, network connection interruption, etc.), the built-in high-availability (HA) mechanism of OpenStack will automatically intervene and generate a virtual machine evacuation request.
[0085] In the embodiments of the present invention, when facing a computing node failure, OpenStack can automatically generate and execute a virtual machine evacuation request through the built-in high-availability mechanism to resume the operation of the virtual machine, thereby ensuring the continuity and stability of the cloud service.
[0086] According to another embodiment of the present application, generating a virtual machine evacuation request based on the high-availability mechanism in response to a failure of a source computing node includes the following execution steps:
[0087] Performing status monitoring on any computing node in the cloud platform based on a first service component to obtain a status monitoring result;
[0088] In response to the status monitoring result indicating that the source computing node has failed, identifying the source virtual machines running on the source computing node based on the high-availability mechanism to obtain an identification result;
[0089] Generating a virtual machine evacuation request corresponding to any source virtual machine based on the identification result.
[0090] In the embodiments of the present invention, the first service component can be understood as the monitoring and status management service component in OpenStack. Exemplarily, the first service component can be the Nova Compute service, which is responsible for monitoring the status of computing nodes, which is not limited herein.
[0091] The status monitoring result can be understood as the result obtained by the continuous status monitoring of the computing nodes in the cloud platform by the first service component (such as Nova Compute). Exemplarily, the status monitoring result includes but is not limited to the hardware status of the node (such as CPU usage, memory usage, disk space, etc.), the software service status (whether it is running normally), the network connection status, and any abnormal situations that may affect the stability of the node and the operation of the virtual machine, which is not limited herein.
[0092] The recognition result can be understood as the detailed information of all virtual machines to be evacuated, such as virtual machine ID, specifications, running status, etc., which are not limited herein.
[0093] Based on the first service component to monitor the status of any computing node in the cloud platform, the obtained status monitoring result can be understood as that OpenStack uses a dedicated service component, namely the first service component (usually refers to the Nova Compute service), to continuously monitor the health and status of each computing node in the cloud platform. The obtained status monitoring result includes the current running status information of the computing node, such as whether the usage of CPU, memory, and disk is normal, whether the network connection is stable, and whether the services on the node are running normally. If any key metrics of the computing node exceed the preset normal range, or a hardware failure or software anomaly is detected, the status monitoring result will reflect that the node is in a failed or unhealthy state, which is not limited herein.
[0094] In response to the status monitoring result indicating that the source computing node has failed, based on the high-availability mechanism to identify the source virtual machines running on the source computing node, the obtained recognition result can be understood as that when the status monitoring result of the first service component shows that the source computing node has failed, the high-availability mechanism of OpenStack (usually implemented by the HA component of Nova) will automatically start the response process. First, it identifies and determines all virtual machines (i.e., source virtual machines) running on the failed node, thereby obtaining the recognition result. The recognition result includes the detailed information of all virtual machines on the failed node, such as virtual machine ID, specifications, running status, etc., so as to determine which virtual machines need to be urgently migrated to avoid data loss and service interruption, which is not limited herein.
[0095] Based on the recognition result, generate a virtual machine evacuation request corresponding to any source virtual machine. It can be understood that after obtaining the recognition result of all source virtual machines running on the source computing node, OpenStack will generate an evacuation request for each of these virtual machines. These requests include the detailed information required for virtual machine evacuation, such as virtual machine specifications, current status, storage requirements, etc., as well as the priority and target requirements of the evacuation operation, which are not limited herein.
[0096] In the embodiment of the present invention, when the source computing node fails, OpenStack continuously monitors the status through the first service component (such as the Nova Compute service), identifies the failure status, and further identifies all virtual machines on the failed node through the HA mechanism, and generates evacuation requests for these virtual machines. The above steps ensure that the cloud platform can quickly respond to failures, protect the integrity of virtual machines and data, and thus maintain high availability.
[0097] According to another embodiment of the present application, the virtual machine evacuation method further includes the following execution steps:
[0098] In response to the failure of creating at least one target virtual machine on a first target computing node, determine a second target computing node from a list of candidate hosts, where the second target computing node is the next computing node of the first target computing node;
[0099] Create at least one target virtual machine on the second target computing node.
[0100] In an embodiment of the present invention, determining a second target computing node from a list of candidate hosts in response to the failure of creating at least one target virtual machine on a first target computing node can be understood as follows: when the OpenStack cloud platform attempts to create a virtual machine on a first computing node selected as an evacuation target, but the creation process fails (possibly due to insufficient resources, hardware failures, software conflicts, etc.), the evacuation operation will not stop. Instead, it will continue to select the next computing node (i.e., the second target computing node) from a pre-screened list of candidate hosts as an alternative target. Exemplarily, this process is continuous, and the OpenStack cloud platform will try each node in the candidate list until it finds a computing node that can successfully create a virtual machine, which is not limited herein.
[0101] Creating at least one target virtual machine on the second target computing node can be understood as follows: when the second target computing node (i.e., the alternative computing node) is determined, the OpenStack cloud platform will attempt to create on this node the virtual machine that failed to be created on the first target computing node before.
[0102] In an embodiment of the present invention, when the initial evacuation attempt fails, the OpenStack cloud platform will automatically try the next computing node in the list until a virtual machine is successfully created. This fault recovery strategy greatly enhances the stability and business reliability of the cloud platform, ensuring that even in the face of node failures, critical virtual machine services can be quickly restored and continue to run.
[0103] According to another embodiment of the present application, the virtual machine evacuation method further includes: in response to the failure of creating at least one target virtual machine on multiple candidate computing nodes, stop evacuating the source virtual machine.
[0104] In an embodiment of the present invention, stopping evacuating the source virtual machine in response to the failure of creating at least one target virtual machine on multiple candidate computing nodes can be understood as follows: when the OpenStack cloud platform attempts to recreate or evacuate the virtual machines affected by the source computing node failure on a series of preselected candidate computing nodes, if the attempts to create target virtual machines on all these candidate nodes end in failure, then the system will stop further evacuation attempts.
[0105] In the embodiments of the present invention, after encountering consecutive creation failures, OpenStack will determine that there are not enough resources or conditions in the current cloud environment to meet the running requirements of virtual machines, and the system will stop further evacuation attempts to prevent the problem from deteriorating further.
[0106] This application provides a solution for optimizing the high availability mechanism of virtual machines in the OpenStack cloud platform. This solution changes the single-node scheduling to multiple candidate host list scheduling. When screening available hosts, multiple candidate hosts are obtained and added to the candidate list. In the case of resource conflicts among candidate hosts, other hosts can be selected from the candidate list to create virtual machines until the virtual machine is successfully created. If the hosts in the candidate list do not meet the resource requests, it indicates that there is no suitable host in the current environment to create virtual machines, and the virtual machine evacuation fails. Figure 4 It is a flowchart of the high availability evacuation process of virtual machines in the OpenStack cloud platform according to the embodiments of this application, as Figure 4 shown. First, a virtual machine evacuation request is received for processing. When a node in the cloud platform fails, the virtual machine HA module triggers the evacuation of the virtual machines on the failed node. The Nova-api service receives the virtual machine evacuation request sent by the virtual machine HA module and hands it over to the nova-condutor service for virtual machine reconstruction. Then, the Nova-condutor service calls the nova-scheduler service to obtain a candidate host list instead of a single host. Through configuration, the number of candidate host lists can be set. During the process of screening candidate hosts by the nova-scheduler service, the available candidate hosts are sorted by weight through filters and weighers, and according to the set number of candidate host lists, an available candidate host list is returned to the nova-condutor service. Finally, the Nova-condutor service sequentially selects a computing node from the candidate host list and calls the nova-compute service to create a virtual machine on this computing node. If the virtual machine is successfully created, the virtual machine evacuation is successful. If the virtual machine creation fails, the nova-compute service calls the nova-condutor service to select another computing node from the candidate host list and calls the nova-compute service on this computing node to create a virtual machine. If the virtual machine creation fails, continue to obtain computing nodes from the candidate host list for creation until the virtual machine is successfully created. If all the nodes in the candidate list are used up and the virtual machine still fails to be created, it indicates that there is no suitable host in the current environment to create virtual machines, and the virtual machine evacuation fails.
[0107] Specifically, first, prepare three control nodes, control1, control2, and control3, and five computing nodes, node1, node2, node3, node4, and node5. The CPU of node1 is 128 cores, the memory is 256G, and the disk is 2T. The CPU of node2 is 48 cores, the memory is 64G, and the disk is 1T. The CPU of node3 is 48 cores, the memory is 64G, and the disk is 1T. The CPU of node4 is 48 cores, the memory is 64G, and the disk is 1T. The CPU of node5 is 256 cores, the memory is 512G, and the disk is 20G. Create 50 virtual machines on node1, with the virtual machine specification of 2 CPUs, 4G memory, and 20G disk (2C4G20G). At the same time, configure the virtual machine HA high-availability mechanism to conduct high-availability tests on the virtual machines.
[0108] Secondly, set the number of candidate host lists to 1 and power off the computing node node1. Since the resource weight of node5 is the highest, most of the initial requests are scheduled to node5 to create virtual machines. However, node5 can only meet the creation of one virtual machine. Therefore, only one virtual machine can be successfully created among these requests, and the virtual machines of other requests will fail to evacuate. Among them, the weight calculation formula: w = weightcpu * number of CPUs + weightmem * size of memory + weightdisk * capacity of disk. Since CPU, memory, and disk belong to different types of resources, it is necessary to normalize the CPU, memory, and disk resources first. The processed result of the CPU is [0.5, 0.5, 0.5, 1], the processed result of the memory is [0.5, 0.5, 0.5, 1], and the processed result of the disk is [1, 1, 1, 0.02]. Set the CPU weight to 2, the memory weight to 1, and the disk weight to 0.1. Thus, the resource weighted value of node2 is 1.6, the resource weighted value of node3 is 1.6, the resource weighted value of node4 is 1.6, the resource weighted value of node5 is 3.002, and the node1 service is in the down state and does not participate in the scheduling calculation. Finally, through statistics, the success rate of virtual machine evacuation is obtained as 68%.
[0109] Then, set the number of candidate host lists to 3, and also power off the computing node node1. According to the log, it can be seen that there are always 3 eligible nodes in the virtual machine scheduling candidate list host. At this time, observe the success rate of virtual machine evacuation on the node1 node. Through statistics, the success rate of virtual machine evacuation is obtained as 96%.
[0110] Finally, analyze and compare the effects of the two virtual machine HA tests. After optimization, the success rate of virtual machine evacuation is much higher than that before optimization, achieving the expected optimization goal.
[0111] In addition, the root cause of resource conflicts during virtual machine creation on the nova-compute server side is that nova-scheduler processes scheduling requests for virtual machine evacuation concurrently. If the concurrency issue of different nova-schedulers obtaining host resources under multiple control nodes can be solved, there is no need to schedule the candidate host list on the nova-compute server side, which can improve the virtual machine evacuation efficiency more efficiently and increase the success rate of the virtual machine HA high-availability mechanism.
[0112] Among them, the execution subject of the above steps can be [servers, terminals], etc., but is not limited thereto.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0114] In this embodiment, a virtual machine evacuation device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0115] Figure 5 is a structural block diagram of the virtual machine evacuation device 500 according to an embodiment of the present application, as Figure 5As shown in the figure, the device includes: an obtaining module 501, configured to obtain a list of candidate hosts in response to receiving a virtual machine evacuation request, where the virtual machine evacuation request is used to indicate that a source computing node in a cloud platform fails, and a source virtual machine is hosted on the source computing node, and the list of candidate hosts includes multiple candidate computing nodes, and the multiple candidate computing nodes in the list of candidate hosts are sorted based on the weights of the computing nodes; a determining module 502, configured to determine a first target computing node to be migrated from the multiple candidate computing nodes based on the list of candidate hosts, where the first target computing node is the candidate computing node with the highest weight in the list of candidate hosts; a creating module 503, configured to create at least one target virtual machine on the first target computing node, and migrate the data stored in the source virtual machine to the target virtual machine, so as to evacuate the source virtual machine.
[0116] According to another embodiment of the present application, the obtaining module 501 is further configured to obtain multiple first computing nodes; screen the multiple first computing nodes to obtain multiple second computing nodes; select multiple candidate computing nodes from the multiple second computing nodes based on a preset number of lists; and generate a list of candidate hosts based on the multiple candidate computing nodes.
[0117] According to another embodiment of the present application, the obtaining module 501 is further configured to screen the multiple first computing nodes based on a filter to obtain multiple third computing nodes, where the filter is used to screen out computing nodes that do not meet preset requirements, and the preset requirements include at least one of the following: the computing node has the network interface and bandwidth required by the virtual machine, the computing node complies with access control, firewall, and encryption standards, the computing node provides the required storage type and capacity, the computing node runs the required operating system and software packages, the computing node hardware supports the image processor architecture and central processor architecture corresponding to the virtual machine; respectively determine the weights of the multiple third computing nodes based on a weigher, where the weigher is used to calculate the weights from at least one of the following aspects: resource utilization rate, load status, geographical location, power status, workload balance between nodes; sort the multiple third computing nodes based on a preset sorting strategy and the weights of the multiple third computing nodes to obtain multiple second computing nodes, where the preset sorting strategy is used to indicate sorting in descending order based on the weights.
[0118] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: all the above-mentioned modules are located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.
[0119] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is set to execute the steps in any one of the above method embodiments when running.
[0120] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media that can store computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs.
[0121] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0122] In an exemplary embodiment, the above electronic device may further include a transmission device and input / output devices. Among them, the transmission device is connected to the above processor, and the input / output devices are connected to the above processor.
[0123] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0124] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0125] An embodiment of the present application also provides a computer program. The computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0126] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0127] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.
[0128] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A virtual machine evacuation method, characterized in that: The method comprises: In response to receiving a virtual machine evacuation request, obtaining a candidate host list, wherein the virtual machine evacuation request is used to indicate that a source computing node in the cloud platform fails, the source computing node carries a source virtual machine, the candidate host list includes multiple candidate computing nodes, and the multiple candidate computing nodes in the candidate host list are sorted based on the weights of the computing nodes; Determine a first target computing node to be migrated from the plurality of candidate computing nodes based on the candidate host list, wherein the first target computing node is the candidate computing node with the highest weight in the candidate host list; At least one target virtual machine is created on the first target computing node, and data stored in the source virtual machine is migrated to the target virtual machine to evacuate the source virtual machine.
2. The method according to claim 1, characterized in that The obtaining of the candidate host list comprises: Obtain multiple first computing nodes; Screening the plurality of first computing nodes to obtain a plurality of second computing nodes; Selecting a plurality of the candidate computing nodes from a plurality of the second computing nodes based on a preset number of lists; The candidate host list is generated based on the plurality of candidate computing nodes.
3. The method according to claim 2, characterized in that The screening of the plurality of first computing nodes to obtain the plurality of second computing nodes comprises: Filtering the plurality of first computing nodes based on a filter to obtain a plurality of third computing nodes, wherein the filter is used to filter out computing nodes that do not meet preset requirements, and the preset requirements include at least one of the following: the computing node has a network interface and bandwidth required by the virtual machine, the computing node complies with access control, firewall and encryption standards, the computing node provides a required storage type and capacity, the computing node runs an operating system and software package required, and the computing node hardware supports an image processor architecture and a central processing unit architecture corresponding to the virtual machine; Determine the weights of the plurality of third computing nodes respectively based on a weighing device, wherein the weighing device is used to calculate the weights from at least one of the following aspects: resource utilization, load status, geographic location, power supply status, and workload balance between nodes; The plurality of third computing nodes are sorted based on a preset sorting strategy and weights of the plurality of third computing nodes to obtain a plurality of second computing nodes, wherein the preset sorting strategy is used to indicate sorting based on a descending order of weights.
4. The method according to claim 1, characterized in that: The method further comprises: In response to a failure of the source computing node, the virtual machine evacuation request is generated based on a high availability mechanism.
5. The method according to claim 4, characterized in that In response to the source computing node failing, generating the virtual machine evacuation request based on a high availability mechanism includes: Performing status monitoring on any computing node in the cloud platform based on the first service component to obtain a status monitoring result; In response to the status monitoring result indicating that the source computing node fails, identifying the source virtual machine running on the source computing node based on the high availability mechanism to obtain an identification result; The virtual machine evacuation request corresponding to any one of the source virtual machines is generated based on the identification result.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: In response to a failure to create at least one of the target virtual machines on the first target computing node, determining a second target computing node from the candidate host list, wherein the second target computing node is a next computing node of the first target computing node; At least one of the target virtual machines is created on the second target computing node.
7. The method according to claim 6, characterized in that The method further comprises: In response to failure to create at least one of the target virtual machines on the plurality of candidate computing nodes, stopping evacuating the source virtual machine.
8. A virtual machine evacuation device, characterized in that: The device comprises: An acquisition module is used to acquire a candidate host list in response to receiving a virtual machine evacuation request, wherein the virtual machine evacuation request is used to indicate that a source computing node in the cloud platform fails, the source computing node carries a source virtual machine, the candidate host list includes multiple candidate computing nodes, and the multiple candidate computing nodes in the candidate host list are sorted based on the weight of the computing nodes; A determination module, configured to determine a first target computing node to be migrated from a plurality of the candidate computing nodes based on the candidate host list, wherein the first target computing node is the candidate computing node with the highest weight in the candidate host list; A creation module is used to create at least one target virtual machine on the first target computing node, and migrate the data stored in the source virtual machine to the target virtual machine to evacuate the source virtual machine.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 are implemented.