Data migration method, system and device of virtual machine and storage medium
By identifying virtual machines with unbalanced resource loads, using machine learning models to select migration targets and dynamically adjust transmission rates, the problem of low virtual machine migration efficiency is solved, and efficient migration with reasonable resource allocation and stable business is achieved.
Patent Information
- Application Number
- CN202511159535.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing virtual machine migration strategies rely on fixed thresholds and resource usage comparisons, resulting in irrational resource allocation, impacting business operations, and low migration efficiency.
By identifying the first virtual machine with resource load higher than the first threshold and the candidate virtual machines with resource load lower than the second threshold, the machine learning model is used to comprehensively consider resources, business data and network link data to select the most suitable migration target virtual machine, and data migration is performed through differential calculation and dynamic adjustment of transmission rate.
It enables more reasonable and accurate migration decisions, avoids resource waste and business performance degradation, and improves the efficiency and stability of virtual machine data migration.
Smart Images

Figure CN120653370A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a data migration method and system, device, and storage medium for a virtual machine. Background Art
[0002] In existing cloud computing environments, with the increasing number of virtual machines and the dynamic changes in business load, the problem of resource load imbalance between virtual machines has become increasingly prominent. Traditional virtual machine migration strategies often rely on fixed thresholds and simple resource utilization comparisons, resulting in irrational resource allocation during the migration process and unnecessary negative impacts on running businesses, leading to inefficient migration.
[0003] Therefore, there is a technical problem in the related art that the data migration efficiency of the virtual machine is low. Summary of the Invention
[0004] The embodiments of the present application provide a data migration method and system, an apparatus, and a storage medium for a virtual machine, so as to at least solve the technical problem of low data migration efficiency of virtual machines in related technologies.
[0005] According to one embodiment of the present application, a method for data migration of a virtual machine is provided, comprising: determining a first virtual machine having a resource load higher than a first threshold from multiple virtual machines, and determining at least one candidate virtual machine having a resource load lower than a second threshold from the multiple virtual machines, wherein the first threshold is higher than the second threshold; determining a second virtual machine for data migration from the at least one candidate virtual machine by utilizing first resource data and business data of the first virtual machine, second resource data of at least one candidate virtual machine, and network link data between the first virtual machine and the at least one candidate virtual machine; and transferring data of the first virtual machine to the second virtual machine.
[0006] According to another embodiment of the present application, a data migration device for a virtual machine is provided, including: a first determination unit, used to determine a first virtual machine whose resource load is higher than a first threshold from multiple virtual machines, and to determine at least one candidate virtual machine whose resource load is lower than a second threshold from multiple virtual machines, wherein the first threshold is higher than the second threshold; a second determination unit, used to determine a second virtual machine for data migration from at least one candidate virtual machine by using first resource data and business data of the first virtual machine, second resource data of at least one candidate virtual machine, and network link data between the first virtual machine and the at least one candidate virtual machine; and a migration unit, used to transfer data of the first virtual machine to the second virtual machine.
[0007] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when running.
[0008] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0009] Through the embodiments provided in this application, a first virtual machine with an excessively high resource load and candidate virtual machines with a resource load below a specific threshold are dynamically identified from multiple virtual machines, ensuring that the migration decision takes into account the unbalanced distribution of resources from the source and avoids blind migration operations. Utilizing the real-time resource and business data of the first virtual machine, the resource status of the candidate virtual machine, and the network link data between the two, an intelligent algorithm (such as a machine learning model) is used to select the most suitable second virtual machine from the candidate list as the migration target. This analysis based on multi-dimensional decision factors can comprehensively consider business needs, network status, and the availability of the target host, thereby making more reasonable and accurate migration decisions, avoiding resource waste and business performance degradation during the migration process, achieving the technical effect of improving the data migration efficiency of virtual machines, and solving the technical problem of low data migration efficiency of virtual machines existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a hardware structure block diagram of a data migration method for a virtual machine according to an embodiment of the present application;
[0011] Figure 2 is a flow chart of a data migration method for a virtual machine according to an embodiment of the present application;
[0012] Figure 3 This is a flowchart of steps for implementing intelligent migration decision-making based on multi-dimensional decision factors according to an embodiment of the present application;
[0013] Figure 4 is a flowchart of implementation steps for dynamic monitoring and adaptive adjustment of a migration process according to an embodiment of the present application;
[0014] Figure 5 This is a structural block diagram of a data migration device for a virtual machine according to an embodiment of the present application. DETAILED DESCRIPTION
[0015] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0016] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0017] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal of a virtual machine data migration method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0018] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data migration method of the virtual machine in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0019] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0020] As an optional solution, virtual machine data migration method, such as Figure 2 As shown, the specific steps include:
[0021] S202, determining a first virtual machine from the plurality of virtual machines whose resource load is higher than a first threshold, and determining at least one candidate virtual machine from the plurality of virtual machines whose resource load is lower than a second threshold, wherein the first threshold is higher than the second threshold;
[0022] S204, determining a second virtual machine for data migration from the at least one candidate virtual machine using the first resource data and service data of the first virtual machine, the second resource data of the at least one candidate virtual machine, and the network link data between the first virtual machine and the at least one candidate virtual machine;
[0023] S206: Transfer the data of the first virtual machine to the second virtual machine.
[0024] Optionally, in this embodiment, a virtual machine (VM) is a software-implemented machine running on a physical server, capable of running an operating system and application programs like an independent computer.
[0025] Optionally, in this embodiment, resource load refers to the usage of computing resources (such as CPU, memory), storage resources (disk I / O), and network resources (bandwidth) consumed by the virtual machine.
[0026] Optionally, in this embodiment, the first threshold and the second threshold are used to define critical points of excessive and low resource loads, respectively. The first threshold is usually set at the warning line of resource utilization, while the second threshold is used to find relatively idle resource carrying capacity.
[0027] Optionally, in this embodiment, the first resource data and the second resource data refer to the real-time resource usage of the first virtual machine and at least one candidate virtual machine, respectively, including CPU usage, memory usage, disk I / O rate, and network traffic, etc.
[0028] Optionally, in this embodiment, the network link data includes indicators such as bandwidth utilization, delay, and packet loss rate of the network connection between virtual machines, which are used to evaluate the network transmission status.
[0029] Optionally, in this embodiment, the resource usage of all virtual machines in the data center is first continuously monitored. When it is detected that the resource load of a certain virtual machine is continuously higher than the set first threshold, while the resource load of other virtual machines is lower than the second threshold, the former is marked as the first virtual machine and the latter is marked as the candidate virtual machine, and resource balancing operations are prepared.
[0030] For example, suppose there are 10 virtual machines in a data center, where the CPU usage of virtual machine A is consistently above 90%, while the CPU usage of virtual machines B and C are both below 30%. In this case, virtual machine A is the first virtual machine, and virtual machines B and C are candidate virtual machines.
[0031] The first VM's resource data and service characteristics are collected, along with the resource status of the candidate VMs. The network link status data between the first VM and the candidate VMs is also obtained. Based on this information, an intelligent algorithm is used to comprehensively analyze the candidate VMs, identify the optimal second VM from the candidate VMs to receive the first VM's data, and calculate the optimal migration time.
[0032] To further illustrate, for virtual machine A, its business type (for example, real-time data processing or batch processing tasks) and current resource usage are analyzed, and then combined with the resource status of virtual machines B and C and the network link data between A and B, and A and C, virtual machine B is determined to be the best migration target after calculation.
[0033] Based on the intelligent decision, the data transfer process is initiated, efficiently transferring the data from the first virtual machine to the selected second virtual machine, achieving online migration of the virtual machine. During data transfer, the system dynamically adjusts the transmission rate based on the real-time status of the network link to ensure the stability of the migration process and minimize the impact on other services.
[0034] To further illustrate, after determining virtual machine B as the migration target, the system begins to transfer the data of virtual machine A to B, monitoring the network status during the process. If the bandwidth is tight, the transmission rate will be automatically reduced to avoid affecting the normal operation of other virtual machines in the data center.
[0035] As can be appreciated, in this embodiment, through real-time monitoring and intelligent analysis, it is possible to accurately identify virtual machines with excessive resource utilization and select the most appropriate migration target among virtual machines with more abundant resources. This method not only optimizes the migration decision-making process but also ensures the efficiency and stability of the migration operation by intelligently controlling the data transmission rate.
[0036] Through the embodiments provided in this application, a first virtual machine with an excessively high resource load and candidate virtual machines with a resource load below a specific threshold are dynamically identified from multiple virtual machines, ensuring that migration decisions take into account the imbalanced distribution of resources from the source and avoiding blind migration operations. Utilizing the real-time resource and business data of the first virtual machine, the resource status of the candidate virtual machines, and the network link data between the two, an intelligent algorithm (such as a machine learning model) is used to select the most suitable second virtual machine from the candidate list as the migration target. This analysis based on multi-dimensional decision factors can comprehensively consider business needs, network status, and the availability of the target host, thereby making more reasonable and accurate migration decisions, avoiding resource waste and business performance degradation during the migration process, and achieving the technical effect of improving the data migration efficiency of virtual machines.
[0037] As an optional solution, before transferring the data of the first virtual machine to the second virtual machine, the method further includes:
[0038] Determining a data migration time for data migration using the first resource data and the service data, as well as the second resource data and the network link data;
[0039] Transferring data from a first virtual machine to a second virtual machine includes:
[0040] The data of the first virtual machine is transferred to the second virtual machine according to the data migration time.
[0041] Optionally, in this embodiment, the data migration time is the optimal time point determined by an intelligent algorithm based on the current system status and expected resource utilization efficiency, and is used to start the data migration process from the first virtual machine to the second virtual machine to reduce the impact on the running business and optimize resource allocation.
[0042] Optionally, in this embodiment, before formally executing the data migration, the migration process is further optimized by utilizing the resource data and business data of the first virtual machine, as well as the second resource data of the candidate virtual machine and the network link data between the two, to determine the optimal data migration time through intelligent analysis based on multi-dimensional factors.
[0043] To further illustrate, assume that during the off-peak hours at night, the network bandwidth and computing resources in the data center are the most abundant. At this time, the system, through real-time monitoring of each virtual machine and combined with historical business data, intelligently determines that 11:00 p.m. is the best time to migrate virtual machine A to virtual machine B. This is because the network latency is low, the data transmission efficiency is the highest, and the number of business requests is small, which will not have a significant impact on the business.
[0044] Once the data migration schedule is determined, the data from the first VM will be transferred to the selected second VM according to this schedule, ensuring a smooth migration. This step ensures that the migration operation occurs within the optimal time window, minimizing disruption to business operations.
[0045] Through the embodiments provided in this application, careful early assessment and planning enable data migration to be performed at the optimal time without impacting business operations, thus avoiding the potential negative impact of performing migration operations during resource constraints or peak business periods. This approach not only improves data migration efficiency but also ensures the rationality of cloud platform resource allocation and the stability of business continuity.
[0046] As an optional solution, before transferring the data of the first virtual machine to the second virtual machine, the method further includes:
[0047] The first resource data and business data, as well as the second resource data and network link data, are input into the first decision model for data migration prediction to obtain a prediction result indicating the second virtual machine and the data migration time, wherein the first decision model is a first neural network model trained using migration data samples of historical virtual machines collected over a historical time period, and the migration data samples include resource data samples, business data samples, network link data samples, and data migration time samples of the historical virtual machines.
[0048] Optionally, in this embodiment, the first decision model refers to a neural network model trained by machine learning technology, which is used to predict the optimal target host (second virtual machine) and optimal migration time for virtual machine data migration to achieve intelligent scheduling and optimization of resources.
[0049] Optionally, in this embodiment, the migration data samples include historical virtual machine resource data, business data, network link data and actual data migration time. These samples are used to train the first decision model so that it can learn the rules and experience of historical migration decisions.
[0050] Optionally, in this embodiment, the historical time period refers to the time range of data collected and used to train the first decision model, which generally includes past resource usage records, business demand changes, network conditions, and results of migration operations.
[0051] Optionally, in this embodiment, before making real-time data migration decisions, a large amount of virtual machine migration data samples are collected over a historical period, including historical virtual machine resource usage, service types, network link status, and actual migration times. These data samples are used to train a first decision model, i.e., a first neural network model, to understand and predict the optimal migration decision based on current resource configurations, service requirements, and network conditions.
[0052] For example, over the past week, the data center recorded the CPU usage, memory usage, data processing type, network bandwidth consumption, and successful migration timing of all virtual machines. This data was used to train the first decision model, enabling it to predict the optimal migration target and time based on similar real-time data.
[0053] When resource load imbalance is detected and virtual machine data migration is required, the resource data and business data of the current first virtual machine, as well as the resource data and network link data of the candidate virtual machines, will be input into the first decision model. The model uses the learned rules to output prediction results, indicating the second virtual machine that is most suitable for receiving data migration and the optimal data migration time.
[0054] For example, if VM A's resource load is detected to be excessive, the system inputs A's CPU usage, memory consumption, and the type of data being processed, as well as the resource status of potential target VMs B and C, and the network conditions between A and B, and between A and C, into the first decision model. The model then predicts that VM B is the optimal migration target and should begin migration at 2:00 AM to avoid impacting services during peak hours.
[0055] Once the prediction result of the first decision model is obtained, the system will follow this instruction and start transferring the data of the first virtual machine to the second virtual machine at the predicted optimal time point to achieve balanced allocation and efficient utilization of resources.
[0056] The embodiments provided herein introduce a first decision model based on machine learning, significantly improving the intelligence and predictive accuracy of virtual machine data migration strategies. By learning from historical data, the model can better understand the complex relationships between dynamic resource load changes, fluctuating business demand, and network conditions that influence migration. This not only helps prepare for migrations before resource usage peaks, but also reduces resource pressure during peak hours.
[0057] As an optional solution, the first resource data and service data, as well as the second resource data and network link data, are input into the first decision model to perform data migration prediction, and a prediction result indicating the second virtual machine and data migration time is obtained, including:
[0058] Using the first decision model, similar calculations are performed on the first resource data and the service data, and the second resource data and the network link data to obtain an adaptation value between the first virtual machine and each candidate virtual machine, wherein the adaptation value is used to indicate a priority of each candidate virtual machine for receiving data from the first virtual machine;
[0059] The candidate virtual machine with the highest adaptation value is determined as the second virtual machine.
[0060] Optionally, in this embodiment, the adaptation value represents the degree of matching between the first virtual machine and the candidate virtual machine, which comprehensively reflects the matching of resource load, service type and network link status, and is used to evaluate the priority of the candidate virtual machine as a data migration target.
[0061] Optionally, in this embodiment, similarity calculation is used as a data processing technology to calculate the similarity or degree of adaptation between the first virtual machine and the candidate virtual machine by comparing the resource data, business data and network link data of the two, providing a basis for intelligent decision-making.
[0062] Optionally, in this embodiment, first, real-time collected first resource data (e.g., CPU usage, memory usage) and service data (e.g., service type, priority) of the first virtual machine, as well as second resource data of at least one candidate virtual machine and network link status data between the first virtual machine and the candidate virtual machine, are input into the first decision model. The model then utilizes a machine learning algorithm, such as a deep neural network, to perform a comprehensive analysis and similarity calculation on this data to determine a compatibility value between the first virtual machine and each candidate virtual machine.
[0063] For example, suppose there are virtual machines A and candidate virtual machines B and C in the system. The resource data, business data of A, and the resource data and network link data of B and C are input into the model. After calculation, the model concludes that the adaptation value of virtual machines A and B is 0.85, and the adaptation value with C is 0.75.
[0064] Optionally, in this embodiment, based on the adaptation value output by the first decision model, the system automatically identifies the candidate virtual machine with the highest adaptation value and determines it as the second virtual machine, that is, the best data migration target. This ensures that the migration operation is performed with the most adapted resources and minimal impact on the business.
[0065] To further illustrate, in the aforementioned adaptation value comparison, virtual machine B has the highest adaptation value, so the system determines it as the best migration target for receiving virtual machine A's data, ie, the second virtual machine.
[0066] Through the embodiments provided by this application, the intelligent decision-making process of virtual machine data migration is further refined by introducing the operation of "similar calculation". Not only the resource load and network link status are taken into consideration, but also the particularity of the business data is taken into consideration, so that the first decision model can evaluate the degree of matching between the first virtual machine and each candidate virtual machine from multiple dimensions and obtain a quantitative adaptation value. The calculation result of the adaptation value directly reflects the comprehensive priority of the candidate virtual machine as the recipient of data migration, thereby ensuring that the system can select the most suitable second virtual machine from many candidates, and realize accurate matching and efficient migration of resources.
[0067] As an optional solution, transferring data from the first virtual machine to the second virtual machine includes:
[0068] Performing differential calculation on the memory data of the first virtual machine and the disk data of the first virtual machine to generate a data difference file;
[0069] Transfer the data difference file to the second virtual machine.
[0070] Optionally, in this embodiment, differential calculation is used to compare the differences between two data sets and generate a data file containing only changed data, thereby reducing the amount of data transmission and improving transmission efficiency.
[0071] Optionally, in this embodiment, the data difference file is generated after the differential calculation, and contains only a collection of memory data and disk data that have changed since the last data synchronization, for efficient migration.
[0072] Optionally, in this embodiment, after determining the target host for data migration (the second virtual machine), the system quickly scans the memory and disk data of the first virtual machine and compares them with the snapshot from the last data synchronization. The differential calculation process identifies all data blocks that have changed since the last synchronization and generates a data difference file containing only these changed data, excluding unchanged data.
[0073] For example, it is detected that virtual machine A needs to be migrated to virtual machine B. Before the migration, a differential calculation is performed on A's memory and disk data. Suppose that 1MB of data in A's memory has changed, and a file on the disk has increased by 50KB. These changes are recorded in the data difference file, and the original data that has not changed is not transferred.
[0074] Subsequently, only the data difference files are transferred from the first virtual machine to the second virtual machine without transferring all memory and disk data, which greatly reduces the time and network resource consumption required for migration.
[0075] To illustrate further, in the above example, the data difference files of virtual machine A (containing 1MB of memory data changes and 50KB of disk data additions) are efficiently transferred to virtual machine B. By applying this difference data to the local virtual machine environment, B quickly completes the complete state replication of A, without the need to additionally transfer A's unchanged memory and disk data, greatly improving migration efficiency.
[0076] The embodiments provided herein significantly improve the efficiency and speed of virtual machine data migration by implementing differential computing and data difference file transfer. Differential computing technology focuses only on the changed portion of the data, and the resulting data difference file is much smaller than the virtual machine's full memory and disk data. This allows for rapid data migration even when network resources are limited or migration time is tight.
[0077] As an optional solution, transferring the data difference file to the second virtual machine includes:
[0078] transmitting the data difference file to the second virtual machine at the first transmission rate;
[0079] When it is detected that the network bandwidth utilization is greater than a third threshold during the data difference file transmission process, the first transmission rate is reduced to a second transmission rate, and the data difference file is transmitted to the second virtual machine at the second transmission rate.
[0080] Optionally, in this embodiment, the first transmission rate is an initially set data difference file transmission speed, which is usually set to a higher value based on the current network conditions and expected data transmission efficiency to speed up the migration process.
[0081] Optionally, in this embodiment, the third threshold is a preset network bandwidth utilization threshold. When it is detected that the network bandwidth utilization exceeds this threshold, it indicates that network resources are tight and the data transmission strategy needs to be adjusted to avoid negative impact on other services.
[0082] Optionally, in this embodiment, the second transmission rate is the data difference file transmission speed automatically reduced by the system when it is detected that the network bandwidth utilization exceeds the third threshold, so as to reduce network resource occupation and ensure the normal operation of other services.
[0083] Optionally, in this embodiment, after the target host (the second virtual machine) for data migration is determined and the data difference file is generated, the system first begins transferring the data difference file at a preset first transfer rate. This rate is typically set based on current network conditions and expected data transfer efficiency, aiming to quickly complete the data migration.
[0084] For example, to transfer the data difference file of virtual machine A to virtual machine B, the initial transmission rate is set to 80% of the link bandwidth. This rate selection is intended to complete the migration as quickly as possible and reduce the downtime of virtual machine A.
[0085] During the data difference file transmission process, the system detects the network bandwidth utilization in real time. If it finds that the bandwidth utilization exceeds the preset third threshold, the system will automatically adjust the transmission strategy, reduce the first transmission rate to the second transmission rate, and slow down the data transmission speed to prevent network congestion and ensure that other services are not affected.
[0086] For example, during the transfer of a data difference file, the system detects that network bandwidth utilization has risen to 85%, exceeding the third threshold of 80%. At this point, the system reduces the transmission rate from 80% of the link bandwidth to 50%, reducing network resource usage and ensuring the normal operation of VMs C and D without being affected by the migration.
[0087] After reducing the transmission rate, the system continues to transfer the data difference file at the second transmission rate until it is completely transferred to the second virtual machine. This strategy ensures smooth data migration while taking into account the rational allocation of network resources.
[0088] To further illustrate, the data difference file of virtual machine A is continued to be transmitted to virtual machine B at 50% of the link bandwidth rate. Although the transmission time is prolonged, the network stability is ensured during the migration process and interference with other virtual machine services is avoided.
[0089] The embodiments provided herein dynamically adjust the transmission rate of data difference files, achieving intelligent and efficient data migration while ensuring reasonable allocation of network resources and business continuity. When network resources are tight, the system can automatically reduce the data transmission speed to prevent network congestion and ensure the normal operation of other services.
[0090] As an optional solution, after transferring the data difference file to the second virtual machine, the method further includes:
[0091] Performing data verification on the data difference file received by the second virtual machine;
[0092] When the result of the data verification indicates that data loss occurs in the data difference file during the data transmission process, the data difference file is retransmitted to the second virtual machine.
[0093] Optionally, in this embodiment, data verification is a process of checking the integrity and accuracy of the transmitted data after the data transmission is completed, so as to ensure that the data is not damaged or lost during the transmission process.
[0094] Optionally, in this embodiment, data loss refers to the failure of some data to successfully reach the receiving end due to network errors, equipment failures or other reasons during the data transmission process, thereby affecting the integrity of the data and the accuracy of the migration.
[0095] Optionally, in this embodiment, after the data difference file is successfully transferred to the second virtual machine (i.e., the data migration target host), the system immediately starts a data verification program to perform an integrity check on the data difference file received by the second virtual machine to check for data loss or damage.
[0096] For example, after the data difference file of virtual machine A is transferred to virtual machine B, the data verification program on the B side runs to compare the received data difference file with the original difference file on the A side to check whether any data blocks were not correctly received during the transmission process.
[0097] If the data verification result shows that data loss occurs in the data difference file during the transmission process, the system will trigger the retransmission mechanism to retransmit the complete data difference file to the second virtual machine to ensure data integrity and migration accuracy.
[0098] To further illustrate, during the data verification process, if it is found that 50KB of data in the data difference file of virtual machine A is not correctly transmitted to virtual machine B, the 50KB of data will be automatically retransmitted until the data verification passes and it is confirmed that the data difference file is completely and successfully received.
[0099] The embodiments provided herein significantly enhance the reliability and accuracy of virtual machine data migration by implementing a data verification and retransmission mechanism. Data verification ensures the integrity of data difference files during transmission, while the retransmission mechanism effectively supplements data loss. Through this series of operations, the system can ensure the smooth completion of data migration even under poor network conditions or unexpected transmission errors.
[0100] As an optional solution, during the process of transferring data from the first virtual machine to the second virtual machine, the method further includes:
[0101] Obtain target operating state data of the second virtual machine and target network link data between the first virtual machine and the second virtual machine;
[0102] determining the target operating state data and the target network link data as load reference data;
[0103] When the load reference data satisfies the target load condition, a target operation corresponding to the target load condition is performed.
[0104] Optionally, in this embodiment, the target operating status data refers to the operating status data that the second virtual machine is expected to reach after receiving the data of the first virtual machine, including but not limited to indicators such as CPU usage, memory usage, disk IO rate, etc., to evaluate the load balancing situation after migration.
[0105] Optionally, in this embodiment, the target network link data is the expected network link quality data when establishing a data transmission channel between the first virtual machine and the second virtual machine, including parameters such as bandwidth, delay, and packet loss rate, which are used to evaluate the feasibility and efficiency of data transmission.
[0106] Optionally, in this embodiment, the load reference data is comprehensive data consisting of target operating state data and target network link data, serving as a reference for determining whether the load of the second virtual machine meets expected standards. The target load conditions are preset operating state and network link quality conditions. When the actual operating state and network link data of the second virtual machine meet these conditions, it indicates that the migration operation meets expectations, and subsequent optimization or adjustment operations can be performed.
[0107] Optionally, in this embodiment, during the data migration process from the first virtual machine to the second virtual machine, target operating status data of the second virtual machine (e.g., expected CPU usage) and target network link data between the first and second virtual machines (e.g., expected network latency) are collected simultaneously. These target data are used to form load reference data to determine whether the load after the migration meets expectations.
[0108] The collected load reference data is compared with the preset target load condition. If the running status and network link data of the second virtual machine both meet the target load condition, it means that the migration operation will not cause target host resource overload or network performance degradation.
[0109] Once it is confirmed that the load reference data meets the target load conditions, the system will perform optimization operations corresponding to the target conditions, such as adjusting virtual machine configurations, optimizing network transmission strategies, etc., to further improve the resource utilization and business response speed of the cloud environment.
[0110] Through the embodiments provided in this application, by collecting and evaluating target operating status data and target network link data in real time during the data migration process, effective load reference data is formed to ensure that the migration operation does not exceed the carrying capacity of the second virtual machine and can be carried out under ideal network conditions. Through this design, the system can not only intelligently judge and execute data migration operations, but also make corresponding optimization adjustments based on real-time feedback from load reference data, ensuring that resource utilization and network performance after migration meet the established goals, improving the intelligent level of cloud resource management and scheduling, and ensuring business continuity and efficient operation.
[0111] As an optional solution, when the load reference data meets the target load condition, executing the target operation corresponding to the target load condition also includes:
[0112] When the load reference data satisfies the first load condition, reducing the data transmission rate of the first virtual machine to the second virtual machine;
[0113] When the load reference data satisfies a second load condition, suspending data transmission from the first virtual machine to the second virtual machine;
[0114] When the load reference data satisfies a third load condition, determining a third virtual machine for data migration from the multiple virtual machines, and transferring the data of the first virtual machine to the third virtual machine;
[0115] The load intensity indicated by the first load condition is lower than the load intensity indicated by the second load condition, and the load intensity indicated by the second load condition is lower than the load intensity indicated by the third load condition.
[0116] Optionally, in this embodiment, the first load condition indicates the upper limit of the data transmission rate that the system can accept when the load of the target virtual machine (the second virtual machine) is relatively low. The second load condition indicates that when the load of the target virtual machine begins to increase but remains within an acceptable range, the system will suspend data transmission to ensure stable operation of the virtual machine. The third load condition indicates that if the load of the target virtual machine exceeds a preset threshold, meaning that data transmission may cause a serious performance degradation of the virtual machine, the system will search for and use another more suitable virtual machine (the third virtual machine) as the data transmission target to redirect the data migration operation.
[0117] Optionally, in this embodiment, load intensity refers to the extent to which a virtual machine consumes resources during operation, including but not limited to CPU usage, memory occupancy, and I / O load, etc., and is used to evaluate the real-time operating pressure of the virtual machine.
[0118] Optionally, in this embodiment, when the load reference data meets the first load condition, that is, the target virtual machine is still in a low load state after receiving the data transmission, the system automatically reduces the data transmission rate. This step can further optimize the allocation of network resources and ensure that other services are not affected.
[0119] For example, if the load intensity of target virtual machine B (such as CPU usage of 50%) is lower than the preset first load condition (such as CPU usage of 60%) when receiving data from A, the system will automatically adjust the data transmission rate from high-speed mode to medium-speed mode to prevent a sudden increase in B's load.
[0120] Once the load reference data reaches the second load condition, that is, the load of the target virtual machine begins to increase but is still controllable, the system will suspend data transmission to avoid further increasing the load of the virtual machine and ensure its operational stability and business continuity.
[0121] For example, when the load intensity of target VM B (e.g., CPU usage 65%) approaches a preset second load condition (e.g., CPU usage 70%), the system will suspend data transmission until the load decreases or B is able to handle the additional load.
[0122] If the load reference data meets the third load condition, that is, the load intensity of the target virtual machine exceeds the preset limit value, the system will select a third virtual machine with a lighter load from multiple candidate virtual machines as the new data migration target to avoid excessive burden on the target virtual machine.
[0123] For example, if the resource usage of the target virtual machine B (such as CPU usage reaches 90%) reaches the preset third load condition, the system will automatically switch to virtual machine C with lower resource usage and continue the data migration operation to ensure the smooth progress of the migration process.
[0124] Through the embodiments provided in this application, by implementing a dynamic data transmission control strategy based on three levels of load conditions, the intelligence and stability of the virtual machine data migration process are significantly enhanced. The system can adjust the data transmission rate according to the real-time load intensity of the target virtual machine, and even suspend transmission or redirect data to another virtual machine, avoiding the resource overload problem caused by data migration operations and ensuring the continuity and efficient operation of business in the cloud environment. This method of dynamically adjusting the migration strategy according to load intensity not only improves the flexibility of resource scheduling, but also optimizes the efficiency of network resource utilization.
[0125] As an optional solution, before executing the target operation corresponding to the target load condition, the method further includes:
[0126] Determining that the load reference data satisfies the first load condition when detecting that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to the fourth threshold, or when detecting that the target network link data indicates that the delay increase is higher than the first increase threshold;
[0127] determining that the load reference data satisfies a second load condition when detecting that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to the fifth threshold, or when detecting that the target network link data indicates that the delay increase is higher than the second increase threshold;
[0128] Determining that the load reference data satisfies a third load condition when detecting that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to the sixth threshold, or when detecting that the target network link data indicates that the delay increase is higher than the third increase threshold;
[0129] Among them, the sixth threshold is higher than the fifth threshold, the fifth threshold is higher than the fourth threshold, the third amplification threshold is higher than the second amplification threshold, and the second amplification threshold is higher than the first amplification threshold.
[0130] Optionally, in this embodiment, the fourth, fifth, and sixth thresholds are successively increasing resource load indicators, corresponding to the trigger thresholds of the first, second, and third load conditions, respectively, indicating different levels of target virtual machine resource usage during data transmission.
[0131] Optionally, in this embodiment, the first, second, and third increase thresholds are preset thresholds for network delay increase, corresponding to the trigger conditions of the first, second, and third load conditions, respectively, and are used to determine whether the network link status has reached a level that requires adjustment of the data transmission strategy.
[0132] Optionally, in this embodiment, resource load refers to the resource consumption of a virtual machine during operation, including CPU utilization, memory usage, disk I / O rate, etc., and is used to assess the virtual machine's carrying capacity and operational efficiency. Latency increase refers to the increase in network latency during data transmission, and is used to assess the impact of data transmission on network performance, ensuring service continuity and user experience.
[0133] Optionally, in this embodiment, if the target operating status data indicates that the resource load of the second virtual machine increases from the second threshold value until reaching a fourth threshold value, or if the target network link data indicates that the delay increase exceeds the first increase threshold value, the system determines that the load reference data meets the first load condition. In this case, the data transmission rate may be appropriately reduced to reduce the load pressure on the target virtual machine and prevent further increases in network delay.
[0134] If the target operating status data indicates that the resource load of the second virtual machine continues to increase and reaches the fifth threshold, or if the target network link data indicates that the latency increase exceeds the second increase threshold, the load reference data is determined to meet the second load condition. In this case, the system will suspend data transmission to avoid overloading the target virtual machine or severely degrading network performance.
[0135] If the target operating status data indicates that the resource load of the second virtual machine has risen to a sixth threshold, or if the target network link data indicates that the latency increase has exceeded a third increase threshold, the system determines that the load reference data meets the third load condition. At this point, the system identifies a third virtual machine from the multiple virtual machines for data migration and transfers the data of the first virtual machine to the third virtual machine to avoid resource overload or network performance issues on the original target host.
[0136] The embodiments provided herein establish different thresholds for resource load and network latency increases, creating a dynamic adjustment mechanism based on load reference data. This mechanism is used to monitor and control resource and network status during data migration in real time. When the target virtual machine's resource usage or network latency reaches a preset threshold, the system automatically determines the current load conditions and takes appropriate action, improving the intelligence and reliability of data migration.
[0137] As an optional solution, after obtaining the target operating state data of the second virtual machine and the target network link data between the first virtual machine and the second virtual machine, the method further includes:
[0138] The target operating status data and the target network link data are input into the second decision model for operation prediction to obtain the target operation output by the second decision model, wherein the second decision model is used to calculate the target operating status data and the target network link data to obtain reference values of multiple reference operations, and the reference values are used to indicate the priority of the multiple reference operations as the target operation, and the target operation is the operation with the highest reference value among the multiple reference operations.
[0139] Optionally, in this embodiment, the second decision model is a machine learning-based model that predicts and prioritizes possible actions during data migration. By analyzing target operational status data and target network link data, it outputs a series of reference values for reference actions to guide subsequent data transfer and optimization strategies.
[0140] Optionally, in this embodiment, the target operation is the operation instruction with the highest priority and the most suitable for the current scenario based on the prediction result of the second decision model, such as adjusting the data transmission rate, pausing transmission, or redirecting data to another virtual machine.
[0141] Optionally, in this embodiment, the reference operation is a set of possible operation options derived from analysis by the second decision model, encompassing resource management and network optimization strategies at different levels. The reference value is a numerical value derived from the second decision model's evaluation of each reference operation, reflecting the suitability of the operation as a target operation. A higher reference value indicates a higher priority.
[0142] Optionally, in this embodiment, the acquired target operating status data (such as the CPU utilization and memory occupancy of the second virtual machine) and target network link data (such as network delay and bandwidth utilization) are used as input and sent to a pre-trained second decision model for calculation and prediction.
[0143] The second decision model uses a predefined algorithm based on the input data to calculate a series of reference values for reference operations. Each reference value represents the rationality and priority of the corresponding operation as a target operation, with a higher reference value indicating a more suitable target operation.
[0144] The second decision model automatically selects the operation with the highest reference value as the target operation based on the calculated reference value, which is used to guide the next step of data migration strategy adjustment.
[0145] Through the embodiments provided by this application, by introducing a second decision model, the operating status data and network link data of the target virtual machine are deeply analyzed and the operation prediction is carried out, which can intelligently output the target operation with the highest priority and guide the dynamic adjustment of data migration. Compared with traditional static preset conditions, the second decision model can dynamically evaluate and decide the optimal operation based on real-time data, greatly improving the efficiency and stability of data migration, and reducing the complexity of resource scheduling and the risk of business interruption in cloud environments. This method combines the predictive capabilities of machine learning with the needs of dynamic resource management, providing strong technical support for the intelligent operation of cloud data centers.
[0146] As an optional solution, before inputting the target operating state data and the target network link data into the operation decision model for operation prediction, the method further includes:
[0147] Collecting N groups of migration records of historical virtual machines in a historical time period, wherein one group of migration records is used to indicate a first resource state of the historical virtual machine before performing a historical target operation on the historical virtual machine during the migration process, a second resource state of the historical virtual machine after performing the historical target operation, and expected parameters corresponding to the migration process;
[0148] Determine the first resource state, the second resource state, the expected parameters, and the historical target operation as a set of operation training samples corresponding to a set of migration records;
[0149] Using N groups of operation training samples corresponding to the N groups of migration records, the initialized second neural network model is trained;
[0150] When the loss function value corresponding to the trained second neural network model is less than a preset loss threshold, the training is determined to be completed, and the trained second neural network model is determined as the second decision model.
[0151] Optionally, in this embodiment, the historical time period is a specific time window for collecting historical virtual machine migration data, which will be used to train the second neural network model to improve its decision-making accuracy.
[0152] Optionally, in this embodiment, N groups of migration records are a collection of multiple virtual machine migration operations recorded within a historical time period. Each group of records contains resource status changes before and after migration and specific expected parameters during the migration process, serving as the data basis for model training.
[0153] Optionally, in this embodiment, the first resource status is the resource usage of the historical virtual machine before the data migration operation begins, including but not limited to CPU utilization, memory occupancy, disk I / O rate, etc. The second resource status is the resource usage of the historical virtual machine after the data migration operation is completed, reflecting the impact of the migration on the virtual machine resources.
[0154] Optionally, in this embodiment, the expected parameters are target parameters preset during the data migration process, such as the expected migration time, expected resource balance, or expected business performance indicators. The operation training sample is a dataset consisting of the first resource state, the second resource state, the expected parameters, and the historical target operations in a set of migration records, and is used to train the second neural network model to improve its prediction and decision-making capabilities. The preset loss threshold is a threshold used to evaluate the gap between the model's prediction accuracy and the expected value during the model training process. When the model reaches or falls below this threshold, the model training is considered complete and can be used for actual decision-making.
[0155] Optionally, in this embodiment, N groups of records of virtual machine migration within a historical time period are automatically or manually collected, each group of records containing the resource status of the historical virtual machine before migration, the resource status after migration, the expected parameters during the migration process, and the historical target operations performed for subsequent model training.
[0156] Each set of collected migration records is converted into corresponding operation training samples, that is, the resource status changes and expected parameters of the historical virtual machines are associated with the historical target operations to form structured data samples for model training.
[0157] The initialized second neural network model is trained using N groups of operation training samples, and the model parameters are continuously adjusted through the back-propagation algorithm, so that the model can predict which historical target operation is most in line with expectations under different resource states and expected parameters.
[0158] During the model training process, the loss function value of the model is continuously monitored. When it drops below the preset loss threshold, the model training is considered sufficient and can be used for actual decision-making. At this time, the trained second neural network model is determined as the second decision-making model.
[0159] Through the embodiments provided in this application, by collecting historical virtual machine migration records, operation training samples are constructed for training the second decision model, which is an intelligent decision-making model based on deep learning. During the model training process, the system not only focuses on historical resource status and operation results, but also integrates the expected parameters of migration, enabling the model to more accurately predict the optimal migration strategy to be adopted under the current resource status and expected parameters. By setting a preset loss threshold, the adequacy and effectiveness of model training are ensured, providing intelligent decision-making support for subsequent virtual machine data migration operations.
[0160] As an optional solution, before determining the first resource state, the second resource state, the expected parameters, and the historical target operation as a set of operation training samples corresponding to a set of migration records, the method further includes:
[0161] Obtaining a migration efficiency parameter, a service impact parameter, and a resource balancing parameter of the post-migration virtual machine corresponding to the migration process, wherein the migration efficiency parameter indicates the migration efficiency of the migration process, the service impact parameter indicates the degree of service performance degradation caused by the migration process, the post-migration virtual machine is the virtual machine that receives the migrated data during the migration process, and the resource balancing parameter indicates the degree of difference between the resource utilization of the post-migration virtual machine and the average resource utilization of multiple virtual machines;
[0162] Perform weighted calculations on migration efficiency parameters, business impact parameters, and resource balancing parameters to obtain expected parameters for the migration process.
[0163] Optionally, in this embodiment, the migration efficiency parameter is used to measure efficiency indicators during the virtual machine migration process, including data transmission rate, migration time, etc., reflecting the speed of the migration operation and resource utilization efficiency.
[0164] Optionally, in this embodiment, the business impact parameter represents the degree of impact of data migration on business performance degradation, such as increased business response time, throughput reduction ratio, etc., and is used to quantify the potential negative impact of the migration operation on business continuity and user experience.
[0165] Optionally, in this embodiment, the resource balancing parameter measures the difference between the resource utilization of the migrated virtual machine and the average resource utilization of the entire virtual machine cluster, ensuring reasonable allocation and utilization efficiency of resources and avoiding load imbalance caused by resource concentration on certain hosts.
[0166] Optionally, in this embodiment, weighted calculation is used to assign different weights to each parameter according to its importance, and the overall effect of the migration operation is comprehensively evaluated through mathematical operations. The obtained expected parameters reflect the comprehensive performance of the migration operation in terms of efficiency, business impact and resource balance.
[0167] Optionally, in this embodiment, after the virtual machine migration is completed, migration efficiency parameters, business impact parameters, and resource balancing parameters are collected. Migration efficiency parameters can be the average data transmission rate or total time consumed during the migration process; business impact parameters reflect the specific impact of the migration on business performance, such as how much the response time is extended; and resource balancing parameters are the difference between the average resource utilization of the virtual machine and the cluster after the migration, reflecting whether the resource allocation is balanced.
[0168] Weighted calculations are performed on each parameter to determine the expected parameters for the migration process. The selection of weights is crucial and should be adjusted based on the actual situation, prioritizing migration efficiency, business impact, and resource balance. This weighted approach allows for a comprehensive assessment of the overall benefits of the migration operation, providing data support for subsequent decision-making.
[0169] For example, assuming the migration efficiency parameter weight is 0.4, the business impact parameter weight is 0.5, and the resource balancing parameter weight is 0.1, substitute these values into the weighted calculation to obtain the overall expected parameters for the migration process.
[0170] The calculated expected parameters, resource status data before and after migration, and historical target operations are combined to form a complete set of operation training samples for subsequent training of the second decision model.
[0171] Through the embodiments provided in this application, by introducing migration efficiency parameters, business impact parameters and resource balancing parameters, a comprehensive evaluation of the overall effect of the virtual machine migration process is carried out, and the expected parameters for measuring the comprehensive performance of the migration operation are obtained through weighted calculation. This parameter integrates the performance of the migration operation in three dimensions: efficiency, business impact and resource balance, and is an important component of constructing operation training samples. By incorporating the expected parameters and historical resource status and operations into the training samples, a more accurate second decision model can be trained, so that when making decisions on migration strategies, it can better balance the speed of virtual machine migration, business continuity and uniformity of resource allocation, thereby improving the operational efficiency and user experience of cloud data centers.
[0172] As an optional solution, obtain the migration efficiency parameters corresponding to the migration process, including:
[0173] Get the total migration duration corresponding to the migration process, and get the elapsed migration duration when executing the target operation during the migration process;
[0174] The difference between the total migration time and the migration time is divided by the total migration time to obtain the migration efficiency parameter.
[0175] Optionally, in this embodiment, the total migration duration is the entire duration from the start of migration of the virtual machine to the complete migration to the target host, including the sum of the time of all migration-related operations such as data transmission and state adjustment.
[0176] Optionally, in this embodiment, the migrated time is the cumulative time from the start of the target operation (such as adjusting the data transmission rate, pausing the transmission, or rescheduling the migration target) to the completion of the operation during the migration process. This time may be affected by the execution of the operation. For example, pausing the transmission will extend the migrated time.
[0177] Optionally, in this embodiment, the migration efficiency parameter is an indicator for measuring the efficiency of the migration operation by calculating the ratio of the difference between the total migration time and the migrated time to the total migration time. The larger the value of this parameter, the smaller the negative impact of the operations during the migration process on the overall migration efficiency, that is, the higher the migration efficiency.
[0178] Optionally, in this embodiment, after the virtual machine data migration is completed, the system automatically records the total time taken for the entire migration process, i.e., the total migration duration; at the same time, it records the cumulative time for executing the target operation (such as adjusting the transmission rate, etc.) during the migration process, i.e., the migration duration.
[0179] Assume that the total time it takes to migrate virtual machine A to B is 120 seconds, and the cumulative time it takes to adjust the transmission rate is 40 seconds.
[0180] Subtract the migrated time (40 seconds) from the total migration time (120 seconds), and divide the difference (80 seconds) by the total migration time (120 seconds). The result (approximately 0.67) is used as the migration efficiency parameter, which indicates the proportion of the target operation's impact on the total migration time in this migration.
[0181] For example, through the above calculation, the migration efficiency parameter is obtained as 0.67, which means that in the total migration time, about 67% of the time is not directly affected by the target operation. It can be regarded as effective migration time and the migration efficiency is relatively high.
[0182] Through the embodiments provided in this application, a migration efficiency parameter is calculated by comparing the total migration duration with the actual migration duration. This parameter calculation method takes into account the actual impact of the target operation on the migration duration during the migration process and can effectively reflect the time efficiency of the migration operation.
[0183] As an optional solution, before determining the trained second neural network model as the operation decision model, the method further includes:
[0184] Obtaining expected parameters corresponding to each group of migration records and target expected parameters corresponding to each group of migration records during the training of the initialized second neural network model;
[0185] The loss function value corresponding to the trained second neural network model is determined based on the difference information between the target expected parameters corresponding to each group of migration records and the expected parameters corresponding to each group of migration records.
[0186] Optionally, in this embodiment, the target expectation parameter is a preset ideal migration result parameter during the machine learning model training process. It reflects the comprehensive expectations of the efficiency, business impact, and resource balance of the migration operation under ideal conditions. It is the target value for model training and is used to assess the gap between the model output and the ideal state.
[0187] Optionally, in this embodiment, the difference information is the difference between the expected parameters output by model training and the target expected parameters, which is usually expressed as a numerical difference or ratio, and is used to quantify the accuracy and effectiveness of the model prediction results.
[0188] Optionally, in this embodiment, the loss function value is an indicator that measures the difference between the model's predicted value and the target value. A smaller loss function value indicates that the model's predicted result is closer to the target expectation and the model training effect is better. In neural network model training, the loss function is a guide to the optimization direction and is used to guide the adjustment of model parameters.
[0189] Optionally, in this embodiment, each time the second neural network model is trained using an operation training sample, the expected parameters output by the model (the migration effect parameters predicted by the current training sample) and the pre-set target expected parameters (the preset ideal migration effect parameters) are recorded.
[0190] For example, suppose that in a certain training round, the model predicts migration efficiency parameters, business impact parameters, and resource balancing parameters as 60%, 10%, and 1%, respectively, while the preset target expected parameters are 65%, 5%, and 0.5%, respectively.
[0191] Based on the difference between the expected parameters output by the training sample in each training session and the preset target expected parameters, the loss function is used to calculate the error or deviation of the model prediction. This value reflects the gap between the current training state of the model and the ideal state.
[0192] For example, using squared error as the loss function to calculate the difference between the two parameter sets above, the resulting loss function value is 0.007. This means that the model's prediction results deviate from the preset target to a certain extent, but are still within an acceptable range overall.
[0193] The progress and effectiveness of model training are evaluated by continuously comparing the loss function values after each training iteration. When the model is sufficiently trained, resulting in the loss function value falling below the preset threshold, it indicates that the model has reached a satisfactory training state and can be used as an operational decision model.
[0194] For example, after multiple iterations, the loss function value stabilizes below 0.0005, which is far lower than the preset loss threshold (such as 0.01). At this time, the system determines that the model training is completed and determines the trained second neural network model as the operation decision model.
[0195] Through the embodiments provided in this application, by comparing the expected parameters predicted by the model with the preset target expected parameters and calculating the value of the loss function, the accuracy and convergence of the model training can be monitored in real time. This process ensures that the operational decision model can make the most ideal migration decision based on historical migration data when it is finally applied. This takes into account both migration efficiency and business impact and resource balance, providing a reliable technical guarantee for the intelligent online migration of virtual machines in cloud data centers.
[0196] As an optional solution, the present application also provides a virtual machine data migration system, which is characterized by including:
[0197] a resource monitoring module, configured to continuously monitor resource loads of the plurality of virtual machines and identify a first virtual machine having a resource load exceeding a first threshold and at least one candidate virtual machine having a resource load below a second threshold, wherein the first threshold is higher than the second threshold;
[0198] an intelligent decision-making module, communicatively connected to the resource monitoring module, configured to receive first resource data and service data of the first virtual machine, second resource data of at least one candidate virtual machine, and network link data between the first virtual machine and the at least one candidate virtual machine, and determine, using a machine learning algorithm, a second virtual machine for data migration from the at least one candidate virtual machine;
[0199] The data transmission module is in communication with the intelligent decision module and is used to execute data transmission according to the selection of the second virtual machine indicated by the intelligent decision module, so as to transmit the data of the first virtual machine to the second virtual machine.
[0200] Optionally, in this embodiment, the resource monitoring module is a system component responsible for real-time monitoring of the resource usage (such as CPU, memory occupancy, etc.) of each virtual machine in the cloud data center, and can automatically identify the first virtual machine with excessively high resource load and the candidate virtual machines with relatively light resource load.
[0201] Optionally, in this embodiment, the intelligent decision-making module is connected to the resource monitoring module to collect monitoring data, analyze the data through an integrated machine learning algorithm, and determine which candidate virtual machines are most suitable for receiving the data migration of the first virtual machine, thereby achieving dynamic balancing and optimization of resources.
[0202] Optionally, in this embodiment, the data transmission module is used to perform specific migration operations according to the instructions of the intelligent decision-making module, transfer the data of the first virtual machine to the selected second virtual machine, and optimize the data transmission strategy, such as adjusting the transmission rate or using differential transmission, to improve migration efficiency and reduce business impact.
[0203] Optionally, in this embodiment, the first threshold and the second threshold are preset resource load level standards used to determine whether a virtual machine needs to be migrated and which hosts can be used as migration targets. Typically, a higher first threshold indicates resource shortage, while a relatively lower second threshold indicates resource abundance.
[0204] Optionally, in this embodiment, the resource monitoring module continuously monitors the CPU and memory usage of all virtual machines in the cloud data center. For example, if the CPU usage of virtual machine X reaches 85%, exceeding a preset first threshold (assuming 80%), it is marked as the first virtual machine with excessive resource load. Meanwhile, the CPU usage of virtual machine Y is only 50%, below a second threshold (assuming 60%), and is considered a candidate virtual machine with low resource load.
[0205] After receiving information from the resource monitoring module, the intelligent decision-making module further collects detailed resource data and service type data for the first VM X, as well as resource data and network link data (such as latency and bandwidth utilization) for candidate VMs Y and Y. Using deployed machine learning algorithms, such as gradient boosting decision trees or deep neural networks, the intelligent decision-making module analyzes this data, assesses the migration suitability of each candidate VM, and ultimately selects VM Y, whose resource status and network conditions best match, as the second VM to receive X's data.
[0206] Once the intelligent decision module identifies the second virtual machine, Y, the data transfer module is activated to migrate data from virtual machine X to Y. This process may be accompanied by dynamic adjustments to the transmission rate. For example, data transmission may be performed at a high speed in the initial stage. When network bandwidth usage becomes too high, the rate is automatically reduced to prioritize the transmission of critical data, ensuring the efficiency and security of the migration operation.
[0207] Through the embodiments provided in this application, the resource monitoring module dynamically monitors the resource status of all virtual machines. The intelligent decision-making module comprehensively analyzes and makes decisions based on the monitoring data using machine learning algorithms, selecting the most suitable candidate virtual machines for data reception. The data transmission module then performs data migration operations and optimizes transmission strategies. This system architecture aims to achieve dynamic resource balancing, improve migration efficiency, and minimize business impact, thereby enhancing the operational efficiency and user experience of cloud data centers.
[0208] As an alternative solution, the aforementioned virtual machine data migration method is applied to the scenario of dynamic and intelligent online migration of virtual machines. With the widespread adoption of cloud computing technology, virtual machines (VMs) have become essential components of data centers and cloud computing environments. However, as services grow and change, the resource requirements of VMs are also constantly changing. To ensure business continuity and efficiency, dynamic online migration of VMs has become an important means of achieving flexible cloud resource scheduling, load balancing, and energy conservation. However, existing dynamic online migration technologies for VMs still have many problems. First, migration decisions lack comprehensive consideration of multiple factors. They are typically based solely on the VM's current CPU and memory resource utilization, failing to fully account for factors such as service type, network bandwidth usage, and real-time load fluctuations on the target host. This can lead to irrational migration decisions, resulting in resource imbalances after migration or excessive migration frequency. Second, data transmission efficiency is low during the migration process. Traditional full data migration methods can consume a large amount of network bandwidth in a short period of time, impacting the normal operation of other services. Third, insufficient monitoring and adjustment of the VM's operating status during the migration process prevents timely response to unexpected situations such as network jitter and temporary resource shortages on the target host, potentially leading to migration failures or business interruptions. Therefore, there is an urgent need for an improved virtual machine dynamic online migration technology to solve the above-mentioned shortcomings.
[0209] This embodiment designs a method for dynamic online migration of virtual machines based on the above-mentioned virtual machine data migration method, proposes an intelligent migration decision model based on multi-dimensional decision factors, determines the optimal migration time based on this, and designs an efficient data migration method based on differential transmission and flow control. While ensuring the efficiency of virtual machine migration, it minimizes the network impact on other businesses and ensures the stability of the migration process. In addition, a dynamic monitoring and adaptive adjustment mechanism is designed during the migration process to ensure the smooth progress of the virtual machine migration process and minimize the impact on running businesses.
[0210] Optionally, in this embodiment, an intelligent migration decision model based on multi-dimensional decision factors is constructed. First, multi-dimensional data is collected, including virtual machine resource usage (CPU utilization, memory utilization, disk I / O rate, network traffic, etc.), service type (real-time requirements, priority, etc.), the resource status of the migration target host (remaining CPU, memory resources, current load change trends), and network link status (bandwidth utilization, latency, packet loss rate). This data is then analyzed and processed using a machine learning algorithm (such as the gradient boosting decision tree algorithm) to train a migration decision model. Based on pre-defined migration trigger conditions (such as source host resource utilization exceeding a threshold, changes in service requirements, etc.), the decision model comprehensively assesses the necessity and feasibility of migration, selects the optimal migration target host, and determines the optimal migration timing, thus avoiding blind migration and unreasonable resource scheduling.
[0211] Optionally, in this embodiment, an efficient data migration method based on differential transmission and flow control is proposed. Before virtual machine migration (within the iterative migration cycle), a differential calculation is first performed on the memory and disk data of the source virtual machine, and only the data that has changed since the last data synchronization is transmitted, thereby reducing the amount of data transmitted. At the same time, the data transmission rate is dynamically adjusted in combination with the network link status. When the network bandwidth is tight, the transmission rate is reduced, and a priority queue is used to sort the data, giving priority to the transmission of critical data that ensures the normal operation of the virtual machine (such as operating system core data and business data being processed); when the network bandwidth is sufficient, the transmission rate is increased to speed up the migration process. In this way, while ensuring the efficiency of virtual machine migration, the network impact on other businesses is reduced, ensuring the stability of the migration process.
[0212] Optionally, in this embodiment, a dynamic monitoring and adaptive adjustment mechanism for the migration process is designed. During the virtual machine migration process, the status of the source host, the target host, and the network link are monitored in real time. The monitoring module is used to collect data such as CPU usage, memory load, network delay, packet loss rate, etc., and these data are fed back to the migration control module in real time. When it is detected that network jitter causes an increase in data transmission delay, or the target host is resource-constrained due to a sudden increase in load, the migration control module dynamically adjusts the migration strategy based on an adaptive adjustment algorithm based on reinforcement learning. For example, suspend some non-critical data transmission to give priority to core data transmission for virtual machine operation; or re-evaluate and switch to a more suitable migration target host to ensure the smooth progress of the virtual machine migration process and minimize the impact on the business.
[0213] Optionally, in this embodiment, the flowchart of the intelligent migration decision implementation steps based on multi-dimensional decision factors is as follows: Figure 3 As shown in the figure, the data collection module regularly collects operating data from Host 1 and Host 2 and then stores this data in a database. Historical data is regularly retrieved from the database to train and optimize the data model, improving the model's decision accuracy. The real-time monitoring module continuously monitors the operating status of hosts (such as Host 1 and Host 2) to determine whether the trigger conditions for virtual machine migration are met. If a host meets the migration trigger conditions, it passes the relevant information to the migration decision module. The migration decision module uses the trained model and, combining multi-dimensional data (such as the resource status of each host), analyzes and selects the optimal host to receive the virtual machine migration. Based on the optimal host selected by the migration decision module, the migration operation is executed for the virtual machine (such as virtual machines 1, 2, 3, and 4 that meet the migration conditions), completing the entire migration process.
[0214] Optionally, in this embodiment, a data collection module is deployed on the cloud data center's management platform to collect virtual machine resource usage data and service type information at fixed intervals (e.g., 30 seconds). It also obtains resource status data and network link status data for the migration target host and stores this data in a database. The migration decision model is trained and optimized using a gradient boosting decision tree using historically collected data on a regular basis (e.g., weekly). Model parameters are adjusted based on actual migration feedback to improve the model's decision accuracy. When the migration trigger condition is met (e.g., the source host's CPU utilization exceeds 80% for five consecutive minutes), the migration decision module invokes the trained model, comprehensively analyzes multi-dimensional data, selects the optimal migration target host from candidate target hosts, and determines the optimal migration start time.
[0215] Optionally, in this embodiment, the steps for implementing efficient data migration based on differential transmission and flow control include the following: After determining the migration target host, the migration control module initiates a data preprocessor to perform differential calculations on the source virtual machine's memory and disk data, generating a data difference file. The data transmission module dynamically adjusts the data transmission rate based on real-time collected network link status data. At the beginning of the migration, if the network bandwidth is sufficient, data is transmitted at a higher rate (e.g., 80% of the link bandwidth); when it is detected that the network bandwidth utilization exceeds 70%, the transmission rate is reduced, and critical data is prioritized according to the data priority queue to ensure basic operation of the virtual machine. After the data is transmitted to the target host, the data is verified. If data loss or errors are found, the corresponding data is promptly retransmitted from the source host to ensure the integrity of the virtual machine data.
[0216] Optionally, in this embodiment, the flowchart of the steps of dynamic monitoring and adaptive adjustment of the migration process is as follows: Figure 4As shown, the data collection module collects real-time data such as CPU usage, memory usage, network latency, and packet loss rate from the monitoring components of host 1 (running virtual machines 1 and 2) and host 2 (running virtual machines 3 and 4). It also performs differential calculations on memory and disk data. During the migration process, the migration control module monitors the resource status of each host and other resources in real time to determine whether resources exceed thresholds. If not, the data processing module generates differential data and passes it to the data transmission module for transmission via the migration data transmission channel to the target host (e.g., host 2). If resources exceed the threshold, the adaptive adjustment phase begins: Action 1: Adjust the transmission rate to optimize data transmission efficiency; Action 2: Complete the migration or terminate and roll it back to ensure system stability; Action 3: Reschedule the migration to a suitable target host to ensure a smooth migration. After the data transmission module transfers data, the data verification module verifies the data and, if any data loss or other issues are detected, provides feedback to ensure data accuracy. Once the data transmission and verification phases are successfully completed, the migration ends. If necessary, the new host can repeat the migration process. The entire process revolves around virtual machine migration, and achieves efficient and reliable virtual machine migration through data collection, control, processing, transmission, verification and adaptive adjustment.
[0217] Optionally, in this embodiment, monitoring sensors and modules are deployed on the source host, target host, and key network link nodes. These sensors collect data such as CPU usage, memory usage, network latency, and packet loss rate at 10-second intervals and transmit it to the migration control module in real time. When the migration control module detects a sudden increase in network latency exceeding a preset threshold (e.g., a 50% increase in latency), or a short-term rise in the target host's memory usage to 90%, an adaptive adjustment algorithm based on reinforcement learning is triggered. The agent selects an appropriate action based on its current state (e.g., reducing the data transmission rate by 30%) and, after executing the action, observes the new state and reward, continuously optimizing and adjusting its strategy to ensure a smooth migration. Once all data transfers are complete and the virtual machine successfully boots and operates normally on the target host, the migration control module confirms the migration is complete and updates the cloud data center's resource management information.
[0218] Optionally, in this embodiment, the virtual machine migration adaptive adjustment algorithm based on reinforcement learning is designed as follows.
[0219] Define resource status: Let S_t be the migration status at time t, which includes the resource status of the source host S_{source,t} (CPU usage u_{source,t}^{CPU}, memory usage u_{source,t}^{memory}, etc.), the resource status of the target host S_{target,t} (remaining CPU resources r_{target,t}^{CPU}, remaining memory resources r_{target,t}^{memory}, etc.), the network link status S_{network,t} (bandwidth utilization u_{network,t}, latency d_{network,t}, packet loss rate l_{network,t}), and the VM migration progress P_t. That is, S_t = [S_{source,t}, S_{target,t}, S_{network,t}, P_t].
[0220] Define execution actions: A_t includes operations such as adjusting the data transmission rate (such as increasing by n% or decreasing by m%), pausing / resuming partial data transmission, and selecting and switching a new migration target host.
[0221] Design reward function: The reward R_t comprehensively considers migration efficiency, business impact and resource utilization. Suppose the total time taken for virtual machine migration is T_{total}, the migration time is T_t, the performance degradation index of the business caused by migration is D_t (such as the increase in response time, the percentage of throughput decrease, etc.), and the resource balance of the target host after migration is B_t (measured by calculating the difference between the target host and the average resource utilization of the cluster).
[0222] The reward function is defined as: R_t=α*(T_{total}-T_t} / {T_{total})-β*D_t+γ*B_t, where α, β, and γ are weight coefficients, which are adjusted according to actual business needs to balance migration efficiency, business performance, and resource balance.
[0223] Reinforcement Policy Learning: A Deep Q-Network (DQN) is used for reinforcement policy learning. The designed agent uses a neural network to predict the Q-value (Q(S_t, A_t)) of each action based on the current state S_t and selects the action with the highest Q-value. After executing the action, it observes the new state S_{t+1} and the reward R_t obtained. The experience (S_t, A_t, R_t, S_{t+1}) is stored in the experience replay buffer D. A batch of experience is randomly sampled from the experience replay buffer D for training. The neural network parameters θ are updated by minimizing the loss function using the mean squared error.
[0224] The loss function is defined as follows: L(θ)=E_(S_t,A_t,R_t,S_{t+1})∼D[(R_t+γmax_A′Q(S_{t+1},A_t′;θ-)-Q(S_t,A_t;θ)) 2 ], which measures the mean squared error between the current Q-network prediction and the target Q-value. The state transition sequence (S_t, A_t, R_t, S_{t+1}) is randomly sampled from the experience replay buffer D, meaning that the samples follow the distribution in D. Through continuous training and learning, the agent gradually optimizes the migration strategy and achieves adaptive adjustment of the virtual machine migration process.
[0225] Through the above innovations and algorithms, more intelligent, efficient and stable dynamic online migration of virtual machines can be achieved, which improves the utilization efficiency of cloud resources, reduces the risk of business interruption, and has important application value and economic benefits.
[0226] The embodiments provided herein enable intelligent migration decisions and determine the optimal migration timing. During data transmission, efficient data migration can be achieved based on differential transmission and flow control. Furthermore, dynamic monitoring and adaptive adjustment mechanisms can be used to ensure smooth virtual machine migration, minimizing the impact on ongoing services. This invention facilitates the dynamic balancing of user data center resources and improves service utilization efficiency.
[0227] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0228] This embodiment also provides a virtual machine data migration device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0229] Figure 5 is a structural block diagram of a data migration device for a virtual machine according to an embodiment of the present application, such as Figure 5 As shown, the device includes:
[0230] A first determining unit 502 is configured to determine a first virtual machine having a resource load higher than a first threshold from the plurality of virtual machines, and to determine at least one candidate virtual machine having a resource load lower than a second threshold from the plurality of virtual machines, wherein the first threshold is higher than the second threshold;
[0231] A second determining unit 504 is configured to determine a second virtual machine for data migration from the at least one candidate virtual machine by using the first resource data and service data of the first virtual machine, the second resource data of the at least one candidate virtual machine, and the network link data between the first virtual machine and the at least one candidate virtual machine;
[0232] The migration unit 506 is configured to transfer data of the first virtual machine to the second virtual machine.
[0233] As an optional solution, the device further includes:
[0234] A first determining module is configured to determine a data migration time for data migration by using the first resource data and the service data, and the second resource data and the network link data before transferring the data of the first virtual machine to the second virtual machine;
[0235] The migration unit 506 includes:
[0236] The first migration module is configured to transfer data of the first virtual machine to the second virtual machine according to a data migration time.
[0237] As an optional solution, the device further includes:
[0238] The first prediction module is used to input the first resource data and business data, as well as the second resource data and network link data, into the first decision model for data migration prediction before transferring the data of the first virtual machine to the second virtual machine, so as to obtain a prediction result indicating the second virtual machine and the data migration time, wherein the first decision model is a first neural network model trained using migration data samples of historical virtual machines collected in a historical time period, and the migration data samples include resource data samples, business data samples, network link data samples and data migration time samples of the historical virtual machines.
[0239] As an optional solution, the first prediction module includes:
[0240] a calculation submodule, configured to perform similar calculations on the first resource data and the service data, and the second resource data and the network link data, using the first decision model, to obtain an adaptation value between the first virtual machine and each candidate virtual machine, wherein the adaptation value is used to indicate a priority of each candidate virtual machine for receiving data from the first virtual machine;
[0241] The first determining submodule is used to determine the candidate virtual machine with the highest adaptation value as the second virtual machine.
[0242] As an optional solution, the migration unit 506 includes:
[0243] A differential calculation module, configured to perform differential calculation on the memory data of the first virtual machine and the disk data of the first virtual machine to generate a data difference file;
[0244] The second migration module is used to transfer the data difference file to the second virtual machine.
[0245] As an optional solution, the second migration module includes:
[0246] a migration submodule, configured to transmit the data difference file to the second virtual machine at the first transmission rate;
[0247] When it is detected that the network bandwidth utilization is greater than a third threshold during the data difference file transmission process, the first transmission rate is reduced to a second transmission rate, and the data difference file is transmitted to the second virtual machine at the second transmission rate.
[0248] As an optional solution, the device further includes:
[0249] a verification module, applied to perform data verification on the data difference file received by the second virtual machine after the data difference file is transmitted to the second virtual machine;
[0250] The retransmission module is used to retransmit the data difference file to the second virtual machine after the data difference file is transmitted to the second virtual machine if the result of data verification indicates that data loss occurs in the data difference file during the data transmission process.
[0251] As an optional solution, the device further includes:
[0252] A first acquisition module is configured to acquire target operating state data of the second virtual machine and target network link data between the first virtual machine and the second virtual machine during the process of transmitting data of the first virtual machine to the second virtual machine;
[0253] a second determining module, configured to determine target operating state data and target network link data as load reference data during the process of transmitting data of the first virtual machine to the second virtual machine;
[0254] The execution module is configured to execute a target operation corresponding to the target load condition when the load reference data satisfies the target load condition during the process of transmitting the data of the first virtual machine to the second virtual machine.
[0255] As an optional solution, the execution module further includes:
[0256] A first execution submodule is configured to reduce a data transmission rate of data transmitted from the first virtual machine to the second virtual machine when the load reference data satisfies a first load condition;
[0257] a second execution submodule, configured to suspend data transmission from the first virtual machine to the second virtual machine when the load reference data satisfies a second load condition;
[0258] a third execution submodule, configured to, when the load reference data satisfies a third load condition, determine a third virtual machine for data migration from the multiple virtual machines, and transfer the data of the first virtual machine to the third virtual machine;
[0259] The load intensity indicated by the first load condition is lower than the load intensity indicated by the second load condition, and the load intensity indicated by the second load condition is lower than the load intensity indicated by the third load condition.
[0260] As an optional solution, the device further includes:
[0261] a third determining module, configured to, before executing the target operation corresponding to the target load condition, determine that the load reference data satisfies the first load condition when it is detected that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to the fourth threshold, or when the target network link data indicates that the delay increase is higher than the first increase threshold;
[0262] a fourth determining module, configured to, before executing the target operation corresponding to the target load condition, determine that the load reference data satisfies the second load condition when it is detected that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to a fifth threshold, or when the target network link data indicates that the delay increase is greater than the second increase threshold;
[0263] a fifth determining module, configured to, before executing the target operation corresponding to the target load condition, determine that the load reference data satisfies the third load condition when it is detected that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to the sixth threshold, or when the target network link data indicates that the delay increase is greater than the third increase threshold;
[0264] Among them, the sixth threshold is higher than the fifth threshold, the fifth threshold is higher than the fourth threshold, the third amplification threshold is higher than the second amplification threshold, and the second amplification threshold is higher than the first amplification threshold.
[0265] As an optional solution, the device further includes:
[0266] The second prediction module is used to input the target operating status data of the second virtual machine and the target network link data between the first virtual machine and the second virtual machine into the second decision model for operation prediction after obtaining the target operating status data of the second virtual machine and the target network link data, so as to obtain the target operation output by the second decision model, wherein the second decision model is used to calculate the target operating status data and the target network link data to obtain reference values of multiple reference operations, and the reference values are used to indicate the priority of the multiple reference operations as the target operation, and the target operation is the operation with the highest reference value among the multiple reference operations.
[0267] As an optional solution, the device further includes:
[0268] a collection module configured to collect N sets of migration records of historical virtual machines in a historical time period before inputting the target operating state data and the target network link data into the operation decision model for operation prediction, wherein one set of migration records is configured to indicate a first resource state of the historical virtual machine before performing a historical target operation on the historical virtual machine during the migration process, a second resource state of the historical virtual machine after performing the historical target operation, and expected parameters corresponding to the migration process;
[0269] a sixth determination module, configured to determine the first resource state, the second resource state, the expected parameters, and the historical target operation as a set of operation training samples corresponding to a set of migration records before inputting the target operating state data and the target network link data into the operation decision model for operation prediction;
[0270] a training module for training an initialized second neural network model using N sets of operation training samples corresponding to the N sets of migration records before inputting the target operating state data and the target network link data into the operation decision model for operation prediction;
[0271] The seventh determination module is used to determine that the training is completed before the target operating status data and the target network link data are input into the operation decision model for operation prediction, and to determine the trained second neural network model as the second decision model when the loss function value corresponding to the trained second neural network model is less than a preset loss threshold.
[0272] As an optional solution, the device further includes:
[0273] a second acquisition module, configured to obtain, before determining the first resource state, the second resource state, the expected parameters, and the historical target operation as a set of operation training samples corresponding to a set of migration records, a migration efficiency parameter, a business impact parameter, and a resource balancing parameter of the post-migration virtual machine corresponding to the migration process, wherein the migration efficiency parameter is used to indicate the migration efficiency of the migration process, the business impact parameter is used to indicate the degree of business performance degradation caused by the migration process, the post-migration virtual machine is a virtual machine that receives the migrated data during the migration process, and the resource balancing parameter is used to indicate the degree of difference between the resource utilization of the post-migration virtual machine and the average resource utilization of multiple virtual machines;
[0274] A weighting module is used to perform weighted calculation on migration efficiency parameters, business impact parameters and resource balancing parameters before determining the first resource state, the second resource state, the expected parameters and the historical target operation as a set of operation training samples corresponding to a set of migration records to obtain the expected parameters corresponding to the migration process.
[0275] As an optional solution, the second acquisition module includes:
[0276] The acquisition submodule is used to obtain the total migration duration corresponding to the migration process, as well as the migration duration when executing the target operation during the migration process;
[0277] The second determining submodule is configured to determine a result obtained by subtracting the migration duration from the total migration duration and dividing the result by the total migration duration as a migration efficiency parameter.
[0278] As an optional solution, the device further includes:
[0279] A third acquisition module is configured to acquire, before determining the trained second neural network model as the operation decision model, expected parameters corresponding to each group of migration records and target expected parameters corresponding to each group of migration records during the training of the initialized second neural network model;
[0280] The eighth determination module is used to determine the loss function value corresponding to the trained second neural network model based on the difference information between the target expected parameters corresponding to each group of migration records and the expected parameters corresponding to each group of migration records before determining the trained second neural network model as the operation decision model.
[0281] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0282] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0283] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0284] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when run.
[0285] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0286] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0287] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0288] An embodiment of the present application further provides a computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores the computer program product, and when the computer program is executed by a processor, the steps of the method in each embodiment of the present application are implemented.
[0289] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0290] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0291] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A data migration method for a virtual machine, characterized in that: include: Determining a first virtual machine from a plurality of virtual machines, whose resource load is higher than a first threshold, and determining at least one candidate virtual machine from the plurality of virtual machines, whose resource load is lower than a second threshold, wherein the first threshold is higher than the second threshold; Determining a second virtual machine for data migration from the at least one candidate virtual machine by using the first resource data and service data of the first virtual machine, the second resource data of the at least one candidate virtual machine, and the network link data between the first virtual machine and the at least one candidate virtual machine; The data of the first virtual machine is transferred to the second virtual machine.
2. The method according to claim 1, characterized in that Before transferring the data of the first virtual machine to the second virtual machine, the method further includes: Determining a data migration time for the data migration by using the first resource data and the service data, and the second resource data and the network link data; The transferring the data of the first virtual machine to the second virtual machine includes: The data of the first virtual machine is transferred to the second virtual machine according to the data migration time.
3. The method according to claim 2, characterized in that Before transferring the data of the first virtual machine to the second virtual machine, the method further includes: The first resource data and the business data, as well as the second resource data and the network link data, are input into a first decision model for data migration prediction to obtain a prediction result indicating the second virtual machine and the data migration time, wherein the first decision model is a first neural network model trained using migration data samples of historical virtual machines collected over a historical time period, and the migration data samples include resource data samples, business data samples, network link data samples, and data migration time samples of the historical virtual machine.
4. The method according to claim 3, characterized in that The step of inputting the first resource data and the service data, as well as the second resource data and the network link data, into a first decision model to perform data migration prediction, and obtaining a prediction result indicating the second virtual machine and the data migration time, includes: Using the first decision model, performing similarity calculations on the first resource data and the service data, and on the second resource data and the network link data, to obtain an adaptation value between the first virtual machine and each candidate virtual machine, wherein the adaptation value is used to indicate a priority of each candidate virtual machine for receiving data from the first virtual machine; The candidate virtual machine with the highest adaptation value is determined as the second virtual machine.
5. The method according to claim 1, characterized in that The transferring the data of the first virtual machine to the second virtual machine includes: Performing differential calculation on the memory data of the first virtual machine and the disk data of the first virtual machine to generate a data difference file; The data difference file is transferred to the second virtual machine.
6. The method according to claim 5, characterized in that The transmitting the data difference file to the second virtual machine includes: transmitting the data difference file to the second virtual machine at a first transmission rate; When it is detected that the network bandwidth utilization is greater than a third threshold during the data difference file transmission process, the first transmission rate is reduced to a second transmission rate, and the data difference file is transmitted to the second virtual machine at the second transmission rate.
7. The method according to claim 5, characterized in that After transferring the data difference file to the second virtual machine, the method further includes: Performing data verification on the data difference file received by the second virtual machine; When the result of the data verification indicates that data loss occurs in the data difference file during the data transmission process, the data difference file is retransmitted to the second virtual machine.
8. The method according to claim 1, characterized in that During the process of transferring the data of the first virtual machine to the second virtual machine, the method further includes: Obtain target operating state data of the second virtual machine and target network link data between the first virtual machine and the second virtual machine; determining the target operating state data and the target network link data as load reference data; When the load reference data satisfies a target load condition, a target operation corresponding to the target load condition is performed.
9. The method according to claim 8, characterized in that When the load reference data satisfies the target load condition, executing the target operation corresponding to the target load condition further includes: When the load reference data satisfies a first load condition, reducing a data transmission rate of data transmitted from the first virtual machine to the second virtual machine; When the load reference data satisfies a second load condition, suspending data transmission from the first virtual machine to the second virtual machine; When the load reference data satisfies a third load condition, determining a third virtual machine for data migration from the multiple virtual machines, and transferring the data of the first virtual machine to the third virtual machine; The load intensity indicated by the first load condition is lower than the load intensity indicated by the second load condition, and the load intensity indicated by the second load condition is lower than the load intensity indicated by the third load condition.
10. The method according to claim 9, characterized in that Before executing the target operation corresponding to the target load condition, the method further includes: Determining that the load reference data satisfies the first load condition when detecting that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to a fourth threshold, or when detecting that the target network link data indicates that the delay increase is higher than the first increase threshold; Determining that the load reference data satisfies the second load condition when detecting that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to a fifth threshold, or when detecting that the target network link data indicates that the delay increase is higher than the second increase threshold; determining that the load reference data satisfies the third load condition when it is detected that the target operating state data indicates that the resource load of the second virtual machine increases from the second threshold to a sixth threshold, or when the target network link data indicates that the delay increase is higher than the third increase threshold; Among them, the sixth threshold is higher than the fifth threshold, the fifth threshold is higher than the fourth threshold, the third amplification threshold is higher than the second amplification threshold, and the second amplification threshold is higher than the first amplification threshold.
11. The method according to claim 8, characterized in that After acquiring the target operating state data of the second virtual machine and the target network link data between the first virtual machine and the second virtual machine, the method further includes: The target operating status data and the target network link data are input into a second decision model for operation prediction to obtain the target operation output by the second decision model, wherein the second decision model is used to calculate the target operating status data and the target network link data to obtain reference values of multiple reference operations, and the reference values are used to indicate the priorities of the multiple reference operations as the target operation, and the target operation is the operation with the highest reference value among the multiple reference operations.
12. The method according to claim 11, characterized in that Before inputting the target operating state data and the target network link data into the operation decision model for operation prediction, the method further includes: Collecting N groups of migration records of historical virtual machines in a historical time period, wherein one group of migration records is used to indicate a first resource state of the historical virtual machine before performing a historical target operation on the historical virtual machine during the migration process, a second resource state of the historical virtual machine after performing the historical target operation, and expected parameters corresponding to the migration process; determining the first resource state, the second resource state, the expected parameter, and the historical target operation as a set of operation training samples corresponding to the set of migration records; Training the initialized second neural network model using N groups of operation training samples corresponding to the N groups of migration records; When the loss function value corresponding to the trained second neural network model is less than a preset loss threshold, the training is determined to be completed, and the trained second neural network model is determined as the second decision model.
13. The method according to claim 12, characterized in that Before determining the first resource state, the second resource state, the expected parameters, and the historical target operation as a set of operation training samples corresponding to the set of migration records, the method further includes: Obtaining a migration efficiency parameter, a service impact parameter, and a resource balancing parameter of the post-migration virtual machine corresponding to the migration process, wherein the migration efficiency parameter is used to indicate the migration efficiency of the migration process, the service impact parameter is used to indicate the degree of service performance degradation caused by the migration process, the post-migration virtual machine is a virtual machine that receives migrated data during the migration process, and the resource balancing parameter is used to indicate the degree of difference between the resource utilization of the post-migration virtual machine and the average resource utilization of the multiple virtual machines; A weighted calculation is performed on the migration efficiency parameter, the service impact parameter, and the resource balancing parameter to obtain the expected parameter corresponding to the migration process.
14. The method according to claim 13, characterized in that The obtaining of a migration efficiency parameter corresponding to the migration process includes: Obtaining a total migration duration corresponding to the migration process, and obtaining an elapsed migration duration when executing the target operation during the migration process; A difference obtained by subtracting the completed migration duration from the total migration duration is divided by the total migration duration to obtain a result, which is determined as the migration efficiency parameter.
15. The method according to claim 13, characterized in that Before determining the trained second neural network model as the operation decision model, the method further includes: Acquire the expected parameters corresponding to each group of migration records and the target expected parameters corresponding to each group of migration records during the training of the initialized second neural network model; The loss function value corresponding to the trained second neural network model is determined based on difference information between the target expected parameters corresponding to each group of migration records and the expected parameters corresponding to each group of migration records.
16. A data migration system for a virtual machine, characterized in that: include: a resource monitoring module, configured to continuously monitor resource loads of the plurality of virtual machines and identify a first virtual machine having a resource load exceeding a first threshold and at least one candidate virtual machine having a resource load below a second threshold, wherein the first threshold is higher than the second threshold; an intelligent decision-making module, communicatively connected to the resource monitoring module, configured to receive first resource data and service data of the first virtual machine, second resource data of the at least one candidate virtual machine, and network link data between the first virtual machine and the at least one candidate virtual machine, and determine, using a machine learning algorithm, a second virtual machine for data migration from the at least one candidate virtual machine; A data transmission module is communicatively connected to the intelligent decision module and is used to execute data transmission according to the selection of the second virtual machine indicated by the intelligent decision module, and transmit the data of the first virtual machine to the second virtual machine.
17. A data migration device for a virtual machine, characterized in that: include: a first determining unit, configured to determine, from the plurality of virtual machines, a first virtual machine having a resource load higher than a first threshold, and to determine, from the plurality of virtual machines, at least one candidate virtual machine having a resource load lower than a second threshold, wherein the first threshold is higher than the second threshold; a second determining unit, configured to determine a second virtual machine for data migration from the at least one candidate virtual machine by using the first resource data and service data of the first virtual machine, the second resource data of the at least one candidate virtual machine, and the network link data between the first virtual machine and the at least one candidate virtual machine; The migration unit is configured to transfer data of the first virtual machine to the second virtual machine.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 15 when executed by a processor.
19. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Load balancing scheduling method, apparatus, and computer-readable storage medium
CN109213595A
Virtual machine scheduling method and device
CN114546602A
Dynamic resource scheduling method and system based on data center and storage medium
CN115373862A
Virtual machine scheduling method, device and equipment and computer storage medium
CN116932134A
Data center resource management system and application
CN117112142A
Cited By
Server release method
CN120994410A