Virtual machine data transmission method and device, equipment and medium

By distinguishing between hot and cold data in virtual machine data transmission and adjusting the transmission strategy according to the virtual machine's health status, the problems of virtual machine performance degradation and redundant transmission are solved, and the stable operation of the cloud platform is achieved.

CN120994312APending Publication Date: 2025-11-21JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511148981.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, frequent snapshot creation leads to decreased virtual machine performance, redundant transmission of cold data and low efficiency in merging incremental data, and rigid bandwidth allocation during concurrent transmission of virtual machines can easily cause congestion.

Method used

After enabling continuous data protection in the virtual machine, it is determined whether the data to be transmitted is hot or cold data, and different transmission strategies are adopted. Hot data is transmitted in real time, while cold data is transmitted incrementally or in full. The transmission strategy is adjusted according to the health status of the virtual machine to avoid redundant transmission.

Benefits of technology

This effectively avoids redundant transmission of cold data, solves the problem of virtual machine performance degradation caused by frequent snapshot creation, and lays the foundation for stable operation of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994312A_ABST
    Figure CN120994312A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual machine data transmission method and device, equipment and a medium, and relates to the technical field of cloud computing, and the method comprises the steps: judging whether to-be-transmitted data is hot data or not after a virtual machine starts continuous data protection; the hot data is data of which the virtual machine disk block access frequency is greater than a target frequency threshold; if the data is hot data, reading and transmitting the first to-be-transmitted data in real time; the first to-be-transmitted data comprises a hot data input / output log and a bitmap in the cache pool; if the data is not the hot data, whether a metadata mark exists in the reflink configuration file or not is judged, and corresponding second to-be-transmitted data is transmitted according to a corresponding judgment result; the metadata mark is a mark used for identifying whether the virtual machine has an effective reflink snapshot chain or not; in the data transmission process, the transmission strategy of the data to be transmitted is adjusted according to the health degree of the virtual machine, so that data transmission is completed. The problem that the performance of the virtual machine is reduced is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing technology, and in particular to a virtual machine data transmission method, apparatus, device, and medium. Background Technology

[0002] With the development of cloud computing, more and more enterprise users are placing higher demands on the efficiency and continuity of business processing when using virtualization platforms. Cloud platforms provide a disaster recovery solution, and existing solutions extract differential data through periodic snapshots. However, frequent snapshots can lead to a decline in virtual machine performance, and the efficiency of incremental data merging is low. Traditional CBT (Changed Block Tracking) only records block changes without combining data access pattern analysis, resulting in the repeated transmission of cold data. Rigid bandwidth allocation during concurrent transmission of virtual machines can easily cause congestion.

[0003] It is evident that avoiding redundant transmission of cold data and the performance degradation of virtual machines caused by frequent snapshot creation are problems that need to be solved by those skilled in the art. Summary of the Invention

[0004] The purpose of this invention is to provide a virtual machine data transmission method, apparatus, device, and medium that avoids redundant transmission of cold data and solves the problem of virtual machine performance degradation caused by frequent snapshot creation. The specific solution is as follows:

[0005] In a first aspect, the present invention discloses a virtual machine data transmission method, comprising:

[0006] After enabling continuous data protection in the virtual machine, determine whether the data to be transmitted is hot data; hot data is data whose virtual machine disk block access frequency is greater than the target frequency threshold.

[0007] If it is hot data, then read and transmit the first data to be transmitted in real time; the first data to be transmitted includes hot data input / output logs and bitmaps in the buffer pool;

[0008] If it is not hot data, then check whether there is a metadata tag in the reflink configuration file, and transmit the corresponding second data to be transmitted according to the corresponding judgment result; the metadata tag is a tag used to identify whether the virtual machine has a valid reflink snapshot chain;

[0009] During data transmission, the transmission strategy of the data to be transmitted is adjusted according to the health status of the virtual machine in order to complete the data transmission; the data to be transmitted includes the first data to be transmitted and the second data to be transmitted.

[0010] Optionally, before enabling persistent data protection for the virtual machine, the following are also included:

[0011] Add the virtual machine to the pre-created protection group to enable the protection function of the protection group and call the underlying asynchronous interface for data transmission to perform asynchronous data processing;

[0012] At the start of asynchronous data processing, the underlying interface is used to determine whether the virtual machine has enabled persistent data protection.

[0013] If continuous data protection is not enabled, perform the corresponding operation based on the virtual machine's state to enable continuous data protection.

[0014] Optionally, perform corresponding operations based on the virtual machine's state to enable continuous data protection, including:

[0015] If the virtual machine is currently powered off, a snapshot is generated directly, and continuous data protection for the virtual machine is enabled through the virtualization management library.

[0016] If the virtual machine is currently powered on, pause the virtual machine, create a reflink snapshot, enable continuous data protection for the virtual machine based on the reflink snapshot, and restart the virtual machine.

[0017] Optionally, virtual machine data transfer methods also include:

[0018] Create the virtual machine's reflink configuration file;

[0019] If the pre-created protection group stops protecting the virtual machine, then continuous data protection for the virtual machine is turned off, data transfer is stopped, and residual data in the snapshot chain is deleted via the application programming interface.

[0020] Optionally, determine whether metadata tags exist in the reflink configuration file, and transmit the corresponding second data to be transmitted based on the determination result, including:

[0021] If there are metadata tags in the reflink configuration file, the difference data between the two snapshots is obtained and transmitted through the underlying interface, and the difference data is merged to complete the transmission of the second data to be transmitted.

[0022] If it does not exist, the second data to be transmitted will be transmitted in full directly; the second data to be transmitted is cold data.

[0023] Optionally, the data transfer strategy can be adjusted based on the health status of the virtual machine to complete the data transfer, including:

[0024] Obtain target monitoring data for virtual machines; target monitoring data includes any one or a combination of several of the following: CPU utilization of the host machine corresponding to the virtual machine, storage processing read and write operation efficiency, disaster recovery backup system's disaster recovery storage write speed, and virtual machine's network bandwidth utilization.

[0025] The scores for each target monitoring data are determined according to pre-created scoring rules;

[0026] Data weights for each target monitoring data are assigned based on each score.

[0027] The overall health of the virtual machine is determined by weighted summation based on the scores and data weights.

[0028] The health level of a virtual machine is determined based on a comprehensive health score according to pre-created health level classification rules; different comprehensive health scores correspond to different health level registrations.

[0029] If the virtual machine's health level is at the highest level, the data to be transferred will be transmitted normally, and corresponding health reports will be generated periodically.

[0030] If a virtual machine is at the second health level, reduce the synchronization frequency of non-critical virtual machines; non-critical virtual machines are those with lower business continuity requirements than preset requirements, or those that allow for brief interruptions or data delays.

[0031] If the virtual machine is at the third health level, then pause the transmission of the second data to be transmitted, and continue transmitting the first data to be transmitted and metadata.

[0032] Optionally, the process of assigning data weights to the monitoring data of each target based on each score also includes:

[0033] If network congestion occurs, increase the data weight of network bandwidth utilization in the target monitoring data;

[0034] If a storage performance alarm is received, the weight of data related to storage processing read and write operation efficiency in the target monitoring data will be increased.

[0035] Secondly, the present invention discloses a virtual machine data transmission device, comprising:

[0036] The first judgment module is used to determine whether the data to be transmitted is hot data after the virtual machine enables continuous data protection; hot data is data whose virtual machine disk block access frequency is greater than the target frequency threshold.

[0037] The data reading module is used to read and transmit the first data to be transmitted in real time if it is hot data; the first data to be transmitted includes hot data input / output logs and bitmaps in the buffer pool;

[0038] The second judgment module is used to determine whether there is a metadata tag in the reflink configuration file if it is not hot data, and to transmit the corresponding second data to be transmitted according to the corresponding judgment result; the metadata tag is a tag used to identify whether the virtual machine has a valid reflink snapshot chain;

[0039] The transmission strategy adjustment module is used to adjust the transmission strategy of the data to be transmitted according to the health status of the virtual machine during the data transmission process in order to complete the data transmission; the data to be transmitted includes first data to be transmitted and second data to be transmitted.

[0040] Thirdly, the present invention discloses an electronic device, comprising:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute computer programs to implement virtual machine data transfer methods as described above.

[0043] Fourthly, the present invention discloses a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements the virtual machine data transmission method as described above.

[0044] As can be seen, after enabling continuous data protection in the virtual machine, this invention determines whether the data to be transmitted is hot data; hot data is data whose virtual machine disk block access frequency is greater than the target frequency threshold; if it is hot data, the first data to be transmitted is read and transmitted in real time; the first data to be transmitted includes hot data input / output logs and bitmaps in the cache pool; if it is not hot data, it determines whether there is a metadata tag in the reflink configuration file, and transmits the corresponding second data to be transmitted according to the corresponding determination result; the metadata tag is a tag used to identify whether the virtual machine has a valid reflink snapshot chain; during the data transmission process, the transmission strategy of the data to be transmitted is adjusted according to the health status of the virtual machine to complete the data transmission; the data to be transmitted includes the first data to be transmitted and the second data to be transmitted.

[0045] Beneficial effects: This invention avoids redundant transmission of cold data by using a heat classification system. Different transmission methods are used for data with different heat levels. Hot data is transmitted through continuous data protection technology, while cold data is transmitted through different methods. This solves the problem of virtual machine performance degradation caused by frequent snapshot creation and lays a solid foundation for the stable operation of the cloud platform. Attached Figure Description

[0046] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A flowchart of a virtual machine data transmission method provided in an embodiment of the present invention;

[0048] Figure 2 A schematic diagram of a scoring rule provided in an embodiment of the present invention;

[0049] Figure 3 A schematic diagram of weight allocation provided in an embodiment of the present invention;

[0050] Figure 4 This is a schematic diagram of a health grading and response strategy provided in an embodiment of the present invention;

[0051] Figure 5 A flowchart illustrating a specific virtual machine data transmission method provided in this embodiment of the invention;

[0052] Figure 6 This is a schematic diagram of a virtual machine data transmission device provided in an embodiment of the present invention;

[0053] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0055] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.

[0056] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0057] In some technical solutions, periodic snapshots are used to extract differential data. However, frequent snapshots can lead to decreased virtual machine performance and low efficiency in merging incremental data. Traditional CBT only records block changes without combining data access pattern analysis, resulting in repeated transmission of cold data. Rigid bandwidth allocation during concurrent virtual machine transmission can easily cause congestion. To address these technical problems, this invention discloses a virtual machine data transmission method, apparatus, device, and medium that can avoid redundant transmission of cold data and solve the problem of decreased virtual machine performance caused by frequent snapshot creation.

[0058] See Figure 1As shown, this embodiment of the invention provides a virtual machine data transmission method, including:

[0059] Step S11: After enabling continuous data protection in the virtual machine, determine whether the data to be transmitted is hot data; hot data is data whose virtual machine disk block access frequency is greater than the target frequency threshold.

[0060] In this embodiment of the invention, before enabling continuous data protection (CDP) on a virtual machine, the virtual machine is added to a pre-created protection group to enable the protection function of the protection group and to call the underlying asynchronous data transmission interface for asynchronous data processing. At the start of asynchronous data processing, the underlying interface determines whether continuous data protection is enabled on the virtual machine. If not, the corresponding operation is performed based on the virtual machine's state to enable continuous data protection. In other words, a protection group is first created, a virtual machine is added to the protection group, and the protection group is enabled. However, CDP is not immediately enabled for the virtual machines within the protection group during the initial protection activation task. A protection group is a logical container used to centrally manage the data protection policies (such as CDP) of a group of virtual machines. Users create protection groups (e.g., through a management interface or API). The target virtual machine is added to the protection group. The protection function of the protection group is enabled, triggering the subsequent CDP process. When enabling protection, the system only marks the protection group status and does not immediately start CDP. Instead, it triggers the underlying data transmission task through an asynchronous interface (such as a message queue) to avoid blocking the main thread. Then, in the call to the underlying asynchronous interface for starting data transmission, the underlying layer needs to determine whether CDP is enabled. Specifically, at the start of asynchronous processing, the underlying interface checks whether the virtual machine has CDP enabled: if it has, data transmission is started directly; if not, different operations are performed based on the virtual machine's state. Additionally, it should be noted that the virtual machine in this invention can be a disaster recovery virtual machine.

[0061] In this embodiment of the invention, before enabling Continuous Data Protection (CDP), if the virtual machine is currently powered off, a snapshot is directly generated, and CDP is enabled through the virtualization management library. If the virtual machine is currently powered on, the virtual machine is paused, a reflink snapshot is created, CDP is enabled based on the reflink snapshot, and the virtual machine is restarted. In other words, for a powered-off virtual machine: a snapshot is directly generated (no need to pause). CDP is started through libvirt (the virtualization management library) to begin capturing subsequent data changes. For a powered-on virtual machine, if CDP is not enabled, the virtual machine is paused, a reflink is created, CDP is enabled, and then the virtual machine is restored. Pausing the virtual machine ensures data consistency (similar to "transaction isolation" in databases). Creating a reflink snapshot utilizes the reflink feature of the file system to quickly generate a lightweight snapshot, sharing data blocks to save space. Enabling CDP starts CDP based on the snapshot. Restoring the virtual machine removes the pause and resumes operation.

[0062] In addition, this invention creates a reflink configuration file for the virtual machine. If the pre-created protection group stops protection, the continuous data protection of the virtual machine is turned off, data transmission is stopped, and residual snapshot chain data is deleted through the application programming interface. That is, a reflink configuration file (e.g., / datastore / xxxx / i-000xx / reflink.conf) is created under the storage pool. If the protection group stops protection, the CDP of the virtual machines within the protection group will be turned off, and ongoing data transmission will stop. An interface needs to be provided at the underlying level to clean up residual data. Specifically, an independent configuration file (e.g., / datastore / xxxx / i-000xx / reflink.conf) is generated for each virtual machine in the storage pool, recording metadata such as snapshot path and CDP status for locating data during subsequent recovery or cleanup. The protection stoppage process involves: turning off the CDP service of all virtual machines; interrupting the current data transmission task (e.g., deleting temporary files or terminating network connections); and resource reclamation, where the underlying level needs to provide an API (Application Programming Interface) (e.g., cleanup_reflink_data()) to delete associated temporary files, residual snapshot chain data, etc., to prevent storage leaks.

[0063] In summary, the above process describes a virtual machine data protection mechanism based on asynchronous interfaces and reflink snapshots. Its core is to achieve efficient and secure CDP functionality through state judgment, pause / resume operations, and metadata management.

[0064] After enabling continuous data protection in the virtual machine, the access frequency of the virtual machine disk blocks is analyzed, and the data is divided into hot data (frequently modified) and cold data (infrequently modified). CDP real-time log transmission is enabled only for hot data, while CBT timed batch synchronization is used for cold data.

[0065] Step S12: If it is hot data, then read and transmit the first data to be transmitted in real time; the first data to be transmitted includes hot data input / output logs and bitmaps in the buffer pool.

[0066] In this embodiment of the invention, if the data in the disaster recovery virtual machine is hot data, then the cache disk (hot data I / O (input / output) logs and bitmaps) are read sequentially from the CDP cache pool. The I / O logs and bitmaps in the cache storage pool are incremental data for the virtual machine and need to be transmitted to the other end.

[0067] Step S13: If it is not hot data, determine whether there is a metadata tag in the reflink configuration file, and transmit the corresponding second data to be transmitted according to the corresponding judgment result; the metadata tag is a tag used to identify whether the virtual machine has a valid reflink snapshot chain.

[0068] In this embodiment of the invention, if the data in the virtual machine is cold data, the system checks if `lastRefink` exists in `reflink.conf`. If it does, the difference between the two reflink snapshots is transmitted directly; otherwise, the cold data is transmitted in full. In other words, if metadata markers exist in the reflink configuration file, the difference data between the two snapshots is obtained and transmitted through the underlying interface, and the difference data is merged to complete the transmission of the second data to be transmitted; if it does not exist, the second data to be transmitted is transmitted in full directly; the second data to be transmitted is cold data. The above process describes a cold data transmission strategy based on the reflink snapshot mechanism. The core logic is to determine the data transmission method (incremental or full) by checking the metadata marker (`lastRefink`) in `reflink.conf`. Cold data refers to virtual machine data that has not been modified for a long time or has extremely low access frequency (such as archived files, historical logs, etc.), characterized by a large data volume but minimal changes. In disaster recovery scenarios, cold data needs to be synchronized from the production environment to the disaster recovery system, requiring a balance between efficiency and resource overhead. The configuration file path is ` / datastore / xxxx / i-000xx / reflink.conf` (as described in previous rounds). It records metadata about the snapshot chain, including the snapshot path and CDP status. When the `lastRefink` field exists: it indicates that the virtual machine has a valid reflink snapshot chain, and synchronization can be completed through two incremental transfers (the difference between the `lastRefink` snapshot and the current data), avoiding the overhead of a full transfer. When the `lastRefink` field does not exist: it indicates that no incremental snapshots are available, and only a full transfer of cold data is possible to ensure data integrity. Incremental transfer condition: Check if `lastRefink` exists in `reflink.conf`. If it exists, obtain the difference data between the two snapshots through the underlying interface (such as libvirt or the storage API) and transfer it to the target. Full transfer condition: When `lastRefink` is not present, start a full data transfer task, directly reading the complete data block from the source. This is suitable for initial synchronization or scenarios where the snapshot chain is broken. Incremental transfer only synchronizes the difference data, significantly reducing network bandwidth and storage resource consumption, and is especially suitable for large-scale migration of cold data. Incremental transmission is a data transfer method that only transmits the data differences that have changed since the last synchronization (or checkpoint), instead of transmitting the entire dataset each time. This method can significantly reduce network bandwidth consumption, storage space usage, and synchronization time, and is particularly suitable for scenarios with large amounts of data or frequent updates (such as virtual machine backups, disaster recovery synchronization, and file synchronization). The reflink snapshot chain ensures that data is transmitted without interruption or loss, meeting the RTO (Recovery Time Objective) requirements of disaster recovery systems.In this way, by dynamically selecting the transmission method through metadata-driven (lastRefink) and combining the lightweight characteristics of reflink snapshots, a balance between efficiency and reliability is achieved in cold data disaster recovery scenarios.

[0069] Step S14: During the data transmission process, adjust the transmission strategy of the data to be transmitted according to the health status of the virtual machine to complete the data transmission; the data to be transmitted includes the first data to be transmitted and the second data to be transmitted.

[0070] In one specific embodiment of this invention, since disaster recovery virtual machine data transmission is affected by multiple factors, such as host, storage, and network performance, this invention introduces the concept of comprehensive health. When determining the transmission strategy, this invention first acquires the target monitoring data of the virtual machine. The target monitoring data includes any one or a combination of several of the following: CPU utilization of the host machine corresponding to the virtual machine, storage processing read / write operation efficiency, disaster recovery backup system's disaster recovery storage write speed, and the virtual machine's network bandwidth utilization. A score is determined for each target monitoring data point according to pre-created scoring rules. Data weights are assigned to each target monitoring data point based on its score. A weighted sum is performed based on each score and data weight, and the overall health of the virtual machine is determined based on the weighted result. The overall health level of the virtual machine is determined based on the overall health level according to pre-created health level classification rules. Different overall health levels correspond to different health level registrations. If the virtual machine's health level is the first health level, the data to be transmitted is transmitted normally, and corresponding health reports are generated periodically. If the virtual machine's health level is the second health level, the synchronization frequency of non-critical virtual machines is reduced. Non-critical virtual machines are those with lower business continuity requirements than preset requirements, or those that allow for brief interruptions or data delays. If the virtual machine's health level is the third health level, the transmission of the second data to be transmitted is suspended, while the transmission of the first data to be transmitted and metadata continues.

[0071] Specifically, the process begins by obtaining the CPU utilization of the host machine hosting the disaster recovery virtual machine, the storage I / O capacity of the production storage pool, the storage write speed of the disaster recovery end, and the network bandwidth utilization used by the disaster recovery virtual machine. Specifically, host CPU, storage I / O, disaster recovery write speed, and network bandwidth are divided into multiple health intervals and assigned scores; indicator weights are dynamically allocated based on real-time business scenarios; a comprehensive health score is calculated through weighted summation, and a tiered response strategy is triggered. The first step, indicator segmentation and scoring, involves dividing each monitoring indicator into health intervals and assigning a score (0-100 points) based on the interval the actual value falls into, intuitively reflecting the status of each individual indicator. Figure 2 As shown, corresponding scoring rules can be preset for each indicator.

[0072] Step 2: Dynamic Weight Allocation: Based on the current business priority and historical fault records, dynamically adjust the weights of each indicator (the sum of the weights is 1), such as... Figure 3 As shown, different weights are assigned to different scenarios. It should be noted that during the process of assigning data weights to each target monitoring data based on various scores, if network congestion occurs, the data weight for network bandwidth utilization in the target monitoring data is increased; if a storage performance alarm is received, the data weight for storage processing read / write operation efficiency in the target monitoring data is increased. In a specific embodiment, the weight allocation strategy includes: when network congestion occurs, the network bandwidth weight is increased to over 50%; when a storage performance alarm is received, the storage I / O weight is increased to over 40%.

[0073] Step 3: Calculate the overall health score: The formula is: Overall Health Score = (C score × wC + I score × wI + W score × wW + B score × wB) / 100. Output: A value between 0 and 1 (1 being optimal). wC, wI, wW, and wB represent the weights of each indicator. Pre-defined health score grading and response strategies are available. Figure 4 As shown. In a specific embodiment, a virtual machine has a CPU (Central Processing Unit) utilization of 75% (score 80), a storage I / O latency of 8ms (score 70), a disaster recovery write speed of 80MB / s (score 60), and a network bandwidth utilization of 65% (score 70), and is currently in a normal transmission period (weight: CPU: 20%, I / O: 20%, storage: 30%, network: 30%). Calculation process: Overall health score = (80 × 0.2) + (70 × 0.2) + (60 × 0.3) + (70 × 0.3) / 100 = 0.69. Result: Health score = 0.69 (sub-healthy). In this case, according to... Figure 4 The system can be identified as being in a sub-optimal state. Based on the response actions, a reasonable solution is provided: adjust the disaster recovery virtual machine transmission strategy by reducing the synchronization frequency of non-critical virtual machines. Non-critical virtual machines are those that are not important; they do not need to be constantly synchronized. Hot and cold data are transmitted normally, but the synchronization frequency of less important virtual machines can be reduced, thus improving transmission performance. This completes the data transmission of virtual machines.

[0074] When reducing the synchronization frequency of non-critical virtual machines, the synchronization mechanism can be manually triggered: pause the default real-time synchronization and switch to manual full or incremental synchronization as needed. Alternatively, the synchronization timing can be controlled through scheduled tasks (such as Cron), for example, performing a full synchronization every morning at midnight. The synchronization mechanism can also be adjusted through parameter tuning: adjust the flushInterval parameter to extend the synchronization interval (e.g., from 1 second to 60 seconds) to achieve near real-time synchronization at the minute level; modify the minpoll and maxpoll parameters in the NTP configuration to reduce the time synchronization frequency. Other optimization methods include resource isolation by deploying non-critical virtual machines in dedicated clusters to avoid competing for bandwidth with critical services; and using caching techniques to reduce I / O operations.

[0075] Beneficial effects: This invention avoids redundant transmission of cold data by using a heat classification system. Different transmission methods are used for data with different heat levels. Hot data is transmitted through continuous data protection technology, while cold data is transmitted through different methods. This solves the problem of virtual machine performance degradation caused by frequent snapshot creation and lays a solid foundation for the stable operation of the cloud platform.

[0076] As described in the previous embodiment, this invention discloses a virtual machine data transmission method that can avoid redundant transmission of cold data through hot-data grading, thus solving the problem of virtual machine performance degradation caused by frequent snapshot creation. The specific virtual machine data transmission method will be described in detail below.

[0077] See Figure 5 As shown, this invention first creates a protection group, moves virtual machines into the protection group, and enables protection for the protection group. However, it does not immediately enable CDP for the virtual machines within the protection group during the task of enabling protection. In the asynchronous interface for calling the underlying start data transfer function, the underlying layer needs to determine whether CDP is enabled. For virtual machines in a shutdown state, a snapshot is taken, and libvirt starts CDP. For virtual machines in a boot state, if CDP is not enabled, the virtual machine is first paused, a reflink is created, CDP is enabled, and then the virtual machine is restored. A reflink configuration file is created under the storage pool: / datastore / xxxx / i-000xx / reflink.conf. If the protection group stops protecting, CDP for the virtual machines within the protection group is disabled, and ongoing data transfer is stopped. The underlying layer needs to provide an interface to clean up residual data.

[0078] Then, the disaster recovery virtual machine is categorized into hot and cold data: the access frequency of virtual machine disk blocks is analyzed to classify data into hot data (frequently modified) and cold data (infrequently modified). If the data in the disaster recovery virtual machine is hot data, the cache disk (hot data IO logs and bitmaps) is read sequentially from the CDP cache pool. If the data in the disaster recovery virtual machine is cold data, the reflink.conf file is checked for the existence of lastRefink. If it exists, the reflink difference data is directly transmitted twice; otherwise, the cold data is transmitted in full.

[0079] Finally, a comprehensive health score is set: Since data transmission for disaster recovery virtual machines is affected by multiple factors, including host, storage, and network performance, a comprehensive health score concept is introduced. This involves acquiring data such as the CPU utilization of the host to which the disaster recovery virtual machine belongs, the storage I / O capability of the production storage pool, the storage write speed of the disaster recovery end, and the network bandwidth utilization used by the disaster recovery virtual machine. Specifically: host CPU, storage I / O, disaster recovery write speed, and network bandwidth are divided into multiple health intervals and assigned scores; indicator weights are dynamically allocated based on real-time business scenarios; the comprehensive health score is calculated through weighted summation, triggering a tiered response strategy. The weight allocation strategy includes: when there is network congestion, the network bandwidth weight is increased to over 50%; when there is a storage performance alarm, the storage I / O weight is increased to over 40%. Step 1: Indicator Segmentation and Scoring: Each monitoring indicator is divided into health intervals, and a score (0-100 points) is assigned based on the interval the actual value falls into, intuitively reflecting the status of each item. Step 2: Dynamic Weight Allocation: Based on the current business priority and historical fault records, dynamically adjust the weights of each indicator (the sum of the weights is 1). Step 3: Comprehensive Health Calculation Formula: Comprehensive Health = (C score × wC + I score × wI + W score × wW + B score × wB) / 100. Output: A value between 0 and 1 (1 is optimal). Adjust the transmission strategy based on the comprehensive health score, and complete the data transmission according to the adjusted strategy.

[0080] Beneficial effects: By using a heat-based classification system to avoid redundant transmission of cold data, different transmission methods are used for data with different heat levels. Hot data is transmitted through continuous data protection technology, while cold data is transmitted through different methods. This solves the problem of virtual machine performance degradation caused by frequent snapshot creation, laying a solid foundation for the stable operation of the cloud platform.

[0081] See Figure 6 As shown, an embodiment of the present invention provides a virtual machine data transmission device, comprising:

[0082] The first judgment module 11 is used to determine whether the data to be transmitted is hot data after the virtual machine enables continuous data protection; hot data is data whose virtual machine disk block access frequency is greater than the target frequency threshold.

[0083] The data reading module 12 is used to read and transmit the first data to be transmitted in real time if it is hot data; the first data to be transmitted includes hot data input / output logs and bitmaps in the buffer pool;

[0084] The second judgment module 13 is used to determine whether there is a metadata tag in the reflink configuration file if it is not hot data, and to transmit the corresponding second data to be transmitted according to the corresponding judgment result; the metadata tag is a tag used to identify whether the virtual machine has a valid reflink snapshot chain;

[0085] The transmission strategy adjustment module 14 is used to adjust the transmission strategy of the data to be transmitted according to the health status of the virtual machine during the data transmission process in order to complete the data transmission; the data to be transmitted includes first data to be transmitted and second data to be transmitted.

[0086] Since the embodiments of the device part correspond to the embodiments described above, please refer to the embodiments described in the method part for the embodiments of the device part, and will not be repeated here.

[0087] Beneficial effects: This invention avoids redundant transmission of cold data by using a heat classification system. Different transmission methods are used for data with different heat levels. Hot data is transmitted through continuous data protection technology, while cold data is transmitted through different methods. This solves the problem of virtual machine performance degradation caused by frequent snapshot creation and lays a solid foundation for the stable operation of the cloud platform.

[0088] Furthermore, embodiments of this application also disclose an electronic device, Figure 7 This is a structural diagram of an electronic device according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. Specifically, the electronic device may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the virtual machine data transmission method disclosed in any of the foregoing embodiments. Furthermore, the electronic device in this embodiment may specifically be an electronic computer.

[0089] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0090] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0091] The operating system 221 is used to manage and control the various hardware devices on the electronic device and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the virtual machine data transfer method executed by the electronic device as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0092] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed virtual machine data transmission method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0093] Furthermore, this application also discloses a computer program product, including a computer program / instructions; wherein, when the computer program / instructions are executed by a processor, they implement the aforementioned disclosed virtual machine data transfer method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0095] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0096] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0097] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0098] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A virtual machine data transmission method, characterized by, The application comprises the following steps: After starting the continuous data protection of the virtual machine, it is judged whether the data to be transmitted is hot data; The hot data is data whose virtual machine disk block access frequency is greater than a target frequency threshold; If the data to be transmitted is hot data, the first data to be transmitted is read and transmitted in real time; the first data to be transmitted comprises hot data input / output logs and bitmaps in a cache pool; If the data to be transmitted is not hot data, it is judged whether there is a metadata mark in a reflink configuration file, and corresponding second data to be transmitted is transmitted according to the corresponding judgment result; the metadata mark is a mark used to identify whether there is a valid reflink snapshot chain of the virtual machine; During the data transmission process, the transmission strategy of the data to be transmitted is adjusted according to the health degree of the virtual machine, so as to complete the data transmission; the data to be transmitted comprises the first data to be transmitted and the second data to be transmitted.

2. The virtual machine data transmission method of claim 1, wherein, Before starting the continuous data protection of the virtual machine, the following steps are further included: The virtual machine is added to a pre-created protection group, so as to start the protection function of the protection group and call an asynchronous interface of underlying data transmission to perform asynchronous data processing; When the asynchronous data processing starts, it is judged through an underlying interface whether the continuous data protection of the virtual machine is started; If the continuous data protection is not started, corresponding operations are performed according to the state of the virtual machine, so as to start the continuous data protection.

3. The virtual machine data transmission method of claim 2, wherein, The corresponding operations performed according to the state of the virtual machine, so as to start the continuous data protection, comprise the following steps: If the current state of the virtual machine is a shutdown state, a snapshot is directly generated, and the continuous data protection of the virtual machine is started through a virtualization management library; If the current state of the virtual machine is a startup state, the virtual machine is suspended, a reflink snapshot is created, the continuous data protection of the virtual machine is started based on the reflink snapshot, and the virtual machine is restarted.

4. The virtual machine data transmission method of claim 1, wherein, Further comprising the following steps: The reflink configuration file of the virtual machine is created; If the pre-created protection group stops protection, the continuous data protection of the virtual machine is closed, data transmission is stopped, and snapshot chain residual data is deleted through an application programming interface.

5. The virtual machine data transmission method of claim 1, wherein, The judgment of whether there is a metadata mark in the reflink configuration file and the transmission of corresponding second data to be transmitted according to the corresponding judgment result comprise the following steps: If there is a metadata mark in the reflink configuration file, difference data between two snapshots is obtained and transmitted through an underlying interface, and the difference data is merged, so as to complete the transmission of the second data to be transmitted; If there is not, the second data to be transmitted is directly transmitted in full amount; the second data to be transmitted is cold data.

6. The virtual machine data transmission method according to any one of claims 1 to 5, wherein, The adjustment of the transmission strategy of the data to be transmitted according to the health degree of the virtual machine, so as to complete the data transmission, comprises the following steps: Target monitoring data of the virtual machine are obtained; the target monitoring data comprise any one or a combination of several of the following: central processing unit usage rate of a host corresponding to the virtual machine, storage processing read / write operation efficiency, disaster recovery end storage writing speed of a disaster recovery backup system, and network bandwidth usage rate of the virtual machine; The scores of the target monitoring data are determined according to pre-created scoring rules; According to each score, a data weight of each target monitoring data is allocated; Based on the weighted summation of each score and data weight, a comprehensive health degree of the virtual machine is determined according to the corresponding weighted result; According to a pre-created health degree grading rule, a health degree of the virtual machine is determined based on the comprehensive health degree; wherein different comprehensive health degrees correspond to different health degree records; If the health degree of the virtual machine is a first health degree level, the to-be-transmitted data is normally transmitted, and a corresponding health report is periodically generated; If the health degree of the virtual machine is a second health degree level, the synchronization frequency of a non-critical virtual machine is reduced; the non-critical virtual machine is a virtual machine with a lower requirement on service continuity than a preset requirement, or a virtual machine allowing temporary interruption or data delay; If the health degree of the virtual machine is a third health degree level, the transmission of the second to-be-transmitted data is suspended, and the transmission of the first to-be-transmitted data and metadata is continued.

7. The virtual machine data transmission method according to claim 6, wherein, In the process of allocating a data weight of each target monitoring data according to each score, the process further includes: If the network is congested, the data weight of the network bandwidth usage in the target monitoring data is increased; If a storage performance alarm is received, the data weight of the storage processing read-write operation efficiency in the target monitoring data is increased.

8. A virtual machine data transfer apparatus, comprising: The method comprises: A first judgment module is configured to judge whether the to-be-transmitted data is hot data after the virtual machine starts the continuous data protection; The hot data is data with a virtual machine disk block access frequency greater than a target frequency threshold; A data reading module is configured to read and transmit the first to-be-transmitted data in real time if the to-be-transmitted data is hot data; the first to-be-transmitted data includes hot data input / output logs and bitmaps in a cache pool; A second judgment module is configured to judge whether there is a metadata mark in a reflink configuration file if the to-be-transmitted data is not hot data, and transmit corresponding second to-be-transmitted data according to the corresponding judgment result; the metadata mark is a mark used to identify whether there is a valid reflink snapshot chain of the virtual machine; A transmission strategy adjustment module is configured to adjust the transmission strategy of the to-be-transmitted data according to the health degree of the virtual machine during data transmission, so as to complete data transmission; the to-be-transmitted data includes the first to-be-transmitted data and the second to-be-transmitted data.

9. An electronic device, comprising: The method comprises: A memory is configured to store a computer program; A processor is configured to execute the computer program to implement the steps of the virtual machine data transmission method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement the steps of the virtual machine data transmission method according to any one of claims 1 to 7.