Multi-channel parallel virtual machine online migration method and system based on memory layering

By layering virtual machine memory into static data and dynamic data, using shared storage direct transmission and multi-channel parallel transmission, an adaptive hybrid migration mechanism and network fault tolerance recovery process are designed, which solves the problems of low migration efficiency and reliability under large memory and high write dirty rates in AI scenarios, and achieves efficient and reliable virtual machine migration.

CN120353542AActive Publication Date: 2025-07-22HENAN AIRPORT GRP CO LTD

Patent Information

Application Number
CN202510483615.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing virtual machine migration technology has low migration efficiency, poor iteration convergence and lack of network fault tolerance in AI scenarios under large memory and high write dirty rates.

Method used

The virtual machine memory is layered into static data and dynamic data, and direct shared storage transmission and multi-channel parallel transmission is adopted. The adaptive hybrid migration mechanism based on iterative convergence factor is designed, and a network fault tolerance recovery process is introduced.

Benefits of technology

It realizes efficient migration of large memory virtual machines, ensures iterative convergence of high-write dirty scenarios and migration reliability in complex network environments, and provides low-latency and high-availability online migration solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353542A_ABST
    Figure CN120353542A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-channel parallel virtual machine online migration method and system based on memory layering, and belongs to the technical field of cloud computing. According to the method, a to-be-migrated virtual machine memory is divided into a static data layer and a dynamic data layer, and the static data layer is directly transmitted to a target end through shared storage; the dynamic data layer is divided into a plurality of areas according to memory addresses, independent transmission channels are distributed for parallel migration, and the sizes of the partitions are dynamically adjusted according to the dirty page rate; calculating an iteration convergence factor in real time, and switching to a post-copy mode and backing up dirty pages to a shared storage when continuous iteration is not converged for multiple times; and recovering data from the shared storage when the network fails, and finally finishing the migration of the remaining dirty pages. Through memory hierarchical optimization, parallel transmission acceleration and a network fault-tolerant mechanism, the problems of low efficiency, iteration divergence and service interruption risks of large-memory and high-write-dirty-rate virtual machine migration in an AI scene are solved, the migration speed and success rate are remarkably improved, and continuous execution of an AI task is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and particularly relates to a multi-channel parallel virtual machine online migration method and system based on memory layering. Background Art

[0002] The virtual machine online migration technology is the core means to realize the dynamic allocation of cloud computing resources, but it faces severe challenges in the artificial intelligence scenario. AI applications (such as large-scale deep learning training, large language model inference) need to process a large amount of memory data, and a large amount of memory dirty writes are generated due to high-frequency parameter updates, resulting in various defects and deficiencies in traditional migration technologies such as the pre-copy mechanism, the post-copy mechanism, the hybrid copy mechanism, memory compression, and hardware acceleration solutions. Among them:

[0003] (1) The pre-copy mechanism relies on multiple rounds of iterative transmission of dirty pages. During AI training, the number of iterations surges due to continuous parameter updates, and the migration time of virtual machine memory data can reach several hours. At the same time, when the dirty page generation rate exceeds the transmission bandwidth, the iteration cannot converge, resulting in migration failure.

[0004] (2) After the target end starts in the post-copy mechanism, high-frequency page fault interrupts are triggered, resulting in a significant reduction in the speed of AI inference tasks and a sharp increase in latency. In addition, a network failure interruption during migration will cause data inconsistency between the source / end and the target end, leading to service collapse.

[0005] (3) The fixed threshold switching strategy of the hybrid copy mechanism cannot adapt to sudden changes in AI loads. For example, when the training batch is switched, the dirty page rate suddenly rises. After switching to the post-copy mechanism, there is still a lack of network fault tolerance mechanism, and there is still a problem of service collapse.

[0006] (4) In addition, the memory compression solution will cause an increase in CPU utilization, exacerbating the contention for AI computing resources. Hardware acceleration solutions (such as Intel VT-d, RDMA) rely on specific devices, with insufficient adaptation rate in heterogeneous cloud environments, and automatic frequency reduction leads to a decrease in the execution efficiency of AI tasks.

[0007] In summary, the existing online migration mechanisms and solutions for memory data cannot effectively balance the migration efficiency and reliability of large memory and high dirty write rate in the AI scenario. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to overcome the defects existing in the existing virtual machine migration technology, such as low migration efficiency, poor iterative convergence, and lack of network fault tolerance in the scenarios of large memory and high write dirty rate. A multi-channel parallel virtual machine online migration method and system based on memory layering are provided. By stratifying the virtual machine memory into static data and dynamic data, direct transfer through shared storage and multi-channel parallel transfer are respectively adopted; an adaptive hybrid migration mechanism based on an iterative convergence factor is designed to dynamically switch the migration mode; and a network fault tolerance recovery process is introduced. The above technologies cooperate to achieve efficient migration of large-memory virtual machines, guarantee iterative convergence in high write dirty scenarios, and improve migration reliability in complex network environments, providing a low-latency and highly available online migration solution for AI training / inference scenarios.

[0009] The multi-channel parallel virtual machine online migration method based on memory layering includes the following steps:

[0010] S1 - Based on the pre-copy mechanism, divide the memory data of the virtual machine to be migrated into a static data layer and a dynamic data layer. Among them, the static data layer is the data in the full-memory copy stage of pre-copy, and the dynamic data layer is the dirty page data generated in the iterative copy stage;

[0011] S2 - Save the static data layer in the form of a snapshot file to the shared storage and transmit it to the target end through the storage network;

[0012] S3 - Divide the dynamic data layer into multiple independent regions according to the memory address space, allocate an independent transmission channel for each region, and dynamically adjust the partition size according to the dirty page rate; transmit the dynamic data through multi-channel parallel, and the target end merges the dirty pages according to the memory address version number and verifies the integrity;

[0013] S4 - Real-time monitor the dirty page generation rate and transmission bandwidth of the dynamic data layer, calculate the iterative convergence factor, and when the continuous mean value of the convergence factor ≥ 1, determine that the iteration does not converge;

[0014] S5 - If it is determined that the iteration does not converge multiple times, trigger the adaptive hybrid migration mechanism, switch to the post-copy mode, and the target end pulls the missing pages through a page fault interrupt request;

[0015] S6 - In the post-copy mode, back up the dirty page data of the last iteration to the shared storage; if a network fault is detected, the target end loads the backup data from the shared storage to complete the migration;

[0016] S7 - When the remaining dirty page transmission time is less than the preset downtime threshold, pause the source virtual machine, transmit the remaining data to the target end and resume operation to complete the migration.

[0017] Further, the saving and transmission of the static data layer in step S2 specifically include,

[0018] - Mark the memory pages of the static data layer using a bitmap (SBitmap);

[0019] - Transmit the snapshot file to the shared storage through the storage network;

[0020] - The target end directly loads the static data layer from the shared storage and merges it with the dynamic data layer.

[0021] Furthermore, the partitioning rules of the dynamic data layer described in step S3 specifically include:

[0022] - Divide it into n regions of equal size according to the memory address space, and each region corresponds to an independent transmission channel;

[0023] - Dynamically adjust the partition size according to the dirty page generation rate, and split the high dirty page rate region into smaller blocks.

[0024] Furthermore, the target end maintains an independent receive buffer for each transmission channel, and only retains the latest dirty page status after sorting by the memory address version number.

[0025] Furthermore, the calculation method of the iterative convergence factor described in step S4 is:

[0026] - λ(i) = R(i) / B(i), where λ(i) is the iterative convergence factor, R(i) is the dirty page generation rate in the i-th round, and B(i) is the transmission bandwidth;

[0027] - When λ(i) ≥ 1 for 5 consecutive times and the average value ≥ 1, it is determined that the iteration does not converge.

[0028] Furthermore, the network fault tolerance mechanism described in step S5 specifically includes:

[0029] - When switching to post-copy, save the dirty page data of the last iteration as a shared storage backup;

[0030] - In case of a network fault, the target end loads the missing pages based on the backup data and verifies the data integrity.

[0031] Furthermore, when calculating the iterative convergence factor in step S5, synchronously monitor the network delay and jitter, and trigger the mode switch in advance when the delay exceeds the threshold.

[0032] Specifically, the method is applicable to the migration of virtual machines with large memory capacity and high write-dirty rate in the AI training scenario; it is applicable to the online resource scheduling of heterogeneous hardware (CPU / GPU / TPU) nodes in the hybrid cloud environment.

[0033] The multi-channel parallel virtual machine online migration system based on memory hierarchy is used to implement the above-mentioned multi-channel parallel virtual machine online migration method based on memory hierarchy; the system includes a source host, a target host, and a shared storage, where,

[0034] The source host includes a memory hierarchy module, a memory hierarchy and multi-channel management module, and an adaptive hybrid migration control module, which are used to perform memory hierarchy, dynamic partitioning, and adaptive migration control. The memory hierarchy module marks the full memory snapshot through a static data bitmap, and the memory hierarchy and multi-channel management module marks the iterative dirty pages through a dynamic data bitmap;

[0035] The target host includes a static data processing module, a dynamic data processing module, and a hybrid migration switching module that correspond one-to-one with the memory hierarchy module, the memory hierarchy and multi-channel management module, and the adaptive hybrid migration control module of the source host, which are used to process static data loading, dynamic data merging, and migration mode switching. The hybrid migration switching module of the target host module responds to a page fault and pulls the missing page from the source host or the shared storage;

[0036] The shared storage is used to provide low-latency static data storage services and support concurrent access and backup recovery.

[0037] A multi-channel parallel virtual machine online migration method and system based on memory hierarchy of the present invention overcomes the defects existing in the existing virtual machine migration technology, such as low migration efficiency, poor iterative convergence, and lack of network fault tolerance in the scenarios of large memory and high write dirty rate. Through the innovative architecture of direct transfer of static data through shared storage, multi-channel concurrent transmission of dynamic data, dynamic switching of iterative convergence factors for migration mode, and backup recovery of network failures, the following beneficial effects are achieved:

[0038] (1) Significantly improved migration efficiency: Through the hierarchical transmission and multi-channel parallel mechanism, the bandwidth limitation of traditional single-channel transmission is broken through, and the migration time of large memory is greatly shortened;

[0039] (2) Breakthrough guarantee of migration success rate: Based on the adaptive switching strategy of dynamic convergence criterion, the iterative failure caused by the dirty page storm is effectively avoided, ensuring the success of migration in high write dirty scenarios;

[0040] (3) Enhanced business continuity: Through the quick recovery of shared storage backup during network failures, the risk of data inconsistency is eliminated, ensuring the seamless migration of AI tasks;

[0041] (4) Resource collaborative optimization: The heterogeneous network division of labor and dynamic partitioning strategy reduce resource contention and adapt to the complex hardware topology of the cross-cloud environment. Description of the Drawings

[0042] The following further describes a method and system for online migration of a multi-channel parallel virtual machine based on memory layering according to the present invention in conjunction with the accompanying drawings:

[0043] Figure 1 It is a structural principle block diagram of the method and system for online migration of a multi-channel parallel virtual machine based on memory layering;

[0044] Figure 2 It is an overall process line block diagram of the method for online migration of a multi-channel parallel virtual machine based on memory layering;

[0045] Figure 3 It is a system architecture diagram of the system for online migration of a multi-channel parallel virtual machine based on memory layering; Specific Embodiments

[0046] In the present invention, unless otherwise clearly specified and defined, terms such as "installation", "connection", "connection", "fixation" and the like shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be directly connected, or indirectly connected through an intermediate medium, and may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0047] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "left", "right", "front", "rear", "top", "bottom", "inside", "outside" and the like are all based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0048] The following further describes the technical solution of the present invention with specific embodiments, but the protection scope of the present invention is not limited to the following embodiments.

[0049] Embodiment 1: As Figure 1 、 2 shown, the method for online migration of a multi-channel parallel virtual machine based on memory layering includes the following steps:

[0050] S1 - Based on the pre - copy mechanism, the memory data of the virtual machine to be migrated is divided into a static data layer and a dynamic data layer. Among them, the static data layer is the data in the full - volume memory copy stage of pre - copy, and the dynamic data layer is the dirty page data generated in the iterative copy stage. Specifically: Add a static data bitmap (StaticDatas Bitmap, SBitmap) data structure in qemu to mark the static data StaticDatas migrated in the first layer. Similarly, mark the dynamic data DynamicDatas to be migrated in the second layer through a dynamic data bitmap (DynamicDatasBitmap, DBitmap) data structure.

[0051] S2 - Save the static data layer in the form of a snapshot file to shared storage and transmit it to the target end through the storage network; S3 - Divide the dynamic data layer into multiple independent regions according to the memory address space. Each region is assigned an independent transmission channel, and the partition size is dynamically adjusted according to the dirty page rate; Transmit the dynamic data through multi - channel parallel transmission. The target end merges the dirty pages according to the memory address version number and verifies the integrity. Specifically: S2 and S3 transmit StaticDatas and DynamicDatas in parallel simultaneously.

[0052] S4 - Real - time monitor the dirty page generation rate and transmission bandwidth of the dynamic data layer, calculate the iterative convergence factor. When the continuous mean value of the convergence factor ≥ 1, it is determined that the iteration does not converge;

[0053] S5 - If it is determined that the iteration does not converge multiple times, trigger the adaptive hybrid migration mechanism and switch to the post - copy mode. The target end pulls the missing pages through a page - fault interrupt request. Specifically: After the target node merges the StaticDatas, it notifies the source node. The source node determines whether to switch to Post - copy according to the value of λ(i) at the beginning of each subsequent iteration. The specific strategy is: Observe the values of λ(i) in the recent 5 times. When all λ(i) are less than 1, it indicates that the migration iteration is in a convergent state, and there is no need to switch at this time. When all λ(i) values are greater than 1 at least 2 times, then calculate the average value λ of λ(i) in the recent 5 times aver ,If λ aver ≥ 1, it indicates that the iteration no longer converges, and at this time, switch to the post - copy mechanism. Based on the average value of the recent convergence factors as the reference value for the switching moment. On the one hand, because the average value of the convergence factor can more accurately represent the convergence trend, and on the other hand, to prevent misjudgment caused by occasional fluctuations in network or application load;

[0054] S6 - In the post - copy mode, back up the dirty page data of the last iteration to the shared storage; if a network failure is detected, the target end loads the backup data from the shared storage to complete the migration. Specifically: If step 5 triggers a switch and the migration method switches to post - copy, determine whether there is a network failure during the Post - copy migration. If there is a failure, obtain the dirty pages generated during the last round of iteration based on the DBitmap and save them in the form of a file to the shared storage, that is Figure 2 the data backup during the post - copy phase in

[0055] When the remaining dirty page transfer time is less than the preset downtime threshold, pause the source virtual machine, transfer the remaining data to the target end and resume operation to complete the migration.

[0056] Embodiment 2: As Figure 1 shown,

[0057] The saving and transmission of the static data layer described in step S2 specifically include

[0058] - Use a bitmap (SBitmap) to mark the memory pages of the static data layer;

[0059] - Transmit the snapshot file to the shared storage through the storage network;

[0060] - The target end directly loads the static data layer from the shared storage and merges it with the dynamic data layer.

[0061] The partitioning rules of the dynamic data layer described in step S3 specifically include

[0062] - Divide it into n regions of equal size according to the memory address space, and each region corresponds to an independent transmission channel;

[0063] - Dynamically adjust the partition size according to the dirty page generation rate, and split the high - dirty - page - rate regions into smaller blocks.

[0064] The target end maintains an independent receive buffer for each transmission channel, and only retains the latest dirty page status after sorting by the memory address version number.

[0065] Specifically: StaticDatas is saved in the shared storage in the form of a file based on SBitmap, and DynamicDatas is iteratively transmitted based on DBitmap. Only the memory pages marked as "dirty" in the previous round are transmitted in each round of iteration. DBitmap is divided into 1 - n regions, and a transmission process is created for each region, corresponding to 1 - n data transmission channels. When the source node finishes saving StaticDatas, the destination node loads StaticDatas from the shared storage and merges it with DynamicDatas. The merging principle is to use DynamicDatas as the latest data for merging, that is, if the corresponding DynamicDatas already exists in the destination virtual machine, no further merging is performed. The remaining steps are as described in Embodiment 1 and will not be repeated here.

[0066] Embodiment 3: As Figure 1 shown, the calculation method of the iterative convergence factor described in step S4 is

[0067] - λ(i) = R(i) / B(i), where λ(i) is the iterative convergence factor, R(i) is the dirty page generation rate in the i - th round, and B(i) is the transmission bandwidth;

[0068] - When λ(i) ≥ 1 for 5 consecutive times and the average value ≥ 1, it is determined that the iteration does not converge.

[0069] Specifically: During the process of transmitting DynamicDatas, it is automatically determined whether to switch the migration mode. It senses whether the iteration converges based on the convergence factor λ(i). λ(i) is related to the memory dirty page generation rate R(i) and the transmission bandwidth B(i). Assume that the transmission time in the i - th round during the iterative migration process is T(i), and the number of memory dirty pages generated is pages(i). Then R(i), B(i), and λ(i) can be expressed as the following formula:

[0070]

[0071] Furthermore, the network fault tolerance mechanism described in step S5 specifically includes

[0072] - When switching during post - copy, save the dirty page data of the last iteration as a shared storage backup;

[0073] - In case of a network fault, the destination end loads the missing pages based on the backup data and verifies the data integrity. The remaining steps are as described in Embodiment 1 and will not be repeated here.

[0074] Embodiment 4: As Figure 1 、 3As shown, when calculating the iterative convergence factor described in step S5, the network delay and jitter are synchronously monitored, and when the delay exceeds the threshold, the mode switch is triggered in advance. The remaining steps are the same as those described in Embodiment 1 and will not be repeated.

[0075] In summary, this method is applicable to the virtual machine migration with large memory capacity and high write-dirty rate in the AI training scenario; it is applicable to the online resource scheduling of heterogeneous hardware (CPU / GPU / TPU) nodes in the hybrid cloud environment.

[0076] Embodiment: As Figure 3 shown, the multi-channel parallel virtual machine online migration system based on memory stratification is used to implement the above-mentioned multi-channel parallel virtual machine online migration method based on memory stratification; the system includes a source host, a target host, and a shared storage, where

[0077] the source host includes a memory stratification module, a memory stratification and multi-channel management module, and an adaptive hybrid migration control module, which are used to perform memory stratification, dynamic partitioning, and adaptive migration control;

[0078] the memory stratification module is responsible for dividing the virtual machine memory into static data (the first layer) and dynamic data (the second layer). The static data is the full amount of memory data in the first round of Pre-copy and is saved to the shared storage in the form of a snapshot file. The dynamic data is the dirty page data transmitted iteratively during the migration process. The memory stratification and multi-channel management module is responsible for dividing the second-layer memory into multiple regions 1-n by address, and each region is assigned an independent transmission channel (channels 1-n), supporting parallel transmission of RDMA or TCP. At the same time, the region size and channel allocation are dynamically adjusted (for example, the region with a high dirty page rate is split into smaller blocks). The adaptive hybrid migration control module is responsible for monitoring the dirty page rate, iteration times, and network delay of the dynamic data, judging the iterative convergence situation based on this, determining whether to trigger the migration mode according to the iterative convergence situation, and sending a switching instruction to the target host at the same time;

[0079] the target host includes a static data processing module, a dynamic data processing module, and a hybrid migration switching module corresponding to the memory stratification module, the memory stratification and multi-channel management module, and the adaptive hybrid migration control module of the source host respectively, which are used to process static data loading, dynamic data merging, and migration mode switching. The hybrid migration switching module of the target host module responds to the page fault interrupt and pulls the missing page from the source host or the shared storage;

[0080] The static data processing module is responsible for loading the snapshot file of the second layer from the shared storage and restoring the basic memory state of the virtual machine; the dynamic data processing module is responsible for receiving multi-channel dirty page data, merging data blocks from different channels according to the memory address, merging data based on the version number, and verifying the integrity of the data; the hybrid migration switching module is responsible for receiving the switching instruction feedback from the source host, and in the Post-copy mode, responding to the page fault interrupt of the virtual machine and pulling the missing page from the source host.

[0081] The shared storage is used to provide low-latency static data storage services, support concurrent access and backup recovery. Specifically, the shared storage provides low-latency and high-throughput first-layer static data storage services, supports concurrent access (writing by the source host and reading by the target host); the shared module can choose NFS (Network File System), Ceph (distributed storage) or S3 (object storage), and can be combined with RDMA acceleration (such as NVMe-oF) to improve the data transmission speed.

[0082] The multi-channel parallel virtual machine online migration method and system based on memory layering overcome the defects existing in the existing virtual machine migration technology, such as low migration efficiency, poor iterative convergence, and lack of network fault tolerance in the scenarios of large memory and high write dirty rate. It provides a multi-channel parallel virtual machine online migration method based on memory layering, which divides the virtual machine memory into static data and dynamic data, and respectively uses direct transmission of shared storage and multi-channel parallel transmission; designs an adaptive hybrid migration mechanism based on the iterative convergence factor to dynamically switch the migration mode; and introduces a network fault tolerance and recovery process. The above technologies cooperate to achieve the efficient migration of large-memory virtual machines, the guarantee of iterative convergence in high-write-dirty scenarios, and the improvement of migration reliability in complex network environments, providing a low-latency and highly available online migration solution for AI training / inference scenarios.

[0083] The above description shows the main features, basic principles, and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments or examples, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, the above embodiments or examples should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0084] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A multi-channel parallel virtual machine online migration method based on memory layering, characterized in that: It includes the following steps: S1 - Based on the pre-copy mechanism, divide the memory data of the virtual machine to be migrated into a static data layer and a dynamic data layer. Among them, the static data layer is the data in the full-memory copy stage of pre-copy, and the dynamic data layer is the dirty page data generated in the iterative copy stage; S2 - Save the static data layer in the form of a snapshot file to the shared storage and transmit it to the target end through the storage network; S3 - Divide the dynamic data layer into multiple independent regions according to the memory address space. Each region is assigned an independent transmission channel, and the partition size is dynamically adjusted according to the dirty page rate; transmit the dynamic data through multi-channel parallel transmission, and the target end merges the dirty pages according to the memory address version number and verifies the integrity; S4 - Real-time monitor the dirty page generation rate and transmission bandwidth of the dynamic data layer, calculate the iterative convergence factor. When the continuous mean value of the convergence factor ≥ 1, it is determined that the iteration does not converge; S5 - If it is determined that the iteration does not converge multiple times, trigger the adaptive hybrid migration mechanism, switch to the post-copy mode, and the target end pulls the missing pages through a page fault interrupt request; S6 - In the post-copy mode, back up the dirty page data of the last iteration to the shared storage; if a network failure is detected, the target end loads the backup data from the shared storage to complete the migration; S7 - When the remaining dirty page transmission time is less than the preset downtime threshold, pause the source virtual machine, transmit the remaining data to the target end and resume operation to complete the migration.

2. The method for online migration of a multi-channel parallel virtual machine based on memory layering according to claim 1, characterized in that: The saving and transmission of the static data layer described in step S2 specifically include - Use a bitmap to mark the memory pages of the static data layer; - Transmit the snapshot file to the shared storage through the storage network; - The target end directly loads the static data layer from the shared storage and merges it with the dynamic data layer.

3. The multi-channel parallel virtual machine online migration method based on memory layering according to claim 1, wherein: The partition rule of the dynamic data layer described in step S3 specifically includes - Divide it into n regions of equal size according to the memory address space, and each region corresponds to an independent transmission channel; - Dynamically adjust the partition size according to the dirty page generation rate, and split the high dirty page rate region into smaller blocks.

4. The multi-channel parallel virtual machine online migration method based on memory layering according to claim 1, wherein: In step S3, the target end maintains an independent receive buffer for each transmission channel and only retains the latest dirty page status after sorting according to the memory address version number.

5. The method for online migration of a multi-channel parallel virtual machine based on memory layering according to claim 1, characterized in that: The calculation method of the iterative convergence factor described in step S4 is - λ(i) = R(i) / B(i), where λ(i) is the iterative convergence factor, R(i) is the dirty page generation rate in the i-th round, and B(i) is the transmission bandwidth; - When λ(i) ≥ 1 for 5 consecutive times and the mean value ≥ 1, it is determined that the iteration does not converge.

6. The multi-channel parallel virtual machine online migration method based on memory layering according to claim 1, characterized in that: The network fault tolerance mechanism described in step S5 specifically includes - When switching to post-copy, save the dirty page data of the last iteration as a shared storage backup; - In case of a network failure, the target end loads the missing pages based on the backup data and verifies the data integrity.

7. The multi-channel parallel virtual machine online migration method based on memory layering according to claim 1, wherein: When calculating the iterative convergence factor in step S5, synchronously monitor the network delay and jitter, and trigger the mode switch in advance when the delay exceeds the threshold.

8. The method for online migration of multi-channel parallel virtual machines based on memory layering according to claim 1, characterized in that: The method is applicable to the migration of virtual machines with large memory capacity and high write-dirty rate in the AI training scenario; It is applicable to the online resource scheduling of heterogeneous hardware nodes in the hybrid cloud environment.

9. A multi-channel parallel virtual machine online migration system based on memory layering, characterized in that: The multi-channel parallel virtual machine online migration system based on memory layering is used to implement the multi-channel parallel virtual machine online migration method based on memory layering according to any one of claims 1-8; the system includes a source host, a target host, and a shared storage, wherein, The source host includes a memory layering module, a memory layering and multi-channel management module, and an adaptive hybrid migration control module, which are used to perform memory layering, dynamic partitioning, and adaptive migration control. The memory layering module marks the full-memory snapshot through a static data bitmap, and the memory layering and multi-channel management module marks the iterative dirty pages through a dynamic data bitmap; The target host includes a static data processing module, a dynamic data processing module, and a hybrid migration switching module corresponding one-to-one to the memory layering module, the memory layering and multi-channel management module, and the adaptive hybrid migration control module of the source host, which are used to process static data loading, dynamic data merging, and migration mode switching. The hybrid migration switching module of the target host module responds to a page fault and pulls the missing page from the source host or the shared storage; The shared storage is used to provide low-latency static data storage services and support concurrent access and backup recovery.

Citation Information

Patent Citations

  • Virtual machine migration method and device

    CN112148421A

  • Hardware marking implementation method for dirty pages of virtual machine of intelligent network card

    CN115586943A

  • Technologies for virtual machine migration

    US20180024854A1

Cited By

  • Virtual machine recovery method and device based on shared memory, medium and program product

    CN120743439A

  • Data migration method and device, virtual machine to be migrated, medium and product

    CN120849025A

  • Optimization method for high-speed read-write of memory

    CN122131986A