Virtual machine migration method and apparatus, offloading card, device, and system

Through uninstalling card and tunneling technology, packet loss and delay problems in virtual machine migration are solved, and efficient virtual machine migration is achieved.

WO2025180084A1PCT designated stage Publication Date: 2025-09-04HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2025/070280
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-01-02
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

During the virtual machine migration process, there are problems such as packet loss and long migration delays, especially in large bandwidth scenarios. The existing technology such as reliability mechanism retransmission of packets or VMM cache packets will lead to network congestion and high storage space requirements.

Method used

Using uninstall card and tunneling technology, the uninstall card serves as a relay device during the virtual machine migration process, and processes and forwards packets in real time to avoid packet loss, and reduces the use of VMM storage space and processor resources, reducing latency.

Benefits of technology

Effectively shortens the delay in virtual machine migration, avoid packet loss, improve resource utilization, reduce dependence on processors, and supports high-performance network protocols such as RoCE.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070280_04092025_PF_FP_ABST
    Figure CN2025070280_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a virtual machine migration method and apparatus, an offloading card, a device, and a system, related to the field of computers. The method comprises: when a virtual machine is migrated from a source device to a destination device, the source device receiving a service packet sent by an external device to the virtual machine in the source device, and an offloading card in the source device transmitting service data comprised by the service packet to an offloading card in the destination device on the basis of a tunnel between the source device and the destination device; the offloading card of the destination device feeding back a response packet by means of the tunnel between the source device and the destination device, or feeding back a response message to the external device on the basis of the tunnel between the destination device and the external device. The packet of the virtual machine is processed and forwarded in real time, preventing a packet loss phenomenon of virtual machine migration, reducing occupation of the storage space of the VMM or the storage space of the source device, reducing participation of a processor in the source device in packet processing, decreasing packet processing time delay, enabling the destination device to restart the virtual machine in a timely manner, and effectively shortening time delay of virtual machine migration.
Need to check novelty before this filing date? Find Prior Art

Description

Virtual machine migration method, device, uninstall card, equipment and system

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 28, 2024, with application number 202410224253.0 and application name “Virtual Machine Migration Method, Device, Uninstall Card, Equipment and System”, all of the contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computers, and in particular to a virtual machine migration method, apparatus, uninstall card, device, and system. Background Art

[0003] Currently, during virtual machine migration, packet loss occurs due to the temporary shutdown of the virtual machine and the suspension of message sending and receiving. In high-bandwidth scenarios, packet loss can be even more severe. Typically, a reliability mechanism is used to retransmit messages, but not all protocols support reliable transmission mechanisms, and retransmitted messages lead to network congestion and longer virtual machine migration delays. Alternatively, a virtual machine monitor (VMM) is used to cache messages, which places a high demand on the virtual machine monitor's storage space, and reading messages from the cache increases virtual machine migration delays. Summary of the Invention

[0004] The present application provides a virtual machine migration method, apparatus, offload card, device and system, which can effectively shorten the delay of virtual machine migration and solve the problem of packet loss in virtual machine migration.

[0005] In a first aspect, a virtual machine migration method is provided, the method comprising: when a virtual machine migrates from a first computer device to a second computer device, the first computer device receives a first service message sent by a third computer device to the virtual machine in the first computer device, and an offload card in the first computer device transmits first service data included in the first service message to the offload card in the second computer device based on a first tunnel between the offload card in the first computer device and the offload card in the second computer device. The third computer device includes a computer device other than the first computer device and the second computer device.

[0006] Compared with the problem of retransmitting messages by using a reliability mechanism or using VMM to cache messages in order to solve the problem of long virtual machine migration time in order to solve the packet loss problem in virtual machine migration, the virtual machine migration method provided by the present application provides a message lossless relay mechanism in the process of virtual machine migration based on the unloading capability of the unloading card and the tunnel technology, that is, after the virtual machine migrates from the first computer device to the second computer device, the unloading card in the first computer device processes the message sent by the third computer device to the virtual machine in the first computer device, and forwards the message to the second computer device through the tunnel. Therefore, during the virtual machine migration process, the message of the virtual machine can be processed and forwarded in real time, avoiding the packet loss phenomenon in virtual machine migration, and eliminating the need for VMM to cache messages or retransmit messages by using a reliability mechanism, thereby reducing the storage space occupied by the VMM or the storage space of the first computer device, reducing the processor in the first computer device from participating in message processing, reducing the message processing delay, enabling the second computer device to restart the virtual machine in time, and effectively shortening the delay of virtual machine migration.

[0007] In one possible implementation, an offload card in a first computer device transmits first service data to an offload card in a second computer device based on a first tunnel, including: the offload card in the first computer device modifies a first service message according to a first tunnel table entry to obtain a first transit service message, and transmits the first transit service message to the offload card in the second computer device based on the first tunnel. The message header of the first transit service message is different from the message header of the first service message. The first transit service message includes the first service data. The first tunnel table entry is used to indicate a transmission rule for data transmission between the offload card in the first computer device and the offload card in the second computer device.

[0008] Since the virtual machine is migrated from the first computer device to the second computer device, after the first computer device receives the message sent to the virtual machine, the first computer device, as a relay device, only needs to modify the message header of the message and send the data sent to the virtual machine to the second computer device through the tunnel, thereby achieving the effect of lossless packet loss during the virtual machine migration process.

[0009] In another possible implementation, the source address in the first business message indicates the address of the third computer device, and the destination address in the first business message indicates the address of the first computer device; the unloading card in the first computer device modifies the first business message according to the first tunnel table entry to obtain a first transit business message, including: modifying the source address in the first business message to the address of the first computer device, and modifying the destination address in the first business message to the address of the second computer device, modifying the virtual extended local area network network identifier VNI in the first business message to the first VNI, and obtaining the first transit business message, the first virtual extended local area network indicated by the first VNI includes the first computer device and the second computer device.

[0010] In another possible implementation, the first service message further includes an identifier of the virtual machine; and the method further includes: an offload card in the first computer device determining a first tunnel table entry based on the identifier of the virtual machine and a first virtual eXtensible Local Area Network Identifier (VNI). The first virtual eXtensible Local Area Network indicated by the first VNI includes the first computer device and the second computer device.

[0011] Because different tunnels are built between different devices, tunnels are managed based on tunnel entries so that devices can forward packets based on the tunnel entries.

[0012] In another possible implementation, before the offload card in the first computer device transmits the first service data to the offload card in the second computer device based on the first tunnel, the method further includes: the offload card in the first computer device establishing the first tunnel.

[0013] By establishing the first tunnel, the first computer device acts as a relay device to forward the message to the second computer device, thereby achieving a lossless packet loss effect during the virtual machine migration process.

[0014] In another possible implementation, after the offload card in the first computer device transmits the first service data to the offload card in the second computer device via the first tunnel, the method further includes: the offload card in the first computer device receives a first transfer response message sent by the offload card in the second computer device via the first tunnel, modifies the first transfer response message according to the second tunnel table entry to obtain a first response message, and transmits the first response message to the third computer device via the second tunnel. The message header of the first transfer response message is different from the message header of the first response message. The first transfer response message includes the first response data. The first response message includes the first response data. The first response data is result data after the offload card in the second computer device processes the first service data. The second tunnel table entry is used to indicate the transmission rules for data transmission between the offload card in the first computer device and the offload card in the third computer device. The second tunnel is used to provide a path for data transmission between the offload card in the first computer device and the offload card in the third computer device.

[0015] Therefore, the first computer device acts as a relay device to modify the header of the transit response message and send the response data of the virtual machine to the third computer device through the tunnel. During the virtual machine migration process, the virtual machine's message can be processed and forwarded in real time, achieving the effect of lossless packet loss during the virtual machine migration process.

[0016] In another possible implementation, the source address in the first transit response message indicates the address of the second computer device, and the destination address in the first transit response message indicates the address of the first computer device; the unloading card in the first computer device modifies the first transit response message according to the second tunnel table entry to obtain a first response message, including: modifying the source address in the first transit response message to the address of the first computer device, and modifying the destination address in the first transit response message to the address of the third computer device, modifying the VNI in the first transit response message to the second VNI, and obtaining a first transit service message, the second virtual extended local area network indicated by the second VNI includes the first computer device and the third computer device.

[0017] In another possible implementation, the first relay response message further includes an identifier of the virtual machine; and the method further includes: an offload card in the first computer device determining a second tunnel entry based on the identifier of the virtual machine and the second VNI. The second virtual extended local area network indicated by the second VNI includes the first computer device and the third computer device.

[0018] In another possible implementation, the method further includes: an uninstall card of the first computer device deleting the first tunnel and the second tunnel according to the instruction.

[0019] Thus, a third tunnel is established between the offload card in the second computer device and the offload card in the third computer device. There is no need for the first computer device to act as a relay device to forward the virtual machine's message. The tunnel can be deleted in time to avoid occupying computing resources and storage resources, thereby improving resource utilization.

[0020] In another possible implementation, when a virtual machine migrates from a second computer device to a first computer device, an offload card in the first computer device receives a first transit service message based on a first tunnel, the first transit service message including first service data sent by a third computer device to the virtual machine in the second computer device, and the first tunnel is used to provide a path for data transmission between the offload card in the first computer device and the offload card in the second computer device; the offload card in the first computer device processes the first service data to obtain first response data; when the offload card in the first computer device and the offload card in the third computer device have not established a third tunnel, the offload card in the first computer device sends a first transit response message to the offload card in the second computer device based on the first tunnel, the first transit response message including first response data, and the third tunnel is used to provide a path for data transmission between the offload card in the second computer device and the offload card in the third computer device; when the offload card in the first computer device and the offload card in the third computer device establish a third tunnel, the offload card in the first computer device sends a first response message to the offload card in the third computer device based on the second tunnel, the first response message including first response data.

[0021] In another possible implementation, the method further includes: the unloading card in the first computer device generates a first transit response message according to the first tunnel table entry, the source address in the first transit service message indicates the address of the first computer device, and the destination address indicates the address of the second computer device.

[0022] In another possible implementation, the method further includes: the offload card in the first computer device generating a first response message according to the second tunnel entry, wherein the source address in the first response message indicates the address of the first computer device, and the destination address indicates the address of the third computer device.

[0023] In a second aspect, a virtual machine migration method is provided, including: when a virtual machine migrates from a first computer device to a second computer device, the first computer device receives a first business message sent by a third computer device to the virtual machine in the first computer device, the first business message including first business data; the unloading card in the first computer device transmits the first business data to the unloading card in the second computer device based on a first tunnel, the first tunnel being used to provide a path for data transmission between the unloading card in the first computer device and the unloading card in the second computer device.

[0024] In one possible implementation, the unloading card in the first computer device transmits the first business data to the unloading card in the second computer device based on the first tunnel, including: the unloading card in the first computer device modifies the first business message according to the first tunnel table entry to obtain a first transit business message, the message header of the first transit business message is different from the message header of the first business message, the first transit business message includes the first business data, and the first tunnel table entry is used to indicate the transmission rules for transmitting data between the unloading card in the first computer device and the unloading card in the second computer device; the unloading card in the first computer device transmits the first transit business message to the unloading card in the second computer device based on the first tunnel.

[0025] In another possible implementation, the method also includes: the offload card in the second computer device receives a first transit service message based on the first tunnel, the first transit service message includes first service data sent by the third computer device to the virtual machine in the first computer device, and the first tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the second computer device; the offload card in the second computer device processes the first service data to obtain first response data; when the offload card in the second computer device and the offload card in the third computer device have not established a third tunnel, the offload card in the second computer device sends a first transit response message to the offload card in the first computer device based on the first tunnel, and the first transit response message includes the first response data; when the offload card in the second computer device and the offload card in the third computer device establish a third tunnel, the offload card in the second computer device sends a first response message to the offload card in the third computer device based on the third tunnel, and the first response message includes the first response data, and the third tunnel is used to provide a path for transmitting data between the offload card in the second computer device and the offload card in the third computer device.

[0026] In another possible implementation, after the offload card in the first computer device transmits the first business data to the offload card in the second computer device based on the first tunnel, the method also includes: the offload card in the first computer device receives a first transit response message sent by the offload card in the second computer device based on the first tunnel, the first transit response message including first response data, and the first response data is result data after the offload card in the second computer device processes the first business data; the offload card in the first computer device modifies the first transit response message according to the second tunnel table entry to obtain a first response message, the message header of the first transit response message is different from the message header of the first response message, the first response message includes the first response data, and the second tunnel table entry is used to indicate a transmission rule for transmitting data between the offload card in the first computer device and the offload card in the third computer device; the offload card in the first computer device transmits the first response message to the third computer device based on the second tunnel, and the second tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the third computer device.

[0027] In another possible implementation, the method also includes: when the offload card in the second computer device establishes a third tunnel with the offload card in the third computer device, the offload card in the second computer device receives a second business message sent by the offload card in the third computer device based on the third tunnel, and the second business message includes second business data sent by the third computer device to the virtual machine in the second computer device; the offload card in the second computer device processes the second business data to obtain second response data; and the offload card in the second computer device sends a second response message to the offload card in the third computer device based on the third tunnel, and the second response message includes the second response data.

[0028] In another possible implementation, the method further includes: when the uninstall card in the second computer device establishes a third tunnel with the uninstall card in the third computer device, the second computer device deletes the first tunnel between the second computer device and the first computer device; and the first computer device deletes the first tunnel and the second tunnel according to the instruction.

[0029] In another possible implementation, the method further includes: the second computer device restarting the virtual machine.

[0030] In a third aspect, a data processing device is provided, the data processing device including modules for executing the virtual machine migration method in the first aspect or any possible design of the first aspect. For example, the data processing device includes a communication module and a processing module.

[0031] The communication module is configured to receive a first service message sent from a third computer device to a virtual machine in the first computer device, where the first service message includes first service data.

[0032] The processing module is used to transmit the first service data to the offload card in the second computer device based on the first tunnel, and the first tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the second computer device.

[0033] In one possible implementation, the processing module is specifically configured to modify a first service message based on a first tunnel table entry to obtain a first transit service message, and the communication module is further configured to transmit the first transit service message to an offload card in a second computer device based on the first tunnel. The message header of the first transit service message is different from the message header of the first service message. The first transit service message includes first service data. The first tunnel table entry is configured to indicate a transmission rule for data transmission between the offload card in the first computer device and the offload card in the second computer device.

[0034] In another possible implementation, the source address in the first business message indicates the address of the third computer device, and the destination address in the first business message indicates the address of the first computer device; the processing module is specifically used to modify the source address in the first business message to the address of the first computer device, and modify the destination address in the first business message to the address of the second computer device, and modify the virtual extended local area network identifier VNI in the first business message to the first VNI to obtain a first transit business message, and the first virtual extended local area network indicated by the first VNI includes the first computer device and the second computer device.

[0035] In another possible implementation, the first service message further includes an identifier of the virtual machine; the processing module is further configured to determine a first tunnel entry based on the identifier of the virtual machine and the first VNI. The first virtual extended local area network indicated by the first VNI includes a first computer device and a second computer device.

[0036] In another possible implementation, the processing module is further configured to establish a first tunnel.

[0037] In another possible implementation, the communication module is further used to receive a first transit response message sent by the offload card in the second computer device based on the first tunnel, the processing module is further used to modify the first transit response message according to the second tunnel table entry to obtain a first response message, and the communication module is further used to transmit the first response message to the third computer device based on the second tunnel. The message header of the first transit response message is different from the message header of the first response message. The first transit response message includes first response data. The first response message includes first response data. The first response data is result data after the offload card in the second computer device processes the first business data. The second tunnel table entry is used to indicate the transmission rules for transmitting data between the offload card in the first computer device and the offload card in the third computer device. The second tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the third computer device.

[0038] In another possible implementation, the first relay response message also includes an identifier of the virtual machine; the processing module is further configured to determine a second tunnel entry based on the identifier of the virtual machine and the second VNI. The second virtual extended local area network indicated by the second VNI includes the first computer device and the third computer device.

[0039] In another possible implementation, the processing module is further configured to delete the first tunnel and the second tunnel according to the instruction.

[0040] In a fourth aspect, an uninstall card is provided, comprising: a processor and a power supply circuit; wherein the power supply circuit is used to power the processor; and the processor is used to execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0041] In a fifth aspect, a computer device is provided, comprising a processor and an offload card, wherein the processor is used to run a virtual machine; when the virtual machine migrates to another computer device, the offload card is used to execute the operating steps of the method in the first aspect or any possible implementation of the first aspect, thereby relaying and forwarding the virtual machine's messages after the virtual machine migrates.

[0042] In a sixth aspect, a communication system is provided, which includes multiple computer devices, and the computer devices include the unloading card as described in the fourth aspect, and the unloading card is used to execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0043] In the seventh aspect, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a processor, the processor executes the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0044] In an eighth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is caused to execute the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0045] The technical effects brought about by any design method in the second to eighth aspects can be referred to the technical effects brought about by the first aspect or different design methods in the first aspect, and will not be repeated here.

[0046] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG1 is a schematic diagram of the architecture of a communication system provided by the present application;

[0048] FIG2 is a schematic diagram of a virtual machine migration architecture provided by this application;

[0049] FIG3 is a schematic diagram of the virtual machine migration process provided by this application, which includes three stages;

[0050] FIG4 is a schematic diagram of a process of relaying a message of a virtual machine during a virtual machine migration process provided by the present application;

[0051] FIG5 is a schematic diagram of a process of relaying a message of a virtual machine during a virtual machine migration process provided by the present application;

[0052] FIG6 is a schematic diagram of a process for transmitting a message of a virtual machine during a virtual machine migration process provided by the present application;

[0053] FIG7 is a schematic structural diagram of a data processing device provided by the present application;

[0054] FIG8 is a schematic structural diagram of a computer device provided in this application. DETAILED DESCRIPTION

[0055] To facilitate understanding, the main terms involved in this application are first explained.

[0056] Virtualization technology abstracts and virtualizes hardware resources such as computing, networking, storage, and input / output (I / O), allowing them to be used in separate, distributed systems. Virtualization can effectively improve hardware resource utilization and flexibility.

[0057] A virtual machine (VM) is a complete computer system that simulates the full functionality of a hardware system through software and runs in a completely isolated environment. When creating a VM on a physical computer, the physical computer's hard disk and memory capacity are used as the VM's. Each VM has independent computing, networking, storage, and operating system resources. A VM can perform the same tasks as a physical computer. Examples include Java virtual machines, Linux virtual machines, and Windows virtual machines.

[0058] Virtual machine migration is the process of moving a virtual machine between different computing or storage resources. For example, migrating a virtual machine from one physical device to another. Virtual machine migration includes cold migration and hot migration.

[0059] Cold migration involves migrating a powered-off or suspended virtual machine from a source device to a destination device. During the migration process, the virtual machine's data is relocated to a new storage location. This virtual machine data includes configuration files and application files. Configuration files contain the virtual machine's configuration data, while application files contain the data of applications running on the virtual machine.

[0060] Hot migration involves migrating a running virtual machine from a source device to a destination device, which then smoothly restarts the virtual machine. In principle, the entire process should be invisible to the VM's tenants and minimize the impact on their services.

[0061] Because a virtual machine's data may not be transmitted to the destination device all at once, after live migration is initiated, the source device migrates the virtual machine's data (for example, application data in memory) to the destination device through multiple iterations of copying. During the final copy of the remaining virtual machine data, the source device shuts down the virtual machine and suspends sending and receiving messages to the virtual machine. Consequently, if external devices send messages to the virtual machine on the source device, packet loss may occur.

[0062] During live migration, the VM shutdown time is minimized to minimize packet loss. After the VM migration is complete, the destination device restarts the VM and relies on network reliability protocols (such as packet retransmission) to recover lost packets. However, packets from transport layer protocols that don't support reliability mechanisms, such as the User Datagram Protocol (UDP), cannot be recovered.

[0063] With the rapid development of cloud platforms, virtualization technology has gradually become one of the key technologies of cloud platforms. With the development of fields such as High Performance Computing (HPC), cloud computing, and big data, as well as data processing units (DPUs), high-performance network protocols such as Remote Direct Memory Access (RDMA) and Remote Direct Memory Access over Converged Ethernet (RoCE) have become widely used. The RoCE protocol does not support selective retransmission and is sensitive to packet loss. In other words, if packet loss occurs, all unacknowledged packets will be retransmitted, which has a significant impact on performance. As high-performance network protocols such as RoCE are gradually deployed in cloud scenarios, it is becoming a trend to include hot migration functions for high-performance network services such as RoCE.

[0064] High-performance network protocols are used in high-bandwidth, low-latency scenarios. They are sensitive to packet loss and latency and support customizable timeouts. However, hot migration solutions are not friendly to high-performance network protocols in the following situations.

[0065] First, during the brief shutdown phase of a virtual machine, some packet loss is inevitable. This loss can be even more severe in high-bandwidth scenarios. If packet loss occurs during the live migration process, it can easily cause timeouts and disconnections, making tenants aware of the issue.

[0066] Secondly, when the virtual machine is restarted, some lost packets are restored through mechanisms such as protocols, but the latency cannot be guaranteed, and the performance of the high-performance network will be affected.

[0067] Finally, in virtualized scenarios, when a VM has many external communication partners (i.e., large tunnel entries), hot migration can struggle to synchronize all the entries to the destination device all at once. Even if the VM restarts the destination device, communication still cannot be quickly and effectively restored.

[0068] Virtual Machine Monitor (VMM): This is the middleware software that runs between the source device and the virtual machine and is used to manage the virtual machine. VMM is also called a hypervisor.

[0069] Virtual eXtensible Local Area Network (VXLAN): A standard technology for Network Virtualization over Layer 3 (NVO3) defined by the IETF, VXLAN is an extension of the traditional VLAN protocol. VXLAN encapsulates Layer 2 (L2) Ethernet frames into UDP packets (i.e., L2 over L4) and transmits them across a Layer 3 (L3) network. Tunnel nodes in VXLAN are called VXLAN Tunnel Endpoints (VTEPs).

[0070] I / O passthrough is a hardware-based virtualization technology that virtualizes a physical network interface card (NIC) into multiple virtual NICs and provides them to virtual machines, giving them the same I / O capabilities as physical devices. High-performance NICs typically support hardware virtualization passthrough, such as single-root I / O virtualization.

[0071] In order to solve the problem of long virtual machine migration delay and packet loss during virtual machine migration. The present application provides a virtual machine migration method, that is, when a virtual machine migrates from a source device to a destination device, the source device receives a service message sent by an external device to the virtual machine in the source device, and the unloading card in the source device transmits the service data included in the service message to the unloading card in the destination device based on the tunnel between the source device and the destination device. The unloading card of the destination device feeds back a response message through the tunnel between the source device and the destination device. After the external device and the destination device establish a tunnel, the external device transmits the service message of the virtual machine to the destination device through the tunnel between the external device and the destination device, and the destination device feeds back a response message to the external device based on the tunnel between the destination device and the external device.

[0072] Compared with the problem of retransmitting messages by using a reliability mechanism or using VMM to cache messages in order to solve the problem of long virtual machine migration time in order to solve the packet loss problem in virtual machine migration, the virtual machine migration method provided by the present application provides a message lossless relay mechanism in the process of virtual machine migration based on the unloading capability of the unloading card and the tunnel technology. That is, after the virtual machine migrates from the source device to the destination device, the unloading card in the source device processes the message sent by the external device to the virtual machine in the source device, and forwards the message to the destination device through the tunnel. Therefore, during the virtual machine migration process, the message of the virtual machine can be processed and forwarded in real time, avoiding the packet loss phenomenon in virtual machine migration, and eliminating the need for VMM to cache messages or retransmit messages by using a reliability mechanism. This reduces the storage space occupied by the VMM or the storage space of the source device, reduces the processor in the source device from participating in message processing, reduces the message processing delay, enables the destination device to restart the virtual machine in time, and effectively shortens the delay of virtual machine migration.

[0073] The virtual machine migration method provided in this application is applicable to high-performance transmission protocols that are sensitive to packet loss and delay, such as the RoCE protocol.

[0074] The virtual machine migration method provided by this application is described in detail below with reference to the accompanying drawings.

[0075] FIG1 is a schematic diagram of the architecture of a communication system provided in this application. The communication system may be a physical architecture of a cloud system. For example, communication system 100 includes multiple computer devices. The multiple computer devices include computer device 110, computer device 120, and computer device 130.

[0076] A computer device can be a device that provides computing functions (such as a server). For example, a computer device may include a central processing unit (CPU), a graphics processing unit (GPU), a DPU, a neural processing unit (NPU), and an embedded neural network processing unit (NPU) with computing capabilities to provide high-performance computing. A computer device may also be referred to as a computing server or application server. Multiple servers that provide computing functions can form a computing cluster.

[0077] Figure 1 is merely a schematic diagram. The embodiments of this application do not limit the number of devices included in the communication system. For example, the communication system may include multiple storage devices to form a storage system to provide a large storage capacity for the computer device. Optionally, the communication system may also include other devices, such as network devices (e.g., switches and routers). Multiple computer devices are connected via the network devices and transmit data through the network devices.

[0078] For example, the communication system 100 further includes a client 140 communicating with multiple computing devices via a network 150. The network 150 may be an intranet within an enterprise, such as a local area network (LAN) or the Internet.

[0079] In this application, virtual machines are deployed on computer devices. This application does not limit the number of virtual machines. In some embodiments, when deploying services such as high-performance computing and artificial intelligence in a system, a large number of virtual machines are running on the computer devices. A virtual machine on a computer device can communicate with virtual machines on multiple computer devices. A high-performance tunnel can be established between any two computer devices using the RDMA protocol and the VXLAN protocol, and data can be transmitted between the computer devices based on the high-performance tunnel.

[0080] In some scenarios, virtual machines need to be migrated. For example, scenarios such as virtual machine migration include maintenance or upgrades of the device where the virtual machine resides, load balancing, and failover.

[0081] For example, Figure 2 is a schematic diagram of a virtual machine migration architecture provided by this application. Communication system 100 includes computer devices 110, 120, and 130. VM1 is deployed on computer device 110, and VM2 is deployed on computer device 130. A tunnel is established between computer devices 110 and 130, for example, a high-performance VXLAN tunnel based on the VXLAN protocol. VM1 on computer device 110 and VM2 on computer device 130 can transmit data through the tunnel.

[0082] In some embodiments, the offload card of the computer device provides the virtual machine with the function of a virtual network card, that is, the offload cards of the two computer devices establish a tunnel to provide data transmission function between the virtual machines of the computer devices.

[0083] An offload card can be a flexible, programmable chip on a chip (Chip) with high computing power and high-performance data processing capabilities. For example, an offload card can be a DPU or a SmartNIC. In future computing, DPUs will be the third largest chip processor after CPUs and GPUs. They are used to offload complex computations that are difficult for general-purpose computing power to handle. They are primarily used in data-centric scenarios such as networking, virtualization, storage, and security.

[0084] For example, the offload card of computer device 110 forwards a message to computer device 130 based on the tunnel entry. The offload card of computer device 130 forwards the message to computer device 110 based on the tunnel entry. Tunnel entries are used to record the transmission rules for data transmission between computer devices. Optionally, the tunnel between computer device 110 and computer device 130 can be a unidirectional tunnel. Computer device 110 establishes a tunnel from computer device 110 to computer device 130, and computer device 110 transmits data to computer device 130 based on the tunnel from computer device 110 to computer device 130. Computer device 130 establishes a tunnel from computer device 130 to computer device 110, and computer device 130 transmits data to computer device 110 based on the tunnel from computer device 130 to computer device 110.

[0085] Optionally, the tunnel table entry described in this application may be a VTEP table entry. The tunnel table entry includes a source address, a destination address, and a virtual eXtensible Local Area Network Identifier (VNI). The source address indicates the VXLAN tunnel endpoint from which the message is sent. The destination address indicates the VXLAN tunnel endpoint from which the message is received. The VNI indicates the virtual eXtensible Local Area Network.

[0086] Furthermore, after receiving a message, the offload card can use I / O passthrough technology to transmit the message to the virtual machine running on the host computer device, allowing the host virtual machine to process the message. Hardware passthrough can avoid VMM involvement, reduce dependencies between components, and improve the success rate of hot migration. For example, if the offload card of computer device 110 receives a message sent to VM1 by computer device 130, the offload card of computer device 110 can transmit the message to VM1 running on host 111, which will then process the message.

[0087] The communication system may further include a management device 160. The management device 160 is used to provide management functions for virtual machines in the system. For example, management functions include deployment, deletion, and migration. In this application, the management device 160 is also used to provide tunnel entry generation and modification functions.

[0088] It should be noted that in actual applications, a single computer device can establish tunnels with multiple other computers to transmit virtual machine data. When a virtual machine migrates from a source device to a destination device, the management device must simultaneously modify the tunnel entries on the computer devices associated with the virtual machine. For example, the address of the device hosting the virtual machine in the tunnel entry must be modified. Due to the large number of tunnel entries and the delay in migrating the tunnel entries from the offload card on the source computer device to the offload card on the destination computer device after the management device issues the command, the resulting delay in switching all tunnel entries is significant, impacting virtual machine restart time.

[0089] In this application, when a virtual machine migrates from a source device to a destination device, the source device can act as a relay, providing a springboard function to transmit the virtual machine's messages to the destination device. For example, after receiving a message sent to the virtual machine, the source device forwards the virtual machine's message to the destination device through the tunnel between the source and destination devices. The destination device then forwards a response message to the source device through the tunnel between the source and destination devices.

[0090] Understandably, since the virtual machine migrates from the source device to the destination device, the virtual machine on the source device suspends sending and receiving messages, and the virtual machine on the destination device has not been restarted. The external device still sends the virtual machine's message to the source device based on the tunnel table entry. The appropriate time for relaying is when the virtual machine on the source device has been shut down, but the binding between the virtual machine on the source device and the virtual network card for I / O direct access on the source device has not been released, and the virtual machine on the destination device has not been restarted, but the binding between the virtual machine on the destination device and the virtual network card for I / O direct access on the destination device has been established. The virtual network card for I / O direct access associated with the virtual machine on the source device and the virtual machine on the destination device has the ability to send and receive messages.

[0091] For example, after VM1 migrates from computer device 110 to computer device 120, computer device 110 can act as a relay device to forward VM1's messages. After receiving the message sent by computer device 130 to VM1, the offload card of computer device 110 transmits the message to the offload card of computer device 120 through the tunnel between the offload cards of computer device 110 and 120, so that the offload card of computer device 120 can process VM1's message. The offload card of computer device 120 can also send a response message back to the offload card of computer device 110 through the tunnel between the offload cards of computer device 110 and 120.

[0092] After the computer device 130 and the computer device 120 establish a tunnel, the computer device 130 and the computer device 120 may transmit messages through the tunnel between the computer device 130 and the computer device 120 .

[0093] As a virtual machine migrates from a source device to a destination device, multiple external devices communicating with the virtual machine need to modify their tunnel entries. In this application, because the source device acts as a relay to forward the virtual machine's messages, the destination device can start the virtual machine before multiple external devices have completed switching their tunnel entries. This effectively shortens the latency of virtual machine migration, enabling real-time processing and forwarding of virtual machine messages during the migration process, and avoiding packet loss during virtual machine migration.

[0094] In addition, the offload card carries the services offloaded from the upper layer and can process messages without the participation of a processor (such as a CPU). There is no need for the VMM to cache messages or use a reliability mechanism to retransmit messages, which reduces the storage space occupied by the VMM or the storage space of the source device and reduces the message processing delay.

[0095] Next, the virtual machine migration process is described in detail with reference to the accompanying drawings. For example, as shown in Figure 3, the virtual machine migration process provided by this application includes three stages. Here, the computer device 110, computer device 120 and computer device 130 in Figure 1 are used as examples for illustration. For ease of description, computer device 110 can be briefly referred to as device A, computer device 120 can be briefly referred to as device B, and computer device 130 can be briefly referred to as device C. Device A can refer to the source device of the virtual machine migration, device B can refer to the destination device of the virtual machine migration, and device C can refer to the external device of the virtual machine migration. Among them, there may be multiple external devices, and one is used as an example for illustration in the figure. Each external device can send a message to the migrated virtual machine with reference to the method provided in this application.

[0096] Phase 1: The virtual machine migrates from the source device to the destination device. A tunnel is established between the source device and the destination device, but no tunnel is established between the external device and the destination device.

[0097] 1. VM1 is migrated from the source device to the destination device. After VM1 data migration is complete and VM1 on the destination device is reactivated, the virtual network adapters (VNICs) associated with VM1 on both the source and destination devices are still able to send and receive packets.

[0098] 2. The source device and the destination device each create a tunnel and set the tunnel to active. The source device can act as a relay device, or alternatively, the source device provides a springboard function. The springboard function is offloaded to the virtual network card for I / O passthrough abstracted by the offload card. The source device creates a VXLAN tunnel (such as VTEP-AB) between the source device and the destination device, which is used to transmit VM1's packets to the destination device. The destination device creates a VXLAN tunnel (such as VTEP-BA) between the destination device and the source device, which is used to transmit VM1's packets to the source device.

[0099] 3. Because the external device's tunnel entry has not been updated, all traffic destined for VM1 continues to be sent from the external device to the source device. For example, if VM2 on device C sends a packet to VM1, device C will continue to send VM1's packet to device A because the tunnel entry stored on device C indicates that the packet should be sent to device A.

[0100] 4. When the springboard is activated, the source device intercepts the message sent to VM1 and forwards the message from VM1 to the destination device based on the tunnel between the source device and the destination device.

[0101] 5. When a tunnel is not established between the external device and the destination device (e.g., VTEP-BC), the message processed by the destination device is still fed back to the source device via the tunnel between the destination device and the source device.

[0102] 6. If the tunnel between the source device and the external device (e.g., VTEP-AC) is not updated, the source device relays the message from VM1 to the external device.

[0103] Phase 2: The external device and the destination device establish a tunnel.

[0104] 7. The destination device establishes a tunnel from the destination device to the external device. The destination device can send VM1's message to the external device based on the tunnel from the destination device to the external device without forwarding VM1's message through the source device.

[0105] For example, device B creates a tunnel entry VTEP-BC from device B to device C. Packets sent from VM1 to VM2 can bypass device A and be sent to VM2 on device C according to the tunnel entry VTEP-BC. Alternatively, device B can only receive packets from VM1 forwarded by device A, but can also reply to VM1's packets without passing through device A.

[0106] 8. The external device establishes a tunnel from the external device to the destination device. The external device can send VM1's message to the destination device based on the tunnel from the external device to the destination device, without forwarding VM1's message through the source device.

[0107] For example, device C creates a tunnel entry VTEP-CB from device C to device B. Packets sent from VM2 to VM1 can be sent to VM1 on device B according to the tunnel entry VTEP-CB, bypassing device A. Alternatively, device B can only receive packets from VM1 forwarded by device C, but no longer send packets from VM1 through device A.

[0108] Phase 3: Closing the tunnel.

[0109] 9. Delete the tunnel between the destination device and the source device, and delete the tunnel between the external device and the source device.

[0110] The tunnel relay solution uses a make-before-break approach, gradually refreshing the tunnel between the destination device and the external device before disconnecting the tunnel between the source and external devices. The key is that during live migration, offload features such as tunnel table lookup, security rules, and connection tracking are implemented on the source device, while message acknowledgment is completed by the destination device.

[0111] Each stage is described below.

[0112] FIG4 is a flow chart of relaying a message of a virtual machine during a virtual machine migration process provided by the present application. Here, the first stage is mainly described in detail. As shown in FIG4 (a), the method includes the following steps.

[0113] Step 410: When the virtual machine VM1 is migrated from the source device to the destination device, the external device sends a first service message to the source device, where the first service message includes first service data.

[0114] The virtual machine management and control platform (e.g., a management device) can instruct the source device to initiate virtual machine migration. For example, the management device instructs device A to migrate VM1 from device A to device B.

[0115] When the data copy of VM1 is completed, the management device instructs the source device and the destination device to temporarily establish a first tunnel dedicated to VM1, so that the source device acts as a relay device and provides a springboard function, that is, the source device transmits VM1's message based on the first tunnel. The springboard function is offloaded to the virtual network card associated with VM1 in the offload card. The first tunnel is used to provide a path for the offload card in the source device to transmit data to the offload card in the destination device. The offload card of the source device and the offload card of the destination device newly store the first tunnel table entry. The first tunnel table entry can be a relay table entry. The first tunnel table entry is used to indicate the transmission rules for the offload card in the source device to transmit data to the offload card in the destination device.

[0116] In some embodiments, device A creates a VXLAN tunnel between device A and device B, namely, VTEP-AB, for device A to transmit packets from VM1 to device B. Device B creates a VXLAN tunnel between device B and device A, namely, VTEP-BA, for device B to transmit packets from VM1 to device A. As shown in Table 1, the offload card on device A adds a tunnel entry for storing VTEP-AB.

[0117] Table 1

[0118] As shown in Table 2, the offload card of device B newly stores the tunnel entry for VTEP-BA.

[0119] Table 2

[0120] As shown in Tables 1 and 2, the source address and destination address are different. The source address indicates the device that sends the message. The destination address indicates the device that receives the message.

[0121] In addition, the management device instructs the external device to refresh the tunnel entry for the migrated virtual machine. For example, the tunnel entry for device C to transmit packets to VM1 indicates that packets from VM1 should be sent to device A. The management device instructs device C to refresh the tunnel entry for device C to transmit packets to VM1. That is, the tunnel entry indicates that packets from VM1 should be sent to device B.

[0122] Because the tunnel entry for sending packets to VM1 recorded by the external device has not been updated, all traffic to VM1 will still be sent from VM1 to the source device. For example, if VM2 on device C sends a packet to VM1, device C will still send VM1's packet to device A because the tunnel entry stored on device C indicates that the packet should be sent to device A.

[0123] Step 420: The offload card in the source device transmits the first service data to the offload card in the destination device based on the first tunnel.

[0124] The source device receives a first service message sent from an external device to VM1 in the source device. The offload card in the source device modifies the VXLAN header in the first service message based on the first tunnel entry to obtain a first transit service message. The header of the first transit service message is different from the header of the first service message, and the first transit service message includes the first service data. Understandably, the payload in the first service message remains unchanged; only the VXLAN header in the first service message is modified. This allows the first transit service message to be transmitted through the first tunnel between the source device and the destination device, and can also increase the message transmission rate.

[0125] In some embodiments, the first service message also includes an identifier of the virtual machine, and the unloading card in the source device determines the first tunnel table entry based on the identifier of the virtual machine and the first VNI. The first virtual extended local area network indicated by the first VNI includes a source device and a destination device. Optionally, the unloading card in the source device stores a correspondence between the first VNI and the second VNI. The unloading card in the source device obtains the second VNI from the first service message, queries the correspondence based on the second VNI, and obtains the first VNI corresponding to the second VNI.

[0126] For example, device A receives the first service message sent by device C to VM1. Device A determines the VTEP-AB tunnel entry shown in Table 1 based on the address of device A, the identifier of VM1 and the first VNI, modifies the VXLAN message header in the first service message, and obtains the first transit service message.

[0127] As shown in Table 3, the VXLAN header in the first service packet and the VXLAN header in the first transit service packet.

[0128] Table 3

[0129] As shown in Table 3, the source address in the VXLAN header of the first service packet indicates the MAC address and IP address of device C, and the destination address indicates the MAC address and IP address of device A. The source address in the VXLAN header of the first transit service packet indicates the MAC address and IP address of device A, and the destination address indicates the MAC address and IP address of device B.

[0130] The offload card in the source device transmits the first transit service message to the offload card in the destination device based on the first tunnel.

[0131] As shown in (b) of FIG4 , the method further includes the following steps.

[0132] Step 430: The offload card in the destination device processes the first service data to obtain first response data.

[0133] The offload card in the destination device receives a first transit service message based on the first tunnel. The first transit service message includes first service data sent by the external device to the virtual machine in the source device.

[0134] Since the data of VM1 has been migrated to the destination device, the offload card in the destination device can obtain the data of VM1, process the first service data according to the data of VM1, and obtain the first response data.

[0135] Step 440: When the offload card in the destination device and the offload card in the external device have not established the third tunnel, the offload card in the destination device sends a first transfer response message to the offload card in the source device based on the first tunnel, where the first transfer response message includes first response data.

[0136] The offload card in the destination device generates a first transfer response message according to the first tunnel entry. The message header of the first transfer service message is different from the message header of the first transfer response message. The first transfer response message includes first response data.

[0137] In some embodiments, the first transit service message also includes an identifier of the virtual machine, and the unloading card in the destination device determines the first tunnel table entry based on the identifier of the virtual machine and the first VNI. The first virtual extended LAN indicated by the first VNI includes a source device and a destination device.

[0138] For example, device B receives a first transit service message sent by device A to VM1. The first transit service message includes the address of device A, the address of device B, the identifier of VM1, and the first VNI. Device B determines the VTEP-BA tunnel table entry shown in Table 2 based on the address of device B, the identifier of VM1, and the first VNI, and generates a first transit response message. The first transit response message includes the address of device A, the address of device B, the identifier of VM1, and the first VNI.

[0139] As shown in Table 4, the VXLAN header in the first transit service packet and the VXLAN header in the first transit response packet.

[0140] Table 4

[0141] As shown in Table 4, the source address in the VXLAN header of the first relay service packet indicates the MAC address and IP address of device A, and the destination address indicates the MAC address and IP address of device B. The source address in the VXLAN header of the first relay response packet indicates the MAC address and IP address of device B, and the destination address indicates the MAC address and IP address of device A.

[0142] The offload card in the destination device transmits the first forwarding response message to the offload card in the source device over the first tunnel. Understandably, if the offload card in the destination device and the offload card in the external device have not established a third tunnel, the destination device may return the first forwarding response message along the original path, so that the external device can receive result data after the offload card in the destination device processes the first service data.

[0143] Step 450: The offload card in the source device modifies the first transit response message according to the second tunnel entry to obtain a first response message.

[0144] The source device receives a first relay response message sent from the destination device to the source device. The offload card in the source device modifies the VXLAN header in the first relay response message based on the second tunnel entry to obtain a first response message. The header of the first relay response message is different from the header of the first response message. The first response message includes first response data. The second tunnel entry is used to indicate a transmission rule for data transmission between the offload card in the source device and the offload card in the external device.

[0145] Understandably, the payload in the first relay response message is not modified, only the VXLAN message header in the first relay response message is modified, so that the first response message is transmitted in the second tunnel between the source device and the external device.

[0146] In some embodiments, the first transit response message also includes the identifier of the virtual machine, and the unloading card in the source device determines the second tunnel table entry based on the identifier of the virtual machine and the second VNI, and the second virtual extended LAN indicated by the second VNI includes the source device and the external device.

[0147] For example, device A receives a first relay response message sent by device B to VM1. Device A determines a VTEP-AC tunnel entry based on the address of device A, the identifier of VM1, and the second VNI, modifies the VXLAN header in the first relay response message, and obtains a first response message. The second virtual extended LAN indicated by the second VNI includes the source device and the external device.

[0148] As shown in Table 5, the VXLAN header in the first relay response message and the VXLAN header in the first response message.

[0149] Table 5

[0150] As shown in Table 5, the source address in the VXLAN header of the first relay response packet indicates the MAC address and IP address of device B, and the destination address indicates the MAC address and IP address of device A. The source address in the VXLAN header of the first response packet indicates the MAC address and IP address of device A, and the destination address indicates the MAC address and IP address of device C.

[0151] Step 460: The offload card in the source device transmits the first response message to the external device based on the second tunnel.

[0152] The second tunnel is used to provide a path for transmitting data between the offload card in the source device and the offload card in the external device.

[0153] Step 470: When the offload card in the destination device establishes a third tunnel with the offload card in the external device, the offload card in the destination device sends a first response message to the offload card in the external device based on the third tunnel. The first response message includes first response data.

[0154] The offload card in the destination device generates a first response message according to the third tunnel entry. The message header of the first transit service message is different from the message header of the first response message, and the first response message includes first response data.

[0155] In some embodiments, the first transit service message also includes the identifier of the virtual machine, the unloading card in the destination device determines the third tunnel table entry based on the identifier of the virtual machine and the second VNI, and the second virtual extended LAN indicated by the second VNI includes an external device and a destination device.

[0156] For example, device B receives a first transit service message sent by device A to VM1. The first transit service message includes the address of device A, the address of device B, the identifier of VM1, and the first VNI. Device B determines the VTEP-BC tunnel table entry based on the address of device B, the identifier of VM1, and the second VNI, and generates a first response message. The first response message includes the address of device B, the address of device C, the identifier of VM1, and the second VNI.

[0157] As shown in Table 6, the VXLAN header in the first transit service packet and the VXLAN header in the first response packet.

[0158] Table 6

[0159] As shown in Table 6, the source address in the VXLAN header of the first transit service packet indicates the MAC address and IP address of device A, and the destination address indicates the MAC address and IP address of device B. The source address in the VXLAN header of the first response packet indicates the MAC address and IP address of device B, and the destination address indicates the MAC address and IP address of device C.

[0160] Understandably, when the offload card in the destination device establishes the third tunnel with the offload card in the external device, the offload card in the destination device transmits the first response message to the offload card in the external device via the third tunnel, so that the external device can receive result data after the offload card in the destination device processes the first service data. The third tunnel is used to provide a path for data transmission between the offload card in the destination device and the offload card in the external device.

[0161] It should be noted that when the tunnel from the unloading card in the external device to the unloading card in the destination device is not established, the external device still sends VM1's message to the source device, and the source device forwards VM1's message to the destination device. When the tunnel from the unloading card in the destination device to the unloading card in the external device is not established, the destination device still sends VM1's response message to the source device, and the source device forwards VM1's response message to the external device. When the tunnel from the unloading card in the destination device to the unloading card in the external device is established, the destination device still sends VM1's response message to the external device, and there is no need for the source device to forward VM1's response message to the external device.

[0162] When a tunnel is established from the offload card in the external device to the offload card in the destination device, the external device still sends VM1's messages to the destination device, and there is no need for the source device to forward VM1's messages to the destination device. When a tunnel is not established from the offload card in the destination device to the offload card in the external device, the destination device still sends VM1's response messages to the source device, and the source device forwards VM1's response messages to the external device. When a tunnel is established from the offload card in the destination device to the offload card in the external device, the destination device still sends VM1's response messages to the external device, and there is no need for the source device to forward VM1's response messages to the external device.

[0163] Figure 5 is a flow chart of a message relaying a virtual machine during a virtual machine migration process provided by the present application. Here, the second stage is mainly described in detail. As shown in Figure 5, the method includes the following steps.

[0164] Step 510: When the offload card in the destination device establishes a third tunnel with the offload card in the external device, the external device sends a second service message to the destination device, where the second service message includes second service data.

[0165] Because the offload card in the destination device establishes a third tunnel with the offload card in the external device, the tunnel entry for packets sent to VM1, recorded by the external device, is updated. For all traffic destined for VM1, the external device sends VM1 packets to the destination device. For example, if VM2 in device C sends a packet to VM1, device C sends VM1's packet to device B because the tunnel entry stored in device C indicates that the packet should be sent to device B.

[0166] For example, the third tunnel entry is as shown in Table 7.

[0167] Table 7

[0168] The VXLAN packet header in the second service packet may include the address of device C, the address of device B, the identifier of VM1, and the second VNI as shown in Table 7.

[0169] Step 520: The offload card in the destination device processes the second service data to obtain second response data.

[0170] The offload card in the destination device receives the second service message sent by the offload card in the external device via the third tunnel. Since the data of VM1 has been migrated to the destination device, the offload card in the destination device can obtain the data of VM1, process the first service data according to the data of VM1, and obtain the first response data.

[0171] Step 530: The offload card in the destination device sends a second response message to the offload card in the external device based on the third tunnel, where the second response message includes second response data.

[0172] The offload card in the destination device generates a second response message according to the third tunnel entry. The message header of the second service message is different from the message header of the second response message, and the second response message includes second response data.

[0173] In some embodiments, the second service message also includes the identifier of the virtual machine, and the unloading card in the destination device determines the third tunnel table entry based on the identifier of the virtual machine and the second VNI, and the second virtual extended LAN indicated by the second VNI includes an external device and a destination device.

[0174] For example, device B receives the second service message sent by device C to VM1. Device B determines the VTEP-BC tunnel entry according to the address of device B, the identifier of VM1, and the second VNI, and generates a second response message.

[0175] As shown in Table 8, the VXLAN header in the second service packet and the VXLAN header in the second response packet.

[0176] Table 8

[0177] As shown in Table 8, the source address in the VXLAN header of the second service packet indicates the MAC address and IP address of device C, and the destination address indicates the MAC address and IP address of device B. The source address in the VXLAN header of the second response packet indicates the MAC address and IP address of device B, and the destination address indicates the MAC address and IP address of device C.

[0178] It should be noted that before the destination device restarts the virtual machine, it can forward VM1's packets through the source device to facilitate processing of the first service data by the offload card in the destination device, thus preventing packet loss during the virtual machine migration. After the destination device establishes a tunnel with the external device, the external device and the destination device can transmit VM1's packets over the tunnel. This solution does not affect the destination device's ability to restart the virtual machine.

[0179] After a tunnel is established between the destination device and the external device, the source device and the destination device can delete the tunnel between them.

[0180] FIG6 is a flow chart of a method for transmitting a message of a virtual machine during a virtual machine migration process provided by the present application. Here, the third stage is mainly described in detail. As shown in FIG6, the method includes the following steps.

[0181] Step 610: The uninstall card of the source device deletes the first tunnel and the second tunnel according to the instruction.

[0182] The source device destroys the tunnel between the source device and the destination device according to the instruction of the management device. The source device's uninstall card deletes the tunnel table entry between the source device and the external device, and the source device's uninstall card deletes the tunnel table entry between the source device and the destination device.

[0183] The offload card of the destination device deletes the tunnel entry between the source device and the destination device.

[0184] It should be noted that, after ensuring that the message processing of VM1 is completed, the source device and the destination device destroy the tunnel between the source device and the destination device.

[0185] The present application provides a virtual machine migration method, (1) based on the offloading capability of the offloading card and the tunnel technology, a lossless relay mechanism for messages is designed. After the virtual machine migrates from the source device to the destination device and establishes an association relationship with the I / O direct virtual network card of the destination device, but before it is started, the association relationship between the source device and the I / O direct virtual network card of the source device is still retained, the offloading capability of the virtual network card is utilized, and the I / O direct virtual network card of the source device is used as a dedicated tunnel relay device to send and receive messages to the I / O direct virtual network card of the destination device associated with the destination device for processing. When the tunnel table entry is not fully synchronized, the I / O direct virtual network card of the destination device can forward the processed message to the external device through the original path of the I / O direct virtual network card of the source device; otherwise, the I / O direct virtual network card of the destination device can forward the processed message directly to the external device.

[0186] (2) Based on the lossless relay mechanism, combined with the offload capability and I / O pass-through capability of the offload card, a springboard capability is implemented at the offload card to support low-latency processing of messages. That is, by utilizing the offload capability of the offload card, the virtual network card with I / O pass-through has the ability to independently process messages and DMA data without the participation of the CPU and VMM. Messages do not need to be stored at the destination VMM, but can be processed in real time at the offload card and forwarded to external devices in a timely manner according to the tunnel table entries. The virtual network card with I / O pass-through of the source device supports the replacement and forwarding of the message tunnel header. There is no need to store messages, nor to wait for the virtual machine to be fully started before reading data from the destination VMM.

[0187] (3) The traffic relay process for lossless hot migration in high-performance network scenarios improves its applicability in high-bandwidth scenarios while retaining the network performance advantages of the hot migration process. To further expand the applicable scenarios of lossless hot migration, a complete traffic relay process is designed under the lossless tunnel relay mechanism. The operation mechanism of the tunnel relay solution is described in various possible stages before the virtual machine is started, while retaining the network performance advantages of the hot migration process. This process decouples the target VMM from participating in storage relay and DMA.

[0188] It is understood that in order to implement the functions in the above embodiments, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0189] The virtual machine migration method provided by the present application is described in detail above with reference to FIG. 1 to FIG. 6 . The data processing device provided by the present application will be described below with reference to FIG. 7 .

[0190] Figure 7 is a schematic diagram of the structure of a possible data processing device provided by this application. These data processing devices can be used to implement the function of uninstalling the card in the source device in the above-mentioned method embodiment, thereby also achieving the beneficial effects of the above-mentioned method embodiment. In this embodiment, the data processing device can be a computer device as shown in Figure 2, or a module (such as a chip) applied to a server.

[0191] As shown in Figure 7, a data processing device 700 includes a communication module 710, a processing module 720, and a storage module 730. The data processing device 700 is used to implement the function of uninstalling the card in the source device in the method embodiment shown in Figure 4 above.

[0192] The communication module 710 is configured to receive a first service message sent from an external device to a virtual machine in a source device, where the first service message includes first service data.

[0193] The processing module 720 is configured to transmit the first service data to the offload card in the destination device based on the first tunnel, where the first tunnel provides a path for transmitting data between the source device and the destination device. For example, the processing module 720 is configured to execute step 420 in FIG. 4 .

[0194] Optionally, the processing module 720 is further configured to modify the first service message according to the first tunnel entry to obtain a first transit service message.

[0195] Optionally, the processing module 720 is configured to modify the first transit response message according to the second tunnel entry to obtain a first response message. For example, the processing module 720 is configured to execute step 450 in FIG. 4 .

[0196] Optionally, the communication module 710 is configured to transmit the first response message to the external device based on the second tunnel. For example, the communication module 710 is configured to execute step 460 in FIG4 .

[0197] The storage module 730 is configured to store tunnel entries.

[0198] It should be understood that the data processing device 700 of the embodiment of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), and the above-mentioned PLD can be a complex programmable logical device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), a DPU, an accelerator card, an offload card, or any combination thereof. When the method shown in Figure 4 is implemented by software, its various modules can also be software modules, and the data processing device 700 and its various modules can also be software modules.

[0199] According to the data processing device 700 of the embodiment of the present application, it can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each unit in the data processing device 700 are respectively for implementing the corresponding processes of each method in Figure 4. For the sake of brevity, they are not repeated here.

[0200] FIG8 is a schematic diagram of the structure of a computer device 800 provided in this application. As shown in FIG8 , computer device 800 includes a processor 810, a bus 820, a storage medium 830, a communication interface 840, a memory 850 (also referred to as a main memory unit), and an offload card 860. Processor 810, offload card 860, storage medium 830, memory 850, and communication interface 840 are connected via bus 820.

[0201] It should be understood that in this embodiment, the processor 810 may be a CPU, but may also be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0202] The computer device 800 may also include a GPU, an NPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application. For example, the offload card 860 may be a DPU.

[0203] The communication interface 840 is used to implement communication between the computer device 800 and external devices or components.

[0204] In this application, when the computer device 800 is used to implement the function of the offload card of the source device shown in Figure 4, the offload card 860 is used to receive messages, so that the offload card 860 is used to forward messages according to the tunnel table entries stored in the memory 850.

[0205] The bus 820 may include a path for transmitting information between the aforementioned components (e.g., the processor 810, the memory 850, and the storage medium 830). In addition to the data bus, the bus 820 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 820 in the figure. The bus 820 may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (UBus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 820 may be divided into an address bus, a data bus, a control bus, etc.

[0206] As an example, computer device 800 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions).

[0207] It is worth noting that FIG8 only uses the example of a computer device 800 including one processor 810 and one offload card 860. Here, the processor 810 and the offload card 860 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined based on business requirements. For example, the computer device 800 may include multiple offload cards.

[0208] Memory 850 may be a volatile memory pool or a non-volatile memory pool, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM (DR RAM). Memory 850 is used to store tunnel table entries, etc.

[0209] The storage medium 830 may correspond to the storage medium used to store data and other information of the virtual machine in the above method embodiment, for example, a disk such as a solid state drive.

[0210] The computer device 800 may be a general-purpose device or a dedicated device. For example, the computer device 800 may be a server or other device with computing capabilities.

[0211] It should be understood that the computer device 800 according to this embodiment may correspond to the data processing device 700 in this embodiment, the unloading card 860 is used to implement the functions of the communication module 710 and the processing module 720 in the data processing device 700, and may correspond to the execution of the corresponding subject in any method in Figure 4, Figure 5 or Figure 6, and the above-mentioned and other operations and / or functions of the communication module 710 and the processing module 720 in the data processing device 700 are respectively for implementing the corresponding processes of each method in Figure 4, Figure 5 or Figure 6, which will not be repeated here for the sake of brevity.

[0212] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in RAM, flash memory, ROM, PROM, EPROM, EEPROM, registers, hard disk, mobile hard disk, CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist in a computing device as discrete components.

[0213] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the embodiment of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid-state drive (SSD). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A virtual machine migration method, characterized in that: include: When a virtual machine is migrated from a first computer device to a second computer device, the first computer device receives a first service message sent by a third computer device to the virtual machine in the first computer device, the first service message including first service data, and the third computer device includes a computer device other than the first computer device and the second computer device; The offload card in the first computer device transmits the first service data to the offload card in the second computer device based on a first tunnel, and the first tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the second computer device.

2. The method according to claim 1, characterized in that The offload card in the first computer device transmits the first service data to the offload card in the second computer device based on the first tunnel, including: The offload card in the first computer device modifies the first service message according to the first tunnel entry to obtain a first transit service message, wherein the message header of the first transit service message is different from the message header of the first service message, the first transit service message includes the first service data, and the first tunnel entry is used to indicate a transmission rule for transmitting data between the offload card in the first computer device and the offload card in the second computer device; The offload card in the first computer device transmits the first transit service message to the offload card in the second computer device based on the first tunnel.

3. The method according to claim 2, characterized in that The source address in the first service message indicates the address of the third computer device, and the destination address in the first service message indicates the address of the first computer device; The offload card in the first computer device modifies the first service message according to the first tunnel entry to obtain a first transit service message, including: The source address in the first service message is modified to the address of the first computer device, and the destination address in the first service message is modified to the address of the second computer device, and the virtual extended local area network identifier VNI in the first service message is modified to the first VNI to obtain the first transit service message. The first virtual extended local area network indicated by the first VNI includes the first computer device and the second computer device.

4. The method according to claim 2 or 3, characterized in that The first service message also includes an identifier of the virtual machine; and the method further includes: The offload card in the first computer device obtains the first tunnel table entry according to the identifier of the virtual machine and the first VNI.

5. The method according to any one of claims 1 to 4, characterized in that Before the offload card in the first computer device transmits the first service data to the offload card in the second computer device based on the first tunnel, the method further includes: An offload card in the first computer device establishes the first tunnel.

6. The method according to any one of claims 1 to 5, characterized in that After the offload card in the first computer device transmits the first service data to the offload card in the second computer device based on the first tunnel, the method further includes: The offload card in the first computer device receives, based on the first tunnel, a first transfer response message sent by the offload card in the second computer device, where the first transfer response message includes first response data, which is result data after the offload card in the second computer device processes the first service data. The offload card in the first computer device modifies the first transfer response message according to the second tunnel entry to obtain a first response message, wherein a message header of the first transfer response message is different from a message header of the first response message, the first response message includes the first response data, and the second tunnel entry is used to indicate a transmission rule for transmitting data between the offload card in the first computer device and the offload card in the third computer device; The offload card in the first computer device transmits the first response message to the third computer device based on a second tunnel, and the second tunnel is used to provide a path for transmitting data between the offload card in the first computer device and the offload card in the third computer device.

7. The method according to claim 6, characterized in that The source address in the first transfer response message indicates the address of the second computer device, and the destination address in the first transfer response message indicates the address of the first computer device; The offload card in the first computer device modifies the first transfer response message according to the second tunnel table entry to obtain a first response message, including: The source address in the first transit response message is modified to the address of the first computer device, and the destination address in the first transit response message is modified to the address of the third computer device, and the VNI in the first transit response message is modified to the second VNI to obtain the first transit service message, and the second virtual extended local area network indicated by the second VNI includes the first computer device and the third computer device.

8. The method according to claim 6 or 7, characterized in that The first transfer response message further includes an identifier of the virtual machine; and the method further includes: The offload card in the first computer device obtains the second tunnel table entry according to the identifier of the virtual machine and the second VNI.

9. The method according to claim 8, characterized in that The method further comprises: The uninstall card of the first computer device deletes the first tunnel and the second tunnel.

10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: When the virtual machine is migrated from the second computer device to the first computer device, the offload card in the first computer device receives a first transit service message based on the first tunnel, where the first transit service message includes first service data sent from the third computer device to the virtual machine in the second computer device, and the first tunnel is used to provide a path for data transmission between the offload card in the first computer device and the offload card in the second computer device; The offload card in the first computer device processes the first service data to obtain first response data; When the offload card in the first computer device and the offload card in the third computer device have not established a third tunnel, the offload card in the first computer device sends a first transfer response message to the offload card in the second computer device based on the first tunnel, where the first transfer response message includes the first response data, a source address in the first transfer response message indicates an address of the first computer device, and a destination address indicates an address of the second computer device, and the third tunnel is used to provide a path for data transmission between the offload card in the second computer device and the offload card in the third computer device; When the offload card in the first computer device establishes the third tunnel with the offload card in the third computer device, the offload card in the first computer device sends a first response message to the offload card in the third computer device based on the second tunnel, where the first response message includes the first response data.

11. A data processing device, characterized in that: When a virtual machine is migrated from a first computer device to a second computer device, the apparatus includes: a communication module, configured to receive a first service message sent by a third computer device to the virtual machine in the first computer device, wherein the first service message includes first service data; The processing module is configured to transmit the first service data to the offload card in the second computer device based on a first tunnel, wherein the first tunnel is configured to provide a path for transmitting data between the first computer device and the second computer device.

12. The device according to claim 11, characterized in that The processing module is specifically configured to modify the first service message according to the first tunnel entry to obtain a first transit service message, wherein the message header of the first transit service message is different from the message header of the first service message, the first transit service message includes the first service data, and the first tunnel entry is used to indicate a transmission rule for transmitting data between the offload card in the first computer device and the offload card in the second computer device; The communication module is specifically configured to transmit the first transit service message to the offload card in the second computer device based on the first tunnel.

13. The device according to claim 12, characterized in that The source address in the first service message indicates the address of the third computer device, and the destination address in the first service message indicates the address of the first computer device; The processing module is specifically used to modify the source address in the first business message to the address of the first computer device, and to modify the destination address in the first business message to the address of the second computer device, and to modify the virtual extended local area network identifier VNI in the first business message to the first VNI, to obtain the first transit business message, and the first virtual extended local area network indicated by the first VNI includes the first computer device and the second computer device.

14. The device according to any one of claims 11 to 13, characterized in that The communication module is further configured to receive, based on the first tunnel, a first transfer response message sent by the offload card in the second computer device, wherein the first transfer response message includes first response data, and the first response data is result data after the offload card in the second computer device processes the first service data; The processing module is further configured to modify the first transfer response message according to the second tunnel entry to obtain a first response message, wherein a message header of the first transfer response message is different from a message header of the first response message, the first response message includes the first response data, and the second tunnel entry is configured to indicate a transmission rule for transmitting data between the offload card in the first computer device and the offload card in the third computer device; The communication module is further configured to transmit the first response message to the third computer device based on a second tunnel, wherein the second tunnel is configured to provide a path for transmitting data between the offload card in the first computer device and the offload card in the third computer device.

15. An uninstall card, characterized in that: include: A processor and a power supply circuit; wherein the power supply circuit is used to supply power to the processor; The processor is configured to execute the operation steps of the method according to any one of claims 1 to 10.

16. A computer device, characterized in that: The computer device includes a processor and an offload card, the processor is used to run a virtual machine; when the virtual machine is migrated to another computer device, the offload card is used to execute the operation steps of the method described in any one of claims 1 to 10 above, thereby relaying and forwarding the virtual machine's messages after the virtual machine is migrated.

17. A communication system, characterized in that: The communication system comprises at least two computer devices, wherein the computer devices comprise the offload card according to claim 15 , and the offload card is configured to execute the operation steps of the method according to any one of claims 1 to 10 .

Citation Information

Patent Citations

  • Method, monitor and system for seamless migration of virtual machine

    CN102185774A

  • Virtual machine migration method, device and system

    CN114691287A

  • Thermal migration method and device of virtual machine, electronic equipment and storage medium

    CN115309503A

  • Tunnel-based service insertion in public cloud environments

    US20200236046A1

Cited By

  • Fttr network roaming method, main gateway, system, storage medium and program product

    CN122496882A