An IP Consistency Method for K3s Service Migration
By implementing synchronous migration of service status and IP addresses in the K3s environment, combined with CRIU technology and Flannel plug-in, the problems of insufficient client transparency, performance delay and fault tolerance in the existing technology are solved, and an efficient and transparent service migration process is achieved.
Patent Information
- Application Number
- CN202411769758.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-04
AI Technical Summary
When the prior art realizes the transparency of K3s service migration to clients, there are problems of insufficient performance delay and failure tolerance, and the introduction of intermediate layers increases system complexity and resource utilization.
An IP consistency method for K3s service migration is proposed. By setting corresponding modules in the source and target nodes, synchronous migration of service status and IP address is realized, combined with CRIU technology, the checkpoint file is generated and restored, and the consistency of service status is ensured, and the IP allocation and routing configuration is modified through Flannel and host-local plug-ins to keep the actual IP address of the Pod unchanged.
It realizes the consistency of service status during service migration and the user's transparent service address maintenance, improves user experience, reduces system complexity and resource usage, and enhances fault tolerance.
Smart Images

Figure CN119652895B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge computing based on K3s, and particularly relates to an IP consistency method for K3s service migration. Background Art
[0002] As an emerging computing model, edge computing deploys computing resources and services to the edge location close to the data source to reduce data transmission latency and network bandwidth pressure, and improve the system's response speed and user experience. In edge computing, K3s, as a lightweight Kubernetes distribution, is widely popular due to its easy deployment and low resource occupancy, enabling edge devices to efficiently manage containerized applications. However, due to the instability and resource constraints of the edge environment, the fault tolerance problem has become an important challenge in edge computing. In this context, service migration technology becomes crucial because it can achieve seamless migration of services in the face of node failures or resource shortages, ensuring the continuous availability and stability of the system, and thus guaranteeing the service quality of users.
[0003] Currently, there have been research works on the migration of containerized services. For example, Docker and CRIU technologies are used to generate checkpoints for containers and then restore their states. These methods can save the running state of containers during the migration process, but they cannot ensure that the service IP addresses before and after migration remain unchanged, which will cause the client connection to be interrupted or require reconnection, and cannot achieve a seamless migration that is transparent to users.
[0004] In the existing edge computing environment based on K3s, the following measures are usually taken to ensure that service migration is transparent to the client:
[0005] 1. Service proxy: Place an intermediate layer between the client and the service, responsible for receiving requests and forwarding them to the appropriate service instance. The service proxy can handle load balancing and failover of traffic. The client always accesses the proxy intermediate layer to request services, ensuring that service migration is invisible to users.
[0006] 2. K3s built-in resource - service: In K3s, a service is an abstract resource used to define the access policy for a group of Pods. It provides a stable way to access application instances running in Pods, solving the problem that the IP address of a Pod may change due to migration or scheduling to different nodes.
[0007] 3. DNS update: By dynamically updating DNS records, ensure that after service migration, the client can access the service through a unified service name. This method allows the actual address of the service to change without modifying the client code, as long as the DNS resolution correctly points to the new service address.
[0008] Although existing technologies provide certain solutions for achieving the transparency of service migration to clients, their main focus is on service discovery through the introduction of intermediate layers such as service proxies, service resources, or DNS resolution, rather than ensuring that the actual IP addresses of Pods remain unchanged before and after service migration. Such technologies have significant drawbacks: First, since requests need to go through additional forwarding and resolution steps, it will increase the latency of service responses and affect the user experience. Second, the introduction of intermediate layers increases the complexity of the system and may lead to more single points of failure. For example, if the service proxy or DNS service fails, it will directly affect the client's access to the service and increase the vulnerability of the system. Finally, intermediate layers usually consume additional computing and network resources, which may lead to performance bottlenecks in resource-constrained edge computing environments and affect system efficiency. Therefore, although existing technologies have solved the transparency problem of service migration to a certain extent, there is still room for improvement in terms of performance and fault tolerance. Summary of the Invention
[0009] To solve the problems in terms of performance and fault tolerance brought by existing technologies for achieving the transparency of service migration to clients, the present invention proposes an IP consistency method for K3s service migration.
[0010] To achieve the above object, the present invention is implemented through the following technical solutions:
[0011] An IP consistency method for K3s service migration includes the following steps:
[0012] S1. Set up a source node service status migration module, a source node IP migration module, and a source node network status migration module on the source node, and set up a target node service status migration module, a target node IP migration module, and a target node network status migration module on the target node;
[0013] S2. When the client requests the source node service status migration module to perform service status migration, the source node service status migration module synchronizes the request information to the target node service status migration module, and the source node network status migration module modifies the network configuration;
[0014] S3. The source node IP migration module obtains the IP information of the source service Pod and writes it into the IP cache file, and NFS synchronizes the IP cache file to the target node IP migration module;
[0015] S4. The target node network plugin Flannel calls the target node IP migration module, reads the IP cache file, and then the target node network plugin Flannel assigns the read IP to the target service Pod;
[0016] S5. The target node service status migration module uses the checkpoint file to restore the status of the target service Pod. At the same time, the target node network status migration module modifies the network configuration of the target node to ensure that the migrated target service Pod can communicate with the resources normally.
[0017] Further, the specific implementation method of the service status migration module in step S2 is as follows:
[0018] S2.1. Generate a status checkpoint on the source node:
[0019] On the source node, represent the container status as:
[0020] S src ={M, P, FD, NC}
[0021] Where M represents the memory status, P represents the process information, FD represents the file descriptor, and NC represents the network connection;
[0022] Call the process checkpoint and recovery tool CRIU to generate a status checkpoint. The expression is:
[0023] C checkpoint =CRIU(S src )→Save as a checkpoint file;
[0024] Where C checkpoint is the checkpoint file of the container;
[0025] The checkpoint file includes the memory status, process information, file descriptor, and network connection data of the container;
[0026] S2.2. Transfer the checkpoint file. Transfer the checkpoint file to the specified path on the target node through the NFS network. The expression is:
[0027] C checkpoint →C checkpoint_target
[0028] Where C checkpoint_target is the checkpoint file of the target node;
[0029] S2.3. Container status recovery: On the target node, the process of Docker starting the container instance and reconstructing the status is;
[0030] Initialize the target status S target :
[0031] S target =Docker_start(C checkpoint_target )
[0032] Use CRIU to read the checkpoint file:
[0033] S target = CRIU_restore(C checkpoint_target ) -> S target = {M, P, FD, NC}
[0034] At the target node, Docker starts the corresponding container instance and uses CRIU to read the checkpoint file to reconstruct the container's state and restore it to the running state before migration;
[0035] S2.4. Complete service state migration: The service state is migrated from the source node to the target node, and the migrated container service state S migrated , the expression is:
[0036] S migrated = S target -> S src
[0037] Finally, the service state of the target node is consistent with that of the source node.
[0038] Furthermore, in step S3, the IP migration module modifies the IP allocation mechanism of the host-local plugin, and sets the function for host-local to call for IP allocation as the GetIter() function. The specific IP allocation process is as follows:
[0039] S3.1. Obtain an iterator: First, call the GetIter() method to obtain an iterator pointing to the IP address range. This iterator will locate to the previously allocated IP address as the starting point for the next allocation;
[0040] S3.2. Check the previously allocated IP address: In GetIter(), the previously allocated IP address will be checked first. If it exists, the iterator will point to this IP address, preparing to allocate the next IP address;
[0041] S3.3. Allocate an IP address: When a request to allocate a new IP address is made, call the Next() method to obtain the next available IP address; if the currently pointed IP address is empty, allocate the starting IP address of the range; if the end of the range has been reached, the iterator will be reset and loop back to the start of the address pool to implement the round-robin strategy;
[0042] During the allocation process, if the current IP address is the subnet gateway or an occupied IP address, the iterator will call its own Next() method to skip the subnet gateway or the occupied IP address until an available IP address is found;
[0043] S3.4. Return result: After finding an available IP address, return a structure containing the available IP address and its corresponding gateway. If no available IP address is found within the entire range, return an error message indicating that there is no available address.
[0044] Furthermore, the IP address allocated last time in step S3.2 is defaultly stored in / var / lib / cni / networks / last_reserved_ip.0.
[0045] Furthermore, the specific implementation method of step S4 is achieved through the host-local plugin of Flannel, including the following steps:
[0046] S4.1. IP pool management: Flannel allocates an IP address pool for each node. The IP address pool is dynamically generated by Flannel according to the CIDR range of the cluster, and each node maintains a local IP address manager.
[0047] S4.2. Allocate an IP address when a Pod starts: When a Pod starts, Flannel calls the host-local plugin to allocate an available IP address from the node's IP address pool. This address is assigned to the Pod and configured through Flannel's network interface. The process of calculating the IP address for the i-th Pod within the CIDR range of the node is as follows:
[0048] IP Pod = CIDR Node + i, i ∈ {1, 2,..., n}
[0049] where, IP Pod is the specific IP address assigned to the Pod, CIDR Node is the CIDR range of the current node, and i is the offset assigned to the i-th Pod.
[0050] S4.3. Set the IP address assigned to the Pod to remain unchanged during the Pod's life cycle until the Pod is deleted or restarted; when the Pod restarts, Flannel reuses the previously assigned IP address to ensure the persistence of the IP address.
[0051] S4.4. Network configuration: Flannel is responsible for setting up routing rules to enable correct routing of traffic inside and outside the cluster, ensuring the communication between Pods and the ability to access external networks.
[0052] Furthermore, in step S5, the target node network status migration module modifies the network configuration of the target node, including IP allocation and modifying the routing table.
[0053] S5.1. IP Assignment: Modify the host-local network plugin to use the IP of the source service Pod to modify the IP of the target service Pod, ensuring that the actual IP address of the service Pod remains unchanged before and after migration;
[0054] S5.2. Modify the routing table: Mark the route of the target Pod in the routing table of the target node to ensure that traffic can reach the new Pod location correctly. The process is described as follows:
[0055] Let R dest be the routing table of the target node, IP migrated be the IP address of the Pod after migration, and local be the local address. The modified routing table rule is:
[0056] R dest (IP migrated ) = local
[0057] means accessing IP migrated directly locally without going through the gateway.
[0058] Furthermore, the source node network configuration of the source node network status migration module includes IP blocking, modifying the routing table, and modifying Iptables.
[0059] Furthermore, the IP blocking is to lock the IP of the source Pod at the beginning of service migration to prevent other new Pods from requesting to use this IP. The process is described as follows:
[0060] Let IP src be the IP address of the source node Pod, t0 be the start time of migration, and t end be the end time of migration. Then, within the time interval [t0, t end , use the following condition to lock the IP:
[0061] Lock(IP src ) = (t0 ≤ t && t ≤ t end )? 1:0
[0062] where Lock(IP src ) = 1 means the IP is locked and cannot be used by other Pods.
[0063] Furthermore, the process of the route modification is as follows:
[0064] Let R src be the routing table of the source node, IP migrated be the IP address of the Pod after migration, and Node dest be the target node. Then the route modification rules are as follows:
[0065] R src (IP migrated ) = Node dest
[0066] Indicates that when accessing the IP on the source node migrated the data packet will be routed to the destination node Node dest .
[0067] Furthermore, the process of modifying the Iptables is described as follows:
[0068] Let IP src be the IP of the source node, IP migrated be the IP address of the Pod after migration, and IP nat be the IP after masquerading. Then when the data packet is sent from the source node to the destination Pod, the masquerading rule is:
[0069] SNAT(IP src ) → IP nat → IP migrated
[0070] Indicates changing the source address IP src to the IP after masquerading IP nat to ensure that the traffic can be correctly routed to the IP address IP migrated of the Pod after migration.
[0071] Advantages of the present invention:
[0072] An IP consistency method for K3s service migration according to the present invention combines the CRIU technology to realize the generation and recovery of the service state checkpoint, ensuring the consistency of the service state during the migration process and further improving the user experience. The present invention modifies the Flannel and host-local network plugins to ensure that the actual IP address of the Pod remains unchanged during the service migration process, avoiding the influence of introducing an intermediate layer in the prior art and ensuring the transparency of the service address for the user during the service migration process. During the implementation process, the IP locking mechanism of the source node effectively prevents the new Pod from occupying the IP of the source Pod during the migration, avoiding IP address conflicts. By dynamically adjusting the network state, it can ensure that the Pod after migration can communicate normally with other resources and continue to provide services for the user. Brief Description of the Drawings
[0073] Figure 1 is a flowchart of an IP consistency method for K3s service migration according to the present invention;
[0074] Figure 2 is a schematic diagram of the basic process of IP migration;
[0075] Figure 3 Overall architecture diagram for service IP migration;
[0076] Figure 4 IP allocation process diagram for the Get() function;
[0077] Figure 5 Process diagram of the modified Get() function;
[0078] Figure 6 Test result diagram of the service address of the present invention;
[0079] Figure 7 Communication test result diagram of the present invention. Detailed implementation manners
[0080] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described specific implementation manners are only a part of the implementation manners of the present invention, rather than all the specific implementation manners. The components of the specific implementation manners of the present invention usually described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other implementation manners.
[0081] Therefore, the detailed description of the specific implementation manners of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed present invention, but only represents the selected specific implementation manners of the present invention. All other specific implementation manners obtained by those skilled in the art based on the specific implementation manners of the present invention without creative efforts belong to the scope of protection of the present invention.
[0082] To further understand the content, features and effects of the present invention, the following specific implementation manners are exemplified and are accompanied by the attached Figure 1 - Attached Figure 7 The details are described as follows: Specific implementation manner one:
[0084] An IP consistency method for K3s service migration, including the following steps:
[0085] S1. Set up a source node service status migration module, a source node IP migration module and a source node network status migration module on the source node, and set up a target node service status migration module, a target node IP migration module and a target node network status migration module on the target node;
[0086] S2. When the client requests the source node service status migration module to perform service status migration, the source node service status migration module synchronizes the request information to the target node service status migration module, and the source node network status migration module modifies the network configuration;
[0087] Further, the specific implementation method of the service status migration module in step S2 is as follows:
[0088] S2.1. Generate a status checkpoint on the source node:
[0089] On the source node, represent the container status as:
[0090] S src ={M, P, FD, NC}
[0091] Where M represents the memory status, P represents the process information, FD represents the file descriptor, and NC represents the network connection;
[0092] Call the process checkpoint and recovery tool CRIU to generate a status checkpoint, and the expression is:
[0093] C checkpoint =CRIU(S src ) → Save as a checkpoint file;
[0094] Where C checkpoint is the checkpoint file of the container, which contains the complete status of the container on the source node, including memory, process, file descriptor, and network connection;
[0095] The checkpoint file includes the memory status, process information, file descriptor, and network connection data of the container;
[0096] First, on the source node, the Docker controller is responsible for stopping the running container and calling CRIU to generate a checkpoint of the container. CRIU will save the data such as the memory status, process information, file descriptor, and network connection of the container as a set of image files to ensure that the complete status of the container is recorded.
[0097] S2.2. Transmit the checkpoint file and transfer the checkpoint file to the specified path on the target node through the NFS network. The expression is:
[0098] C checkpoint →C checkpoint_target
[0099] Where C checkpoint_target is the checkpoint file of the target node, that is, the checkpoint file transferred to the target node through the network, which contains all the status information required for container recovery during the migration process;
[0100] S2.3. Container state recovery: On the target node, the process of Docker starting the container instance and reconstructing the state is as follows;
[0101] Initialize the target state S target :
[0102] S target = Docker_start(C checkpoint_target )
[0103] Use CRIU to read the checkpoint file:
[0104] S target = CRIU_restore(C checkpoint_target ) -> S target = {M, P, FD, NC}
[0105] On the target node, Docker starts the corresponding container instance and uses CRIU to read the checkpoint file to reconstruct the container state and restore it to the running state before migration;
[0106] S2.4. Complete service state migration: The service state is migrated from the source node to the target node to obtain the migrated container service state S migrated , that is, after the migration is completed, the container state on the target node has been restored to the original state of the source node, and the expression is:
[0107] S migrated = S target -> S src
[0108] Finally, the service state of the target node is consistent with that of the source node.
[0109] S3. The IP migration module of the source node obtains the IP information of the source service Pod and writes it into the IP cache file, and NFS synchronizes the IP cache file to the IP migration module of the target node;
[0110] In the network architecture of K3s, Flannel, as a container network interface (CNI) plugin, is responsible for providing network connections and IP assignments for Pods. The working mode of Flannel is to set up a virtual network in the cluster so that each Pod can communicate in a flat network space.
[0111] Furthermore, in step S3, the IP migration module modifies the IP assignment mechanism of the host-local plugin and sets the function called by host-local for IP assignment to the GetIter() function, that is, the available IP address iteration function. The specific IP assignment process is as follows:
[0112] S3.1. Obtain an iterator: First, call the GetIter() method to obtain an iterator pointing to the IP address range. This iterator will locate the previously allocated IP address as the starting point for the next allocation.
[0113] S3.2. Check the previously allocated IP address: In GetIter(), first check the IP address allocated last time. If it exists, the iterator will point to this IP address, preparing to allocate the next IP address.
[0114] S3.3. Allocate an IP address: When a request is made to allocate a new IP address, call the Next() method to obtain the next available IP address. If the currently pointed IP address is empty, allocate the starting IP address of the range. If the end of the range has been reached, the iterator will be reset and loop back to the start of the address pool to implement the round-robin strategy.
[0115] During the allocation process, if the current IP address is the subnet gateway or an occupied IP address, the iterator will call its own Next() method to skip the subnet gateway or the occupied IP address until an available IP address is found.
[0116] S3.4. Return the result: When an available IP address is found, return a structure containing the available IP address and its corresponding gateway. If no available IP address can be found throughout the range, return an error message indicating that there are no available addresses.
[0117] In this way, the round-robin algorithm can effectively allocate IPs in the IP address pool, ensuring fair use of resources while reducing the possibility of conflicts and duplicate allocations.
[0118] The modification process for fixed IPs is mainly reflected in the handling of migrated IPs and the adjustment of IP allocation logic. The modified logic is as Figure 4 shown. Generate the corresponding file path according to the Pod name and check if there is a file saving the source Pod IP under this path. If the IP file exists, read the IP address in it and allocate it to the current Pod. If the source Pod IP file does not exist, enter the regular IP allocation process.
[0119] S4. The Flannel network plugin on the target node calls the IP migration module on the target node, reads the IP cache file, and then the Flannel network plugin on the target node allocates the read IP to the target service Pod.
[0120] Furthermore, the specific implementation method of step S4 is implemented through the host-local plugin of Flannel and includes the following steps:
[0121] S4.1. IP Pool Management: Flannel assigns an IP address pool to each node. The IP address pool is dynamically generated by Flannel based on the CIDR range of the cluster. Each node maintains a local IP address manager;
[0122] S4.2. Assigning IP Addresses When a Pod Starts: When a Pod starts, Flannel calls the host-local plugin to allocate an available IP address from the node's IP address pool. This address is assigned to the Pod and configured through Flannel's network interface. Within the node's CIDR range, the process of calculating the IP address for the i-th Pod is as follows:
[0123] IP Pod = CIDR Node + i, i ∈ {1, 2,..., n}
[0124] Where, IP Pod is the specific IP address assigned to the Pod, CIDR Node is the CIDR range of the current node, and i is the offset assigned to the i-th Pod;
[0125] S4.3. Ensure that the IP address assigned to the Pod remains unchanged during the Pod's lifecycle until the Pod is deleted or restarted; when the Pod restarts, Flannel re-uses the previously assigned IP address to ensure the persistence of the IP address;
[0126] S4.4. Network Configuration: Flannel is responsible for setting up routing rules to enable proper routing of traffic within and outside the cluster, ensuring communication between Pods and the ability to access external networks.
[0127] S5. The target node service status migration module uses checkpoint files to restore the status of the target service Pod. Meanwhile, the target node network status migration module modifies the network configuration of the target node to ensure normal communication between the migrated target service Pod and resources.
[0128] Furthermore, in step S5, the target node network status migration module modifies the network configuration of the target node, including IP assignment and modifying the routing table;
[0129] S5.1. IP Assignment: Modify the host-local network plugin to use the IP of the source service Pod to modify the IP of the target service Pod, ensuring that the actual IP address of the service Pod remains unchanged before and after migration;
[0130] S5.2. Modify the routing table: Mark the route of the target Pod in the routing table of the target node to ensure that traffic can reach the new Pod location correctly. The process is described as follows:
[0131] Let R dest be the routing table of the target node, IP migrated be the IP address of the Pod after migration, and local be the local address. The modified routing table rule is:
[0132] R dest (IP migrated ) = local
[0133] It means that when accessing IP migrated on the target node, access directly locally without passing through the gateway.
[0134] Furthermore, the source node network configuration of the source node network status migration module includes IP blocking, modifying the routing table, and modifying Iptables.
[0135] When performing service migration, in order to prevent users from deploying new service Pods during the migration process from occupying the IP address of the source Pod, an IP locking mechanism is implemented.
[0136] The IP blocking is to lock the IP of the source Pod at the beginning of the service migration to prevent other new Pods from requesting to use this IP. The process is described as follows:
[0137] Let IP src be the IP address of the source node Pod, t0 be the start time of the migration, and t end be the end time of the migration. Then, within the time interval [t0, t end , use the following condition to lock the IP:
[0138] Lock(IP src ) = (t0 ≤ t && t ≤ t end )? 1:0
[0139] Among them, when Lock(IP src ) = 1, it means that the IP is locked and cannot be used by other Pods.
[0140] Communication between source Pods on the source node does not require going through the gateway, and they can directly access each other. However, when these Pods are migrated to the target node, in order to ensure correct access to these migrated Pods, corresponding markings need to be made in the network routing. This means that when accessing the IP address of the migrated Pod, the network router needs to know that these requests should be redirected to the target node instead of the source node. In this way, it can be ensured that even if the physical location of the Pod changes, the source node can still access these migrated Pods through the correct routing.
[0141] The process of modifying the routing is as follows:
[0142] Let R src be the routing table of the source node, IP migrated be the IP address of the migrated Pod, and Node dest be the target node. Then the rules for modifying the routing are as follows:
[0143] R src (IP migrated ) = Node dest
[0144] It means that when accessing IP migrated on the source node, the data packet will be routed to the target node Node dest .
[0145] Since accessing source Pods on the source node does not require address masquerading through Iptables, but when the source Pods are migrated to the target node, address masquerading is required when accessing these Pods again to ensure that the traffic can be correctly routed to the new location. Therefore, it is necessary to mark these migrated Pods in Iptables to avoid the situation where address masquerading is not performed during access, thus ensuring that the target node can correctly identify the data packets accessing the migrated IP address. This mechanism ensures that the traffic is correctly processed.
[0146] The process of modifying Iptables is described as follows:
[0147] Let IP src be the IP of the source node, IP migrated be the IP address of the migrated Pod, and IP nat be the IP after masquerading. Then when the data packet is sent from the source node to the target Pod, the masquerading rule is:
[0148] SNAT(IP src ) → IP nat → IP migrated
[0149] It means that the source address IP srcChange to the IP address after disguise nat , to ensure that traffic can be correctly routed to the IP address of the Pod after migration migrated .
[0150] The experimental results are verified as follows:
[0151] The service IP migration technology under K3s was tested. Considering that the master node does not run business containers, three virtual machines with the same software and hardware configuration were selected as cluster nodes for the experiment. One was the K3s master node, and the other two were K3s worker nodes and served as the source node and target node for service migration. The hardware configuration of the virtual machine nodes is shown in Table 1:
[0152] Table 1 Hardware configuration of virtual machine nodes
[0153]
[0154] In addition to the hardware configuration of the nodes, software tools such as K3s and Docker are also required to build the test environment. The selected software and its versions are shown in Table 2:
[0155] Table 2 Software tools used in the test environment
[0156]
[0157] The network service was selected as the test case in the test. The functional test of service IP migration was carried out in the above test environment, which can better test the network consistency. The client periodically (1ms) sends UDP packets containing time information to port 12345 of the container. The server running in the container listens on port 12345 to receive the packets and parse them, and then writes the parsed results into a file. The network service was deployed in K3s. The client periodically sends packets to simulate user requests. The network state consistency before and after migration was verified by checking whether the external address of the service changes before and after migration and whether the target Pod can communicate normally after migration.
[0158] The test results of service address consistency are as Figure 6 shown. The network service was successfully migrated from edge node 1 to edge node 2 after migration, and the actual IP of the Pod was always 10.240.2.7 before and after migration, realizing the transparency of the service address for users during service migration.
[0159] The test results of communication are as Figure 7 shown. Although the migrated Pod and edge node 2 are not in the same network segment, after the network state migration, it is ensured that the target Pod can communicate normally with other resources.
[0160] Abbreviation and definition of key terms
[0161] 1. K8s (Kubernetes): An open-source container orchestration platform designed to automate the deployment, scaling, and management of applications. It provides an efficient way to manage the lifecycle of containerized applications and supports large-scale microservices architectures.
[0162] 2. K3s (Lightweight Kubernetes): A lightweight distribution of K8s, designed for resource-constrained environments and edge computing, providing a simplified installation and running experience.
[0163] 3. DNS (Domain Name System): Provides the mapping from domain names to IP addresses.
[0164] 4. Pod: A unit of container collection, the smallest schedulable unit in K3s, usually sharing network and storage resources. Services are deployed in the form of Pods in the K3s cluster for client access.
[0165] 5. Docker: A containerization technology platform that allows developers to package an application and all its dependencies into a standardized unit (container), ensuring that the application can run consistently in any environment.
[0166] 6. Flannel: The default network plugin for K3s, used for container networking in the K3s cluster, providing a simple overlay network solution for intra-cluster communication.
[0167] 7. host-local: An IP address allocation mode of Flannel that allocates IP addresses according to the local configuration file of each node, suitable for small clusters or simple deployments.
[0168] 8. CRIU (Checkpoint / Restore In Userspace): A tool for process checkpointing and restoration. It allows a running process to be frozen, saved to disk, and restored at a later point in time, usually used for container migration and restoration.
[0169] 9. SNAT (Source Network Address Translation): A network address translation technology that converts the source IP address of a data packet to another IP address.
[0170] 10. NFS (Network File System): A protocol that allows computers to share files and directories over a network.
[0171] 11. Iptables: A user-space tool in the Linux kernel used to set, maintain, and check IP packet filtering rules in the Linux kernel.
[0172] 12. CIDR (Classless Inter-Domain Routing): A standard for allocating IP addresses and routing.
[0173] 13. UDP (User Datagram Protocol): A network communication protocol used to transmit data in an IP network.
[0174] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0175] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. An IP consistency method for K3s service migration, characterized in that: The steps include: S1. Set the source node service state migration module, the source node IP migration module and the source node network state migration module at the source node, set the target node service state migration module, the target node IP migration module and the target node network state migration module at the target node; S2. When the client requests the source node service state migration module to migrate the service state, the source node service state migration module synchronizes the request information to the target node service state migration module, and the source node network state migration module modifies the network configuration; S3. The source node IP migration module obtains the IP information of the source service Pod and writes it into the IP cache file. NFS synchronizes the IP cache file to the target node IP migration module. S4. The target node network plug-in Flannel calls the target node IP migration module to read the IP cache file, and then the target node network plug-in Flannel assigns the read IP to the target service Pod; S5. The target node service state migration module uses the checkpoint file to restore the state of the target service Pod. At the same time, the target node network state migration module modifies the network configuration of the target node to ensure that the target service Pod communicates normally with the resources after migration.
2. According to claim 1, an IP consistency method for K3s service migration is characterized in that: The specific implementation method of the service state migration module in step S2 is as follows: S2.
1. Generate a state checkpoint at the source node: On the source node, the container status is represented as: S src = {M,P,FD,NC} Among them, M represents memory status, P represents process information, FD represents file descriptor, and NC represents network connection; The process checkpoint and recovery tool CRIU is called to generate a state checkpoint. The expression is: C checkpoint = CRIU (S src ) → Save as checkpoint file; Among them, C checkpoint Checkpoint file for the container; The checkpoint file includes the memory status, process information, file descriptors and network connection data of the container; S2.
2. Transfer the checkpoint file. Transfer the checkpoint file to the specified path of the target node through the NFS network. The expression is: C checkpoint → C checkpoint_target Among them, C checkpoint_target Checkpoint file for the target node; S2.
3. Container state recovery: On the target node, Docker starts the container instance and rebuilds the state as follows; Initialize the target state S target : S target = Docker_start (C checkpoint_target ) Use CRIU to read the checkpoint file: S target = CRIU_restore (C checkpoint_target ) → S target = {M,P,FD,NC} On the target node, Docker starts the corresponding container instance and uses CRIU to read the checkpoint file to rebuild the state of the container and restore it to the running state before the migration; S2.
4. Complete service status migration: The service status is migrated from the source node to the target node, and the migrated container service status S is obtained. migrated , the expression is: S migrated = S target → S src Eventually, the service status of the target node is consistent with the service status of the source node.
3. According to claim 2, an IP consistency method for K3s service migration is characterized in that: In step S3, the IP migration module modifies the IP allocation mechanism of the host-local plug-in and sets the function called by host-local to allocate IP to the GetIter() function. The specific IP allocation process is as follows: S3.
1. Get iterator: First, call the GetIter() method to get an iterator pointing to the IP address range. The iterator will locate the last allocated IP address as the starting point for the next allocation. S3.
2. Check the last allocated IP address: In GetIter(), the last allocated IP address is checked first. If it exists, the iterator will point to it and prepare to allocate the next IP address. S3.
3. Allocate IP address: When a new IP address is requested, call the Next() method to obtain the next available IP address; if the currently pointed IP address is empty, the starting IP address of the allocation range is used; If the end of the range has been reached, the iterator will reset and loop back to the beginning of the address pool, implementing a round-robin strategy; During the allocation process, if the current IP address is a subnet gateway or an occupied IP address, the iterator will call its own Next() method to skip the subnet gateway or the occupied IP address until an available IP address is found; S3.
4. Return result: When an available IP address is found, a structure containing the available IP address and its corresponding gateway is returned. If no available IP address is found in the entire range, an error message is returned, indicating that there is no available address.
4. According to claim 3, an IP consistency method for K3s service migration is characterized in that: The last allocated IP address in step S3.2 is stored in / var / lib / cni / networks / last_reserved_ip.0 by default.
5. According to claim 4, an IP consistency method for K3s service migration is characterized in that: The specific implementation method of step S4 is implemented through the host-local plug-in of Flannel, including the following steps: S4.
1. IP pool management: Flannel allocates an IP address pool to each node. The IP address pool is dynamically generated by Flannel based on the CIDR range of the cluster. Each node maintains a local IP address manager. S4.
2. Allocating IP addresses when Pods start: When a Pod starts, Flannel calls the host-local plugin to allocate an available IP address from the node's IP address pool. This address is assigned to the Pod and configured through Flannel's network interface. The process of calculating the IP address for the i-th Pod within the node's CIDR range is as follows: IP Pod = CIDR Node + i , i ϵ {1,2,...,n} Among them, IP Pod The specific IP address assigned to the Pod, CIDR Node is the CIDR range of the current node, and i is the offset assigned to the i-th Pod; S4.
3. Set the IP address assigned to the Pod to remain unchanged during the life cycle of the Pod until the Pod is deleted or restarted; when the Pod is restarted, Flannel reuses the previously assigned IP address to ensure the continuity of the IP address; S4.
4. Network Configuration: Flannel is responsible for setting routing rules so that traffic can be correctly routed within and outside the cluster, ensuring communication between Pods and the ability to access external networks.
6. The IP consistency method for K3s service migration according to claim 5, characterized in that: In step S5, the target node network state migration module modifies the network configuration of the target node, including IP allocation and modification of the routing table; S5.
1. IP allocation: Modify the host-local network plug-in and use the IP of the source service Pod to modify the IP of the target service Pod to ensure that the actual IP address of the service Pod remains unchanged before and after the migration; S5.
2. Modify the routing table: Mark the route of the target Pod in the routing table of the target node to ensure that the traffic can correctly and directly reach the new Pod location. The process is described as follows: Assume R dest The routing table of the target node, IP migrated is the IP address of the Pod after migration, local is the local address, and the modified routing table rules are: R dest (IP migrated ) = local Indicates access to the IP on the target node migrated Directly access locally without going through a gateway.
7. The IP consistency method for K3s service migration according to claim 6 is characterized in that: The source node network configuration of the source node network state migration module includes IP blocking, modifying the routing table, and modifying Iptables.
8. The IP consistency method for K3s service migration according to claim 7, characterized in that: The IP blocking is to lock the IP of the source Pod at the beginning of the service migration to prevent other new Pods from requesting to use the IP. The process is described as follows: Set IP src is the IP address of the source node Pod, t0 is the time when the migration starts, and t end is the end time of migration, then in the time interval [t0, t end ], use the following conditions to lock the IP: Lock (IP src ) = (t0 ≤ t && t ≤ t end ) ? 1 : 0 Among them, Lock (IP src ) = 1, it means that the IP is locked and cannot be used by other Pods.
9. The IP consistency method for K3s service migration according to claim 8, characterized in that: The process of modifying the routing table by the source node is as follows: Assume R src is the routing table of the source node, IP migrated is the IP address of the Pod after migration, Node dest If it is the target node, the rules for route modification are as follows: R src (IP migrated ) = Node dest Indicates access to the IP address on the source node migrated When the data packet is routed to the target node Node dest .
10. The IP consistency method for K3s service migration according to claim 9, characterized in that: The Iptables modification process is described as follows: Set IP src is the IP address of the source node, IP migrated is the IP address of the Pod after migration, IP nat is the disguised IP, then when the data packet is sent from the source node to the target Pod, the disguise rule is: SNAT (IP src ) → IP nat → IP migrated Indicates that the source address IP src Change to the disguised IP address nat , to ensure that traffic can be correctly routed to the IP address of the migrated Pod migrated .
Citation Information
Patent Citations
Edge service migration method based on docker container
CN110351336A
UD service mode live migration method based on RDMA application container
CN118193131A