High Availability Network Address Translation
The method ensures low-latency network address translation by synchronizing NAT and TCP state tables between active and standby nodes, facilitating seamless failover and uninterrupted network sessions in cloud computing environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-18
- Publication Date
- 2026-03-12
AI Technical Summary
Network address translation (NAT) in cloud computing environments experiences long interruptions and data loss during node failures, particularly in applications running on virtualized customer premises equipment, due to the time-consuming process of reestablishing network sessions.
Implementing a method for maintaining consistent NAT and TCP state tables between active and standby nodes, allowing seamless failover by synchronizing data structures and routing tables, ensuring uninterrupted network sessions through shared sub-interfaces and synchronized databases.
Enables low-latency network address translation with minimal disruption, maintaining active network and application sessions without the need for reestablishment, reducing data loss and latency during node failures.
Smart Images

Figure 0007828533000001 
Figure 0007828533000002 
Figure 0007828533000003
Abstract
Description
[Technical Field]
[0001] It is often beneficial to have a public address used by an entity that is different from the private address used by the entity. [Background technology]
[0002] Public addresses can be used as source and destination addresses for packets sent and received on an external network. Private addresses can be used as source and destination addresses for packets sent and received on an internal network. The translation between public and private addresses, known as network address translation (NAT), can be performed by network elements such as routers, switches, network gateways, or other computing devices. Summary of the Invention [Problem to be solved by the invention]
[0003] Network address translation is particularly useful for applications running in cloud computing environments, which are network environments where application execution is virtualized. Applications running on customer premises equipment can also benefit from address translation. [Means for solving the problem]
[0004] To solve the above problem, the method claimed in the present application is provided.
[0005] In order that the advantages of the present invention may be readily understood, a more particular description of the invention, briefly described above, will now be set forth with reference to specific embodiments thereof which are illustrated in the accompanying drawings, which are to be understood as illustrating only typical embodiments of the invention and are therefore not intended to limit the scope of the invention, and which will be described in detail and with additional particularity using the accompanying drawings. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a schematic block diagram of a network environment implementing low latency network address translation according to an embodiment of the present invention. [Figure 2] FIG. 1 is a schematic diagram illustrating the execution of network address translation failover between an active node and a standby node according to an embodiment of the present invention. [Figure 3] FIG. 2 is a process flow diagram of a method for performing network address translation failover between an active node and a standby node according to an embodiment of the present invention. [Figure 4] FIG. 2 is a schematic block diagram of components implementing an alternative approach for network address translation failover between an active node and a standby node according to an embodiment of the present invention. [Figure 5] FIG. 10 is a process flow diagram of an alternative method for performing network address translation failover between a working node and a standby node that perform network address translation according to an embodiment of the present invention. [Figure 6] FIG. 10 is a process flow diagram of another alternative method for performing network address translation failover between a working node and a standby node that perform network address translation according to an embodiment of the present invention. [Figure 7] FIG. 1 is a schematic diagram of a computer system suitable for implementing methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0007] It will be readily understood that the components of the present invention, as illustrated and described in the accompanying drawings, could be designed and arranged in a wide variety of different configurations. Thus, the following more detailed description of the illustrated embodiments of the invention is not intended to limit the scope of the invention as set forth in the claims, but merely to illustrate specific examples of embodiments in accordance with the present invention being discussed herein. The embodiments described herein can best be understood by referring to the drawings, wherein like parts are designated with like reference numerals throughout the specification and drawings.
[0008] Embodiments in accordance with the present invention may be embodied as an apparatus, a method, or a computer program product. Accordingly, the present invention may be embodied in the form of an entirely hardware form, an entirely software form (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "module" or "system." Furthermore, the present invention may take the form of a computer program product embodied in any tangible medium of expression having computer-usable program code embodied in the medium.
[0009] Any combination of one or more computer usable or computer readable media may be utilized. For example, computer readable media may include one or more of a portable computer diskette, a hard disk, a random access memory (RAM) device, a read only memory (ROM) device, an erasable programmable read only memory (EPROM or flash memory) device, a portable compact disc read only memory (CDROM), an optical storage device, and a magnetic storage device. In selected embodiments, computer readable media may include any non-transitory medium that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0010] Computer program code for carrying out operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the C programming language or similar programming languages, and can also use script or markup languages such as HTML, XML, JSON, etc. The program code can be executed entirely on the computer system as a standalone software package, on a standalone hardware unit, partially on a remote computer located some distance from the computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0011] The present invention is described below with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions or code. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, and can produce a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, generate means for implementing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0012] These computer program instructions may also be stored on a non-transitory computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instruction means that implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0013] Computer program instructions are also loaded onto a computer or other programmable data processing apparatus and cause a series of operational steps to be performed on the computer or other programmable data processing apparatus to create a computer-implemented process, such that the instructions, when executed on a computer or other programmable data processing apparatus, provide a process for implementing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0014] FIG. 1 illustrates an example of a network environment 100 capable of performing low-latency network address translation (NAT). The network environment 100 may include multiple virtual private clouds (VPCs) 102 running on a cloud computing platform 104. The cloud computing platform 104 may be any cloud computing platform known to those skilled in the art, such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud, or other cloud computing platforms. As known to those skilled in the art, the cloud computing platform 104 provides virtualized computing and storage resources that can be accessed by customers over a network 106. The virtual private clouds 102 may function as logically isolated parts of the cloud computing platform 104, and the network, including Internet Protocol (IP) addresses, subnets, routing tables, network gateways, and other network elements, are isolated from other parts of the cloud computing platform and therefore do not need to be globally unique.
[0015] Each virtual private cloud 102 may run one or more nodes 108, 110, 112. Each node 108, 110, 112 may be an application running within the virtual private cloud 102, such as a virtual machine, container, or execution environment running within the virtual private cloud 102. Each node 108, 110, 112 may function as a network element, providing a gateway between the virtual private cloud 102 and other virtual private clouds 102, as well as to external networks 106 connected to the cloud computing platform 104.
[0016] In the illustrated embodiment, the one or more nodes are a hub node 108 connected to the external network 106 and other spoke nodes 110, 112 that can send and receive traffic to and from the external network 106 through the hub node 108. In the illustrated embodiment, the hub node 108 is located in a different virtual private cloud than the spoke nodes 110, 112. In some embodiments, the spoke nodes 110, 112 each have a virtual private network (VPN) session established with the hub node 108. For example, the VPN session may be via Layer 2 Tunneling Protocol (L2TP).
[0017] Each virtual private cloud 102 may further execute one or more workloads 114. Each workload 114 may communicate with workloads in other virtual private clouds 102 and with the external network 106 by way of the nodes 110, 112 within the virtual private cloud that hosts the workload 114. Thus, each workload 114 may have a network connection, such as a virtual local area network (VLAN) connection, to the nodes 110, 112 of its virtual private cloud. Each workload 114 may be an application, a daemon, a network service, an operating system, a container, or any computational process executable by a computer.
[0018] In the illustrated embodiment, one node 110 may be active for one or more workloads 114, and another node 112 may be a backup or standby node 112 for node 110. The active node 110 and the standby node 112 may be located within the same virtual private cloud 102, and a workload may connect to either node 110, 112 through the virtual private cloud's 102 internal virtual network.
[0019] As described herein, the current node 110 can maintain network sessions with components outside the virtual private cloud 102 and possibly outside the cloud computing platform 104. The network sessions can be Transmission Control Protocol (TCP) sessions with a TCP server 116 connected to the external network 106. The network sessions can additionally or alternatively be application sessions established between an external application and a workload 114 connecting to the TCP server 116.
[0020] The approach described herein allows a workload 114 to maintain a session even if a current node 110 fails. In some cases, a TCP session failure results in a long interruption before the session participants perform a handshake to reestablish the session. Similarly, an application session may experience a long interruption before the application attempts to reestablish a new application session. The process of reestablishing a new application session can be time-consuming and may result in data loss. Therefore, the ability to handle the failure of a node 110, 112 without interrupting the network session is of great benefit.
[0021] 2, the current node 110 may maintain various data structures that describe the state of the network interfaces managed by the current node 110. This may include a network address translation (NAT) table 200a, which contains mappings between private IP addresses (within VPC 102) and public IP addresses (addresses used outside VPC 102) of each workload 114. The allocation of one or both of the public and private IP addresses may be dynamic, with allocations being made and released by the workloads according to their communication requirements.
[0022] For example, the hub node 108 can function as a Dynamic Host Configuration Protocol (DHCP) server. The current node 110 can operate as a dynamic Network Address Translation (NAT) server. Thus, a public IP address leased to the current node 110 by the DHCP server can be dynamically mapped by the current node 110 to a private IP address of a workload 114. When a first workload completes transmission or a predetermined period of time expires, the current node 110 can map the public IP address to a different workload 114. The assignment of public IP addresses by the hub node 108 can be stored in the NAT table 200a of the current node 110 for each assignment of a public IP address, and the public IP address can be mapped to the private IP address of the workload 114 that was assigned the public IP address by the current node 110 when performing dynamic network address translation on behalf of the workload 114.
[0023] In some embodiments, a private IP address can be assigned to a sub-interface of the VPN session connecting the working node 110 to the hub node 108. For example, the hub node 108 can create a sub-interface for the workload 114 that includes the workload's MAC (Media Access Control) address. The sub-interface can be further assigned one or both of the public and private IP addresses assigned to that workload 114. For each sub-interface created on the working node 110, the hub node 108 can create another sub-interface on the VPN connection between the hub node 108 and the standby node 112 that has the same MAC address and the same public and private IP addresses. In some embodiments, a sub-interface with the same MAC address is created on the standby node 112, which then attempts to acquire one or both of the interface's public and private IP addresses. Since the MAC address of standby node 112 is the same as the interface on working node 110, hub node 108 then assigns the same public and private IP addresses to its sub-interface via DHCP as working node 110.
[0024] The current node 110 updates the routing table 204 so that the default route for the workload 114 points to the network interface of the current node 110. The routing table 204 may also reference the current node 110 as the network gateway for the VPC 102.
[0025] The working node 110 may maintain a TCP state 202a. The TCP state 202a may maintain the TCP connection state for the public IP address of the workload 114. As known to those skilled in the art, a TCP connection may implement a state machine that changes due to interactions between components connected by the TCP connection. Thus, a TCP state 202a may be this state machine for each TCP connection of each workload 114.
[0026] Standby node 112 may maintain its own copies 200b and 202b of NAT table 200a and TCP state 202a. As TCP state 202a changes, working node 110 may communicate changes to TCP state 202a to standby node 112, and standby node 112 may update TCP state 202b with these updates. In some embodiments, communication of TCP state updates is performed by working node 110 and sent via hub node 108. In other cases, updates are performed by direct communication between nodes 110, 112.
[0027] In addition to sharing updates to the NAT table 200a and TCP state 202a, the nodes 110, 112 can communicate with each other to facilitate detection of a failure of the current node 110. This communication can be performed using the same connection used to share updates to the NAT table 200a and TCP state 202a, and can be a direct connection within the VPC 102 or via the hub node 108. This communication can include an "alive" message sent from the current node 110 to the standby node 112 at predetermined intervals (e.g., 100 microseconds to 2 seconds). Thus, the standby node 112 can detect a failure of the current node 110 in response to not receiving an active message within a threshold period from the last active message received. In some embodiments, the standby node 112 can periodically (e.g., every 100 microseconds to 2 seconds) send a query to the current node 110 and detect a failure of the current node 110 in response to not receiving a response to the query within a threshold period from the time the query was sent.
[0028] Referring to Figure 3, and also to Figure 2, the current node 110 and the standby node 112 can implement the illustrated method 300. The method 300 includes maintaining 302 consistency between the NAT tables 200a, 200b and between the TCP states 202a, 202b. As discussed above, this includes sending updates from the current node 110 to the standby node 112 when the NAT table 200a and the TCP state 202a change. These updates can be sent directly or via the hub node 108. The updates include adding or deleting entries in the NAT table 200b when corresponding entries are created or deleted in the NAT table 200a. The updates also include creating or deleting state machines for TCP connections when state machines for TCP connections are created or terminated on the current node 110. Step 302 further includes creating or deleting a subinterface for the workload 114 on the standby node 112 as the corresponding subinterface is created or deleted on the working node, as described above.
[0029] The method 300 includes monitoring 304, by the standby node 112, the status of the active node 110. This includes monitoring for active messages sent by the active node 110. This may additionally or alternatively include monitoring for responses to queries sent by the standby node 112.
[0030] The method 300 includes the standby node 112 detecting 306 a failure of the active node 110. As discussed above, this may include not receiving an active message within a threshold period from the last active message, or not receiving a response to a query within a threshold period from the time the query was sent.
[0031] When a failure is detected 306, the method 300 includes updating 308 the routing table 204 to replace references to the active node 110 with references to the standby node 112. In particular, this includes referencing the standby node 112, e.g., the private IP address of the standby node 112, as the default gateway for the VPC 102.
[0032] The standby node 112, which is now the working node of the VPC 102, then processes 310 traffic received from the external network 106 and workload 114 according to the copied NAT table 200b and TCP state 202b. In particular, the standby node 112 can perform the functions attributed to the working node 110 acting as a network gateway, including performing NAT, managing the TCP state machine, and other functions attributed to the working node 110.
[0033] During the transition from the active node 110 to the standby node 112, some packets may be lost. However, the TCP protocol provides for retransmission of lost packets. Therefore, TCP sessions remain active and do not need to be re-established. Similarly, because each workload 114 on the standby node 112 can use the same subinterface, any application sessions can continue to operate uninterrupted. Because the NAT table 200a remains the same, applications configured to communicate with the workload's public IP address do not need to obtain a new address and establish a TCP session and new application for the new address. In some embodiments, after the standby node 112 becomes the active node, the standby node 112 can also function as a NAT server for the workloads in the VPC 102.
[0034] As described above, sub-interfaces to the VPN connection between nodes 110, 112 and hub node 108 can be created for each workload 114, and for each sub-interface created on the VPN connection between working node 110 and hub node 108, a corresponding sub-interface with the same public and private IP addresses and MAC addresses is created on the VPN connection between stand-by node 112 and hub node 108. Thus, step 310 involves sending traffic (e.g., TCP packets) for each workload 114 on a sub-interface of stand-by node 112 that has the same public and private IP addresses and MAC addresses as the sub-interface of working node 110 previously used by the respective workload. Thus, delays due to creating new sub-interfaces can be avoided in the event of a failure.
[0035] The change in routing of traffic to / from workload 114 occurs by standby node 112 becoming the new default gateway and by workload 114's private IP address and MAC address remaining the same, and traffic received by standby node 112 that references workload 114's private IP address or MAC address is routed through the appropriate sub-interface associated with the private IP address and MAC address.
[0036] In some embodiments, when the active node 110 resumes operation after a failure, the active node 110 may function as a standby node, i.e., receive a duplicate sub-interface on the hub node 108's VPN connection to the standby node 112, and / or duplicate information for NAT table 200b, TCP state 202b, as described above. When this information is up-to-date for the standby node 112, the active node 110 becomes active again, and the standby node 112 resumes functioning as a standby node.
[0037] The embodiment of Figures 1-3 is shown as being implemented within a cloud computing platform 104. This approach can be implemented by any computing node, including customer premises equipment, connected such that it can act as a hub node to the spoke nodes described above.
[0038] 4 illustrates a configuration in which the active node 110 and the standby node 112 are not connected by the hub node 108. The nodes 110, 112 may run on separate computing devices within the VPC 102 or on an on-premises network. In the illustrated configuration, the active node 110 and the standby node 112 are assigned static pools of IP addresses, and the active node 110 acts as a NAT server and a DHCP server for the workload 114.
[0039] In the illustrated embodiment, each node 110, 112 includes a control plane 400 that can implement logic to perform the functions of the respective node 110, 112 as the active node 110 and the standby node 112 described herein, respectively. In either node 110, 112 that is active, the control plane can act as a DHCP and NAT server for workloads 114 connected to the external network 106 through the active node 110.
[0040] Each node 110, 112 includes or has access to a database 402. The databases 402 are synchronized and updated so that the database 402 of the standby node 112 is the same as the database 402 of the active node 110. For example, the databases 402 can be REDIS databases configured to synchronize with each other.
[0041] The current node 110 creates a NAT table, such as a secure NAT (SNAT) table, that maps private addresses to the MAC addresses of workloads 114 and maps private addresses assigned to workloads 114 to public addresses assigned to that workload 114. The SNAT mappings may also be recorded in the kernel IP table of the device (real or virtual) running the current node 110. As mentioned above, there may be a static pool 408 of IP addresses managed by the current node 110, and public and / or private IP addresses are returned to the pool 408 and subsequently assigned to a second workload 114 after the public and / or private IP addresses assigned to the first workload complete their tasks or after the lease for the public and / or private IP addresses expires.
[0042] The nodes 110, 112 may further include a forwarding information base (FIB) 406 or other data structure that defines the routing of packets received by the nodes 110, 112. In particular, the forwarding information base 406 may define on which output port a packet received on a particular input port should be output. Thus, the forwarding information base 406 may be configured to route packets destined for an external IP address to the external network 106 and to route received packets destined for a public IP address to the private IP address of a workload 114 that has been assigned a public IP address in the SNAT table 404.
[0043] The SNAT table 404, the forwarding information base 406 of the current node 110, and other information such as TCP state information can be written to the database 402 of the current node 110. The database 402 then synchronizes with the database 402 of the standby node 112. The standby node 112 can then populate its SNAT table 404 and forwarding information base 406 according to the database 402 in preparation for a failure of the current node 110.
[0044] When a particular node 110, 112 is an active node, it receives ingress traffic 410, translates it according to the SNAT table, and then outputs it as output traffic 412 to an egress port or to the kernel of a computing device (real or virtual) running the node 110, 112 defined in the forwarding information base 406.
[0045] Referring to Figure 5, the illustrated method 500 can be performed using the system illustrated in Figure 4. The method 500 can be performed by the working node 110, except for the operations attributed to the standby node 112.
[0046] The method 500 includes allocating 502 an IP address to each workload 114 from a static IP pool. This can be performed by DHCP or other IP configuration protocol. Step 502 further includes making an entry in a SNAT table 404. The method 500 further includes writing 504 the entry in the SNAT table 404 into the synchronized database 402. As a result, the data in the database 402 is replicated to the database 402 of the standby node 112.
[0047] For each workload 114 assigned an IP address in step 502, the current node 110 further creates 506 a sub-interface for the workload 114 that is assigned a static IP address (e.g., a static public IP address) by referencing the MAC address of the workload 114. Thus, traffic to and from the workload 114 is routed by the current node 110 through the sub-interface. A corresponding sub-interface is also created on the standby node 112 by referencing the MAC address of the workload 114 and the public IP address assigned to the workload 114. The private IP address of the workload 114 is also associated with the sub-interface on the nodes 110, 112.
[0048] The inverse of steps 504 and 506 is also performed. As workload 114 relinquishes its private and / or public IP addresses, the corresponding entries in SNAT table 404 are deleted, and the sub-interfaces of workload 114 are similarly deleted. To maintain consistency, the corresponding entries in the sub-interfaces and SNAT table 404 on standby node 112 are similarly deleted. These updates can be communicated by updating database 402 of active node 110, resulting in an update of database 402 of standby node 112 to indicate the deleted information.
[0049] Method 500 further includes monitoring 508 the status of the current node 110 and detecting 510 a failure of the current node 110. This may be performed as described above in steps 304 and 306 of method 300 using periodic up-to-date messages or queries.
[0050] When a failure is detected 510, traffic is routed 512 to the standby node 112 instead of the active node 110. The routing change is implemented by modifying the routing table 204 in the VPC 102 that includes the nodes 110, 112. The routing change includes configuring the workload 114 to use the standby node 112 as a default gateway.
[0051] The standby node 112 then processes 514 the received traffic according to the SNAT table 404, subinterfaces, forwarding information base 406, TCP state, or other data received from the current node 110 before the failure. In particular, the standby node 112 can perform the functions attributed to the current node 110 acting as a network gateway, including performing NAT, managing the TCP state machine, routing according to the forwarding information base 406, and other functions attributed to the current node 110.
[0052] Similar to the embodiment of Figures 1-3, the standby node has already configured all or some of the workload's sub-interfaces, SNAT table 404, and forwarding information base 406 before the failure occurs, so that it can route traffic to and from the workload 114 without disrupting higher-level network sessions, such as TCP sessions and application sessions.
[0053] 1-3, the approaches of Figures 4 and 5 can be implemented on cloud computing platform 104 or customer premises equipment, with each node 110, 112 running on a different computing device. Workload 114 can run on the same customer premises equipment as nodes 110, 112, or on different customer premises equipment.
[0054] 1-3, when active node 110 resumes operation after a failure, it can function as a standby node, i.e., receive replica information of sub-interfaces, forwarding information bases, and / or SNAT table 404 from standby node 112, as described above. Once this information is up to date for standby node 112 and the corresponding sub-interfaces have been created on active node 110, active node 110 becomes active again, and standby node 112 again functions as a standby node.
[0055] FIG. 6 illustrates a method 600 that can be used to perform a failover between an active node 110 and a standby node 112 due to the lack of a synchronized database 402 in each node 110, 112.
[0056] In method 600, the active node 110 and the standby node 112 establish a connection between each other 602. In the illustrated embodiment, this connection is a User Datagram Protocol (UDP) connection. The active node 110 then notifies the standby node 112 on this connection 604. The notification may include a notification of sufficient information to enable the standby node 112 to recreate the sub-interface for the workload 114 that was created on the active node 110. The notification may include information such as the private IP address of the workload 114, the public IP address mapped to the public IP address in the SNAT table, and the MAC address. Thus, the standby node 112 creates a sub-interface with the private IP address, public IP address, and MAC address indicated in the notification.
[0057] The notification may also include notification that an interface has been deleted or an entry in a SNAT table has been deleted due to the workload terminating a network session or relinquishing a private and / or public IP address. Accordingly, the standby node 112 deletes the interface referenced in the notification and / or updates the SNAT table to delete the entry referenced in the notification.
[0058] Method 600 further includes monitoring 606 the status of the current node 110 and detecting 608 a failure of the current node 110. This may be performed as described above in steps 304 and 306 of method 300 using periodic up-to-date messages or queries.
[0059] When a failure is detected 608, traffic is routed 610 to the standby node 112 instead of the active node 110. The routing change is implemented by modifying the routing table 204 in the VPC 102 that includes the nodes 110, 112. The routing change includes configuring the workload 114 to use the standby node 112 as a default gateway.
[0060] The standby node 112 then processes 612 the traffic according to the SNAT table 404, forwarding information base 406, TCP state, and / or sub-interface received from the working node 110. In particular, the standby node 112 may perform the functions attributed to the working node 110 acting as a network gateway, including performing NAT, managing the TCP state machine, routing according to the forwarding information base 406, and other functions attributed to the working node 110.
[0061] 1-3, the approach of FIG. 6 can be implemented on cloud computing platform 104 or customer premises equipment, with each node 110, 112 running on a different computing device. Workload 114 can run on the same customer premises equipment as nodes 110, 112, or on different customer premises equipment.
[0062] 1-3, when active node 110 resumes operation after a failure, it can function as a standby node, i.e., receive replica information of sub-interfaces, forwarding information bases, and / or SNAT table 404 from standby node 112, as described above. Once this information is up to date for standby node 112 and the corresponding sub-interfaces have been created on active node 110, active node 110 becomes active again, and standby node 112 again functions as a standby node.
[0063] 7 is a block diagram illustrating an exemplary computing device 700 that can be used to implement the methods and systems disclosed herein. In particular, nodes 108, 110, 112 according to any of the above-described embodiments can have all or some of the attributes of computing device 700. Similarly, a cloud computing platform can be comprised of devices that have all or some of the attributes of computing device 700.
[0064] Computing device 700 can be used to perform various processes as described herein. Computing device 700 can function as a server, a client, or other computing entity. The computing device can perform various monitoring functions described herein and can execute one or more application programs, such as the application programs described herein. Computing device 700 can be any of a wide variety of computing devices, such as a desktop computer, a laptop, a server computer, a handheld computer, a tablet, etc.
[0065] Computing device 700 includes one or more processors 702, one or more memory devices 704, one or more interfaces 706, one or more mass storage devices 708, one or more input / output (I / O) devices 710, and a display device 730, all connected to a bus 712. Processor 702 includes one or more processors or controllers and executes instructions stored in memory device 704 and / or mass storage device 708. Processor 702 may also include various types of computer-readable media, such as cache memory.
[0066] The memory device 704 includes a variety of computer-readable media, such as volatile memory (e.g., random access memory (RAM) 714) and / or non-volatile memory (e.g., read-only memory (ROM) 716). The memory device 704 may also include re-writable ROM, such as flash memory.
[0067] The mass storage device 708 includes a variety of computer-readable media, such as magnetic tape, magnetic disks, optical disks, solid-state memory (e.g., flash memory), etc. As illustrated in Figure 7, a particular mass storage device is a hard disk drive 724. Various drives may also be included within the mass storage device 708 to allow reading from and / or writing to the various computer-readable media. The mass storage device 708 includes removable media 726 and / or non-removable media.
[0068] Input / output devices 710 include various devices that allow data and / or other information to be input to or obtained from computing device 700. Exemplary input / output devices 710 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, lenses, CCD or other imaging devices, etc.
[0069] Display device 730 includes any type of device capable of displaying information to one or more users of computing device 700. Examples of display device 730 include a monitor, a display terminal, a video projection device, etc.
[0070] The interface 706 includes various interfaces that allow the computing device 700 to exchange information with other systems, devices, or computing environments. An exemplary interface 706 includes any number of different network interfaces 720, such as interfaces to a local area network (LAN), a wide area network (WAN), a wireless network, and the Internet. Other interfaces include a user interface 718 and a peripheral interface 722. The interface 706 may also include one or more user interface elements 718. The interface 706 may also include one or more peripheral interfaces, such as interfaces for a printer, a pointing device (mouse, trackpad, etc.), a keyboard, etc.
[0071] The bus 712 allows the processor 702, memory device 704, interface 706, mass storage device 708, and input / output device 710 to communicate with each other, as well as other devices or components connected to the bus 712. The bus 712 may represent one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, a USB bus, etc.
[0072] For purposes of illustration, programs and other executable program components are shown herein as separate blocks, but it is understood that such programs and components may reside at various times in different storage components of computing device 700 and be executed by processor 702. Alternatively, the systems and processes described herein may be implemented in hardware, or in a combination of hardware, software, and / or firmware. For example, one or more application-specific integrated circuits (ASICs) may be programmed to perform one or more of the systems and processes disclosed herein. [Explanation of symbols]
[0073] 100 Network Environment 300 ways 410 ingress traffic 412 outgoing traffic 500 ways 600 ways 700 computing devices 712 Bus
Claims
1. providing a plurality of workloads executing within a computing environment including a plurality of computing devices, each of the computing devices including a processing unit and a memory unit, each of the plurality of workloads being an application, a daemon, a network service, an operating system, a container, or a computer-executable computational process; providing a first node executing within the computing environment, the first node programmed to act as a first gateway between the computing environment and an external network by performing network address translation (NAT), the computing environment configured to cause the plurality of workloads to communicate with the external network through the first node; providing a second node executing within the computing environment programmed to act as a second gateway between the computing environment and the external network by performing network address translation (NAT); the second node is configured to mirror the NAT state of the first node; creating a first interface to the plurality of workloads on the first node and a second interface to the plurality of workloads on the second node that is identical to the first interface; the second node detects a failure of the first node; by the second node in response to detecting a failure of the first node; configuring a computing environment to cause the plurality of workloads to communicate with the external network through the second node utilizing the second interface, the second interface being created prior to the failure of the first node; performing network address translation according to the NAT state of the first node; This includes: The method, wherein the first interface is a sub-interface to a first VPN (Virtual Private Network) connection and the second interface is a sub-interface to a second VPN connection.
2. the computing environment defines a routing table; 2. The method of claim 1, further comprising configuring a computing environment such that the plurality of workloads programmed to communicate with the external network through the second node cause a reference to the first node in a routing table to be replaced with a reference to the second node.
3. 2. The method of claim 1, wherein the NAT state is a NAT table that contains mappings between private IP addresses of the plurality of workloads and public IP addresses of the plurality of workloads.
4. 2. The method of claim 1, wherein the first interface references a media access control (MAC) address of the plurality of workloads, and the second interface references a MAC address of the plurality of workloads.
5. the first node is connected to a hub node by the first VPN connection, the hub node providing connectivity between the first node and the external network; 2. The method of claim 1, wherein the second node is connected to the hub node by the second VPN connection, the hub node providing connectivity between the second node and the external network.
6. 6. The method of claim 5, wherein the computing environment is a cloud computing environment, and the first node and second node execute within a virtual private cloud (VPC) within the cloud computing environment.
7. The method further comprises: maintaining a protocol state machine on a first node of the plurality of workloads; copying the protocol state machine to the second node; 7. The method of claim 6, further comprising: in response to detecting a failure of the first node, continuing to route traffic to the plurality of workloads by the second node in accordance with the protocol state machine.
8. 8. The method of claim 7, wherein each of the protocol state machines is a TCP (Transmission Control Protocol) state machine.
9. 10. The method of claim 8, further comprising configuring the computing environment to cause the plurality of workloads to communicate with the external network through the second node without causing the plurality of workloads to create new TCP sessions.
10. 2. The method of claim 1, further comprising configuring the computing environment to cause the plurality of workloads to communicate with the external network through the second node without causing the plurality of workloads to create new application sessions.
11. The method further comprises: maintaining, by said first node, a first database containing said NAT state; maintaining, by said second node, a second database; synchronizing the first database and the second database by the first node and the second node; 2. The method of claim 1, comprising:
12. executing a plurality of workloads in a computing environment, each workload of the plurality of workloads being an application, a daemon, a network service, an operating system, a container, or a computer-executable computational process; running a first node within the computing environment, the first node connected to the plurality of workloads and managing network communications between the plurality of workloads and an external network outside the computing environment; creating, by the first node, first network interfaces for the plurality of workloads for communication with the external network; creating, by a second node executing within the computing environment, a second network interface having the same address as the first network interface and used by the plurality of workloads; detecting, by the second node, a failure of the first node; and configuring the computing environment, in response to detecting a failure of the first node, to cause the second node to cause the plurality of workloads to communicate with the external network through the second node and the second network interface, the second network interface being created before the failure of the first node, the first network interface being a sub-interface to a first VPN (Virtual Private Network) connection, and the second network interface being a sub-interface to a second VPN connection. A method comprising:
13. The method further comprises: performing, by the first node, network address translation (NAT) using a first NAT table that maps private IP addresses of the plurality of workloads used within the computing environment to public IP addresses used within the external network; maintaining, by the second node, a second NAT table identical to the first NAT table; responsive to detecting a failure of the first node, performing network address translation by the second node using the second NAT table; 13. The method of claim 12, comprising:
14. The first VPN connection is a connection to a hub node that connects the first node to the external network; 13. The method of claim 12, wherein the second VPN connection is a connection to the hub node that connects the second node to the external network.
15. The method further comprises: managing, by the hub node, network address translation (NAT) by creating a first NAT table on the first node, the first NAT table mapping private IP addresses of the plurality of workloads used within the computing environment to public IP addresses used within the external network; 15. The method of claim 14, further comprising creating a second NAT table on the second node by the hub node that is identical to the first NAT table.
16. 13. The method of claim 12, wherein the computing environment is a cloud computing platform.
17. 17. The method of claim 16, wherein the first node and the second node execute within a virtual private cloud (VPC) within the cloud computing platform.
18. 17. The method of claim 16, wherein the first node and the second node execute on different computing devices connected to the same local network.
Citation Information
Patent Citations
Address conversion system, address conversion duplication method and program
JP2017017465A
JPP6579608B
Redundancy support for network address translation (NAT)
US20100254255A1