Host routed overlay with deterministic host learning and localized integrated routing and bridging
A host-routed overlay with localized IRB on CE routers on bare-metal servers addresses operational complexities in data center routing by providing deterministic host learning and efficient routing, enhancing scalability and mobility in data center networks.
Patent Information
- Application Number
- JP2025134822
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-08-23
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-05
AI Technical Summary
Existing data center overlay routing architectures, such as centralized gateway and distributed anycast gateway architectures, face issues with operational complexity, single points of failure, and inefficient host learning due to reliance on ARP-based methods, leading to scalability bottlenecks and unpredictable network behavior.
Implementing a host-routed overlay solution with localized integrated routing and bridging (IRB) on virtual customer edge (CE) routers on bare-metal servers, which provides deterministic host learning and eliminates the need for Layer 2 bridging and ARP-based learning, enabling flexible workload placement and mobility across stretched subnets.
This approach simplifies network operations by eliminating complex Layer 2 functions on leaf nodes, reduces the risk of single points of failure, and enhances network scalability and efficiency through deterministic host learning and routing, supporting IP unicast inter- and intra-subnet VPN connectivity and virtual machine mobility.
Smart Images

Figure 2025166157000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 722,003, entitled "DATABASE SYSTEMS METHODS AND DEVICES," filed August 23, 2018, which is incorporated herein by reference in its entirety, including but not limited to, the portions specifically set forth below, with the following exception: In the event that any portion of said application conflicts with this application, the contents of this application take precedence over said application.
[0002] The present disclosure relates to computing networks, and more particularly to network routing protocols. [Background technology]
[0003] Network computing is a means by which multiple computers, or nodes, work together and communicate with each other over a network. These include wide area networks (WANs) and local area networks (LANs). Local area networks are typically used for smaller, localized networks, such as homes, businesses, and schools. Wide area networks cover larger areas, such as cities, and can even connect computers in different countries. Local area networks are typically faster and more secure than wide area networks, but wide area networks allow for broader connectivity. Local area networks are typically owned, controlled, and managed within the organization in which they are deployed, while wide area networks typically require the connection of two or more constituent local area networks, either via the public Internet or private connections established by a telecommunications provider.
[0004] Local and wide area networks connect computers together and allow the transfer of data and other information. Both local and wide area networks require a means of determining the path along which data should be passed from one computing instance to another. This is also known as routing. Routing is the process of selecting a path for traffic within a network, between networks, or across networks. The routing process typically directs forwarding based on routing tables, which maintain records of routes to various network destinations. Routing tables may be specified by an administrator, learned by monitoring network traffic, or constructed with the assistance of a routing protocol.
[0005] One network architecture is the multi-tenant datacenter. The multi-tenant datacenter defines an end-to-end system suitable for service deployment in a public cloud-based or private cloud-based model. The multi-tenant datacenter may include a wide area network, multiple provider datacenters, and tenant resources. The multi-tenant datacenter may include a multi-layer hierarchical network model. The multi-layer hierarchy may include a core layer, an aggregation layer, and an access layer. The multiple layers may include a Layer 2 overlay and a Layer 3 overlay with an L2 / L3 boundary. Summary of the Invention [Problem to be solved by the invention]
[0006] One data center overlay routing architecture is the centralized gateway architecture. Another data center overlay routing architecture is the distributed anycast gateway architecture. These architectures have many drawbacks, as further explained below. [Means for solving the problem]
[0007] In view of the foregoing, systems, methods, and devices for improved routing architectures are disclosed herein.
[0008] Non-limiting and non-exhaustive embodiments of the present disclosure are described with reference to the following figures, in which like reference numerals refer to like parts throughout the figures unless otherwise specified. Advantages of the present disclosure will be more clearly understood by referring to the following description and the accompanying drawings. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a schematic diagram of a system of networked devices communicating over the Internet. [Figure 2] FIG. 1 is a schematic diagram of a leaf-spine network topology with a centralized gateway datacenter overlay routing architecture as known in the prior art. [Figure 3] FIG. 1 is a schematic diagram of a leaf-spine network topology with a distributed anycast gateway datacenter overlay routing architecture as known in the prior art. [Figure 4]1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server. [Figure 5] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating host learning at boot-up. [Figure 6] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary, which is pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating local forwarding state. [Figure 7] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary, which is pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating remote forwarding states. [Figure 8A] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating intra-subnet server local flows. [Figure 8B] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating inter-subnet server local flows. [Figure 9A] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating an intra-subnet overlay flow form address 12.1.1.4 to address 12.1.1.2. [Figure 9B]FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating an inter-subnet overlay flow from address 12.1.1.4 to address 10.1.1.2. [Figure 10] FIG. 1 is a schematic diagram of a data center fabric architecture with overlay routing at the L2-L3 boundary pushed to a virtual customer edge (CE) router gateway on a bare metal server, illustrating a server link failure. [Figure 11] FIG. 1 is a schematic diagram illustrating components of an exemplary computing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Disclosed herein are systems, methods, and devices for a routed overlay solution for Internet Protocol (IP) subnet stretching using localized integrated routing and bridging (IRB) on host machines. The disclosed systems, methods, and devices provide a virtual first-hop gateway on a virtual customer edge (CE) router on a bare-metal server. The virtual CE router provides localized East-West integrated routing and bridging (IRB) services for local hosts. In one embodiment, a default routed equal-cost multipath (ECMP) provides uplinks from the virtual CE router to leaf nodes for north-south and east-west connectivity.
[0011] The systems, methods, and devices disclosed herein realize numerous networking advantages. The systems enable deterministic host learning without the need for Address Resolution Protocol (ARP)-based learning of routes. The improved systems described herein eliminate age-out, probing, and synchronization, and do not require media access control (MAC) entries on leaf nodes. The improved systems also eliminate the need for complex multi-chassis link aggregation (MLAG) bridging functions in leaf nodes. Furthermore, the virtual CE routers described herein store local Internet Protocol (IP) addresses and media access control (MAC) addresses along with default ECMP routes to leaf nodes. Furthermore, the improved systems described herein provide host routing in leaf nodes for stretched subnets, enabling host mobility.
[0012] In an Ethernet virtual private network (EVPN)-enabled multiple tenant data center overlay, the architecture features a distributed anycast Layer-3 (L3) gateway on the leaf node, providing the first-hop gateway function for workloads. This pushes the service layer L2-L3 boundary down to the leaf node. In other words, all inter-subnet virtual private network (VPN) traffic from workload host virtual machines is routed on the leaf node. Virtual machine mobility and flexible workload placement are achieved by stretching the Layer 2 overlay across the routed network fabric. Intra-subnet traffic across the stretched Layer 2 domain is overlay bridged on the leaf node. The leaf node provides EVPN-IRB services for directly connected host virtual machines. This routes all inter-overlay subnetwork VPN traffic and bridges all intra-overlay subnetwork VPN traffic through the routed fabric underlay.
[0013] The embodiments described herein eliminate the need to support overlay bridging functionality on leaf nodes. Furthermore, the embodiments described herein eliminate the need for Layer 2 MLAG connections and associated complex procedures between leaf nodes and hosts. Furthermore, the embodiments described herein eliminate the need for data plane and ARP-based host learning on leaf nodes. The embodiments disclosed herein achieve these benefits while also providing IP unicast inter- and intra-subnet VPN connectivity, virtual machine mobility, and flexible workload placement across stretched IP subnets.
[0014] Embodiments of the present disclosure separate local Layer 2 switching and IRB functions from leaf nodes and localize them to small virtual CE routers on bare metal servers. This is achieved by running small virtual router VMs on the bare metal servers, which act as first-hop gateways for host virtual machines and provide local IRB switching between virtual machines local to the bare metal servers. This virtual router operates as a traditional CE router that can be multihomed to multiple leaf nodes via Layer 3 routed interfaces on the leaf nodes. Leaf nodes in the fabric function as pure Layer 3 VPN PE routers without any Layer 2 bridging or IRB functionality. To enable flexible placement and mobility of Layer 3 endpoints across the DC overlay while providing optimal routing, traffic can be host-routed on the leaf nodes rather than subnet-routed. This is also the case with EVPN-IRB.
[0015] The improved routing architecture described here (see Figures 4-10) can provide the benefits of a completely routed network fabric. However, the EVPN overlay must provide both routing and bridging functions on the leaf nodes. Connections from leaf nodes to hosts are via Layer 2 ports, and the leaf nodes must provide local Layer 2 switching. The leaf nodes must support dedicated MLAG or EVPN-LAG functionality to allow multihomed hosts across two or more leaf nodes. Furthermore, to bootstrap host learning, ARP requests must first be flooded across the overlay.
[0016] Multi-chassis link aggregation (MLAG) and Ethernet virtual private network link aggregation (EVPN-LAG)-based multihoming introduce the need to support complex Layer 2 functions on leaf nodes. Host MAC addresses must be learned in the data plane of any leaf node and synchronized across all redundant leaf nodes. Host ARP bindings must be learned via ARP flooding on any leaf node and synchronized across all redundant leaf nodes. Furthermore, physical loops resulting from MLAG topologies must be prevented for broadcast, unknown-unicast (BUM) traffic via a split-horizon filtering mechanism across redundant leaf nodes. Furthermore, a designated forwarder election mechanism must be supported on leaf nodes to prevent duplicate BUM packets from being forwarded to multihomed hosts. While EVPN procedures are specified for each of the above, the overall implementation and operational complexity of an EVPN-IRB-based solution may not be desirable for all use cases.
[0017] To facilitate understanding of the present disclosure, some of the many networking computing devices and protocols will be described.
[0018] In a computer network environment, networking devices such as switches or routers can be used to transmit information from one destination to a final destination. In one embodiment, data packages and messages may be generated at a first location, such as a computer in a person's home. The data packages and messages may be generated by an individual interacting with a web browser to request or provide information from a remote server accessible via the Internet. For example, the data packages and messages may be information entered by an individual into a form accessible on an Internet-connected web page. The data packages and messages may need to be transmitted to a remote server that is geographically far away from the individual's computer. There is likely no direct communication between the router in the individual's home and the remote server. Thus, the data packages and messages must "hop" through different networking devices before reaching their final destination at the remote server. The router in the individual's home must transmit the data packages and messages through several different devices connected to the Internet to determine the route the data packages and messages must take to reach their final destination at the remote server.
[0019] Switches (also called switching hubs, bridging hubs, or MAC bridges) create networks. Most internal networks use switches to connect computers, printers, phones, cameras, lights, and servers within a building or campus. Switches act as controllers, allowing networked devices to communicate efficiently with each other. Switches connect devices on computer networks using packet switching, which allows data to be received, processed, and forwarded to the destination device. Network switches are multi-port network bridges that process and forward data at the data link layer (Layer 2) of the Open Systems Interconnection (OSI) model using hardware addresses. Some switches can also process data at the network layer (Layer 3) by incorporating additional routing functionality. Such switches are commonly called Layer 3 switches or multi-layer switches.
[0020] Routers connect networks. Switches and routers perform similar functions, but they perform different functions on a network. Routers are network devices that forward data packets between computer networks. Routers perform traffic directing functions on the Internet. Data sent over the Internet, such as web pages, email, or other forms of information, is sent in the form of data packets. Packets are typically forwarded from one router to another through the networks that make up an internetwork (e.g., the Internet), until they finally reach their destination node. Routers are connected to multiple data lines from different networks. When a data packet arrives on one of the lines, the router reads the network address information in the packet to determine its ultimate destination. The router then uses information in the router's routing table or routing policy to send the packet to the next network on its path. A BGP speaker is a router that has the Border Gateway Protocol (BGP) enabled.
[0021] A customer edge router (CE router) is a router located on the customer premises that provides the interface between the customer's LAN and the provider's core network. CE routers, provider routers, and provider edge routers are components of the Multiprotocol Label Switching architecture. Provider routers are located in the core of a provider's or carrier's network. Provider edge routers are located at the edge of the network. Customer edge routers connect to provider edge routers, which in turn connect to other provider edge routers.
[0022] A routing table or Routing Information Base (RIB) is a data table stored in a router or network computer that lists routes to specific network destinations. Routing tables may include route metrics such as distance and weight. Routing tables contain information about the topology of the network in the immediate vicinity of the router where they are stored. Building a routing table is the primary purpose of a routing protocol. Static routes are entries created in a routing table by non-automatic means; they are fixed and are not the result of some network topology discovery procedure. A routing table can contain at least three information fields, including network ID, metric, and next hop fields. The network ID is the destination subnet. The metric is the routing metric for the path the packet will be sent. The route proceeds toward the gateway with the smallest metric. The next hop is the address of the next station on the packet's way to its final destination. A routing table can also contain the quality of service associated with the route, a link to a list of filtering criteria associated with the route, the interface of an Ethernet card, etc.
[0023] In hop-by-hop routing, each routing table lists, for every reachable destination, the address of the next device along the path to that destination, or next hop. Assuming the routing tables are consistent, an algorithm that relays packets to the destination's next hop is sufficient to deliver data anywhere in the network. Hop-by-hop routing is a feature of the IP internetwork layer and the Open Systems Interconnection (OSI) model.
[0024] Some network communication systems are large, enterprise-level networks with thousands of processing nodes. These thousands of processing nodes share bandwidth from multiple Internet Service Providers (ISPs) and can handle large volumes of Internet traffic. Such systems can be very complex and must be properly configured to provide acceptable Internet performance. If the system is not properly configured for optimal data transmission, Internet access speeds can be slowed and system bandwidth consumption and traffic can increase. To address this issue, a set of services can be implemented to eliminate or mitigate these concerns. This set of services is also known as routing control.
[0025] One embodiment of the routing control mechanism is comprised of hardware and software. The routing control mechanism monitors all outgoing traffic through connections with Internet Service Providers (ISPs). The routing control mechanism assists in selecting the optimal path for efficient transmission of data. The routing control mechanism calculates the performance and efficiency of all ISPs and can select only those ISPs that perform optimally in the applicable area. The route control device can be configured according to predefined parameters regarding cost, performance, and bandwidth.
[0026] Equal cost multipath (ECMP) routing is a routing scheme in which next-hop packet forwarding to a single destination can occur over multiple "best paths." The multiple best paths are equivalent based on a routing metric calculation. Because routing is a hop-by-hop decision limited to a single router, multipath routing can be used with many routing protocols. Multipath routing can significantly increase bandwidth by load-balancing traffic across multiple paths. However, ECMP routing has many known problems when deploying the strategy in practice. Disclosed herein are systems, methods, and devices for improved ECMP routing.
[0027] To promote an understanding of the principles underlying the present disclosure, reference will be made to illustrated embodiments and specific language will be used to describe the same, without intending to limit the scope of the present disclosure. Any changes and further modifications of the features of the present disclosure exemplified herein, and any additional applications of the principles of the present disclosure exemplified herein, will be readily apparent to those skilled in the art based on the present disclosure, and are encompassed by the appended claims.
[0028] Before disclosing and describing structures, systems, and methods for tracking the lifecycle of objects in a network computing environment, it is to be understood that the present disclosure is not limited to the particular structures, configurations, process steps, and materials disclosed herein, as such structures, configurations, process steps, and materials may vary. Furthermore, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure is limited only by the claims and their equivalents.
[0029] In describing and claiming the subject matter of the present disclosure, the following terminology will be used in accordance with the definitions set out below.
[0030] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0031] As used herein, the terms "comprise," "have," "include," "featuring," and their grammatical equivalents are non-exclusive or open terms that do not exclude additional, unrecited elements or method steps.
[0032] As used herein, the phrase "consisting of" and its grammatical equivalents excludes any element or step not recited in a claim.
[0033] As used herein, the phrase "consisting essentially of" and its grammatical equivalents limit the scope of a claim to the materials or steps specified, and to materials or steps that do not materially affect the basic and novel properties or characteristics of the claimed disclosure.
[0034] The following description refers to the drawings, in which FIG. 1 is a schematic diagram of a system 100 for connecting devices to the Internet. The system 100 is presented as background information for explaining certain concepts described herein. The system 100 includes multiple local area networks 110 connected by a switch 106. Each of the multiple local area networks 110 is connectable to one another via the public Internet by a router 112. The exemplary system 100 shown in FIG. 1 has two local area networks 110. However, more local area networks 110 may be connected to one another via the public Internet. Each local area network 110 includes multiple computing devices 108 connected to one another via the switch 106. The multiple computing devices 108 may include, for example, desktop computers, laptops, printers, servers, etc. The local area network 110 can communicate with other networks via the public Internet by a router 112. The router 112 connects the multiple networks to one another. The router 112 is connected to an Internet service provider 102, which is connected to one or more network service providers 104. The network service provider 104 communicates with other local network service providers 104 as shown in FIG.
[0035] The switch 106 connects devices within the local area network 110 using packet switching, allowing data to be received, processed, and forwarded to a destination device. The switch 106 can be configured to receive data from a computer, for example, destined for a printer. The switch 106 can receive the data, process the data, and send the data to the printer. The switch 106 may be a Layer 1 switch, a Layer 2 switch, a Layer 3 switch, a Layer 4 switch, a Layer 7 switch, etc. A Layer 1 network device forwards data but does not manage the traffic passing through the device. An example of a Layer 1 network device is an Ethernet hub. A Layer 2 network device is a multi-port device that processes and forwards data at the data link layer (Layer 2) using hardware addresses. A Layer 3 switch can perform some or all of the functions typically performed by a router. However, some network switches are limited to supporting a single type of physical network, usually Ethernet, while a router may support different types of physical networks on different ports.
[0036] The router 112 is a networking device that forwards data packets between computer networks. In the exemplary system 100 shown in FIG. 1, the router 112 forwards data packets between local area networks 110. However, the router 112 does not necessarily need to forward data packets between local area networks 110, but may also be used to forward data packets between wide area networks, etc. The router 112 performs traffic direction functions on the Internet. The router 112 may have interfaces for different types of physical layer connections, such as copper cable, optical fiber, or wireless transmission. The router 112 can support different network layer transmission standards. Each network interface can be used to forward data packets from one transmission system to another. The router 112 may also be used to connect two or more logical groups of computer devices, called subnets, each with a different network prefix. As shown in FIG. 1, the router 112 can provide connectivity within an enterprise, between an enterprise and the Internet, or between Internet service provider networks. Some routers 112 may be configured to interconnect various Internet service providers or may be used within large corporate networks. Smaller routers 112 typically provide connectivity for home and office networks to the Internet. The routers 112 shown in Figure 1 can represent any router suitable for network transmission, such as an edge router, a subscriber edge router, an inter-provider border router, a core router, an Internet backbone, a port forwarding router, a voice / data / fax / video processing router, etc.
[0037] An Internet Service Provider (ISP) 102 is an organization that provides services for access, use, or participation in the Internet. ISPs 102 can be organized in a variety of forms, such as commercial, community-owned, non-profit, or privately owned. Internet services typically provided by ISPs 102 include Internet access, Internet transit, domain name registration, web hosting, Usenet services, and colocation. ISP 102 as shown in FIG. 1 can represent any suitable ISP, such as a hosting ISP, transit ISP, virtual ISP, toll-free ISP, wireless ISP, etc.
[0038] A network service provider (NSP) 104 is an organization that provides bandwidth or network access by providing direct Internet backbone access to Internet service providers. A network service provider may provide access to a network access point (NAP). A network service provider 104 is also called a backbone provider or Internet provider. Network service providers 104 can include telecommunications providers, data carriers, wireless communication providers, Internet service providers, and cable television operators that offer high-speed Internet access. Network service providers 104 can also include information technology providers.
[0039] The system 100 shown in Figure 1 is merely exemplary, and many different configurations and systems can be constructed for transmitting data between networks and computing devices. Because networking is highly customizable, it is desirable to provide greater customizability in determining the best route for transmitting data between computers or networks. In light of the above, disclosed herein are systems, methods, and devices for offloading best path computation to an external device to provide greater customizability in determining a best path algorithm that best suits a particular grouping of computers or a particular enterprise.
[0040] FIG. 2 is a schematic diagram of an architecture 200 having a centralized gateway known in the prior art. The architecture 200 includes a spine node and a leaf node in a leaf / spine network topology. Inter-subnet routing is performed at the spine node or the aggregation layer. The leaf node is connected to multiple virtual machines. The centralized gateway architecture 200 may include a spine layer, a leaf layer, and an access layer. There may be an L2-L3 boundary at the aggregation layer, and there may be a datacenter perimeter at the core layer. In the architecture shown in FIG. 2, the spine layer including spine S1 and spine S2 may function as the core layer. Above the leaf (aggregation) layer, there is Layer 2 extension via Ethernet Virtual Private Network (EVPN).
[0041] The centralized gateway architecture 200 has many drawbacks. It has an L2-L3 boundary at the spine layer, which can create a scale bottleneck. This further introduces a single point of failure for the architecture 200. Furthermore, there are numerous operational complexities at the leaf nodes of the centralized gateway architecture 200. One of the complexities is that the architecture 200 must deal with the unpredictable nature of MAC and ARP age-outs, probes, silent hosts, and mobility. Furthermore, the architecture 200 must be configured to flood overlay ARPs and populate both IP and MAC forwarding entries for all hosts across the overlay bridge. Furthermore, the architecture 200 must be configured to synchronize MAC addresses and ARPs for MLAGs and perform filtering and selection for MLAGs.
[0042] FIG. 3 is a schematic diagram of an architecture 300 with a distributed anycast L3 gateway on a distributed anycast router, as known in the prior art. The architecture 300 provides first-hop gateway functionality for workloads. As a result, the service layer on the L2-L3 boundary is serviced by the distributed anycast router on the leaf node. In other words, all inter-subnet VPN information traffic from workload host virtual machines is routed at the distributed anycast router. Virtual machine mobility and flexible workload placement are achieved by stretching a Layer 2 overlay across the routed network fabric. Intra-subnet traffic across the stretched Layer 2 domain is overlay bridged on the leaf node. The distributed anycast router may provide EVPN-IRB services for directly connected host virtual machines, route all overlay inter-subnet VPN traffic, and bridge all overlay intra-subnet VPN traffic across the routed fabric underlay.
[0043] Architecture 300 also illustrates an exemplary architecture for providing a fully routed network fabric. However, the architecture 300 illustrated in FIG. 3 has several drawbacks. For example, it again requires the EVPN overlay to provide both routing and bridging functions on the distributed anycast routers. Furthermore, connections from the distributed anycast routers to hosts are via Layer 2 ports, and the leaf nodes must provide local Layer 2 switching. The leaf nodes must support dedicated MLAG or EVPN-LAG functionality to enable multihomed hosts across two or more distributed anycast routers. Furthermore, to bootstrap host learning, ARP requests must first be flooded across the overlay.
[0044] In particular, MLAG or EVPN-LAG-based multihoming requires the distributed anycast router to support complex Layer 2 functions. For example, host MAC addresses must be learned in the data plane of one of the leaf nodes and synchronized across all redundant distributed anycast routers. Similarly, host ARP bindings must be learned via ARP flooding at one of the distributed anycast routers and synchronized across all redundant distributed anycast routers. It is necessary to prevent physical loops resulting from MLAG topologies for BUM traffic via a split-horizon filtering mechanism across redundant distributed anycast routers. Furthermore, it is necessary to support a designated forwarder election mechanism on the distributed anycast router to prevent duplicate BUM packets from being forwarded to multihomed hosts.
[0045] While EVPN procedures are defined for each of the above, the overall implementation and operational complexity of an EVPN-IRB-based solution may not be desirable for all use cases. Therefore, an alternative solution is provided and described herein. For example, it eliminates the need to support overlay bridging functionality in the distributed anycast router. Similarly, this architecture provides IP unicast inter- and intra-subnet VPN connectivity, VM mobility, and flexible workload placement across stretched IP subnets, while eliminating the need for Layer 2 MLAG connections and associated complex procedures between the distributed anycast router and hosts, and the need for data plane and ARP-based host learning on the distributed anycast router.
[0046] FIG. 4 is a schematic diagram of an architecture 400 for a host-routed overlay with deterministic host learning and localized integrated routing and bridging on host machines. The architecture 400 includes a virtual customer edge (CE) router with leaf node links that serve as Layer 3 interfaces. There is no Layer 2 PE-CE. The Layer 3 subnet addresses of the leaf nodes on the virtual CE router are locally scoped and are not redistributed in Border Gateway Protocol (BGP) routing. As shown, the virtual CE router resides on a bare metal server and communicates with one or more virtual machines that are also located on the bare metal server. In one embodiment, the virtual CE router and one or more virtual machines are located on the same physical bare metal server. The virtual CE router communicates with one or more virtual machines that are also located on the bare metal server. The virtual CE router communicates with one or more leaf nodes in a leaf-spine network topology. Each leaf node that communicates with the virtual CE router has a dedicated communication line to the virtual CE router, as shown in FIGS. 4-10. The Layer 2 and Layer 3 boundary (L2 / L3 boundary) exists at the virtual CE router.
[0047] In the example shown in Figure 4, two bare metal servers each have one virtual CE router. The bare metal servers also contain multiple virtual machines. One virtual CE router has two subnets, including anycast gateway MACs (AGMs) 10.1.1.1 / 24 and 12.1.1.1 / 24. The anycast gateway MAC (AGM) box is internal to the virtual CE router. The interface between the virtual CE router and one or more virtual machines on the bare metal server can be created using a Linux hypervisor. The virtual CE router includes physical connections to leaf nodes. In the example shown in Figure 4, one virtual CE router includes physical connections to leaf nodes L1 and L2. This is shown by the physical connection to leaf L1, which has address 15.1.1.1, terminating at a virtual CE router with address 15.1.1.2. This is further shown by the physical connection to leaf L2, which has address 14.1.1.1, terminating at a virtual CE router with address 14.1.1.2. This is further illustrated by the physical connection to leaf L3, with address 15.1.1.1, terminating at a virtual CE router with address 15.1.1.2, and by the physical connection to leaf L4, with address 14.1.1.1, terminating at a virtual CE router with address 14.1.1.2.
[0048] The architectures shown in Figures 4-10 have many advantages over architectures known in the prior art, including those shown in Figures 2 and 3. Traditionally, a Layer 2 link is created between the server and the leaf node. This Layer 2 link causes many problems in architectures known in the prior art. The architectures shown in Figures 4-10 move the L2-L3 boundary to the virtual CE router, avoiding many of the problems known to exist in the architectures shown in Figures 2 and 3. For example, by placing the virtual CE router and virtual machines on the same server box, functionality is localized, eliminating the Layer 2 link from the server to the leaf node known in the prior art. The architectures shown in Figures 4-10 introduce Layer 3 router links from the bare metal server to each of multiple leaf nodes. This simplifies the leaf node functionality and achieves the same functionality without requiring Layer 2 termination at each leaf node.
[0049] The architecture 400 includes spine nodes S1 and S2 that communicate with leaf nodes L1, L2, L3, and L4. The leaf node L1 has an address of 15.1.1.1, the leaf node L2 has an address of 14.1.1.1, the leaf node L3 has an address of 15.1.1.1, and the leaf node L4 has an address of 14.1.1.1. Nodes L1 and L2 communicate with a virtual customer edge (CE) router. The virtual CE router is deployed on a bare metal server along with virtual machines. Nodes L3 and L4 communicate with the virtual customer edge (CE) router. The L2-L3 boundary exists at the virtual CE router level. The virtual CE router communicates with multiple virtual machines, including VM-a, VM-b, VM-c, VM-d, VM-e, VM-f, VM-g, and VM-h, as shown in the figure.
[0050] Host VM IP-MAC bindings are traditionally learned on the first-hop gateway via ARP. However, in stretched subnet scenarios, ARP-based learning requires flooding ARP requests across the overlay to bootstrap host learning at the local virtual CE router. This requires a Layer 2 overlay flood domain. To avoid relying on Layer 2 overlay and ARP-based host learning across leaf nodes, host VM IP and MAC bindings configured on the VM external interface must be passively learned by the L3DL on the server by exposing them to the hypervisor. This ensures that directly connected host VM bindings are always known in advance. This also eliminates the need for glean processing and flooding. Local VM IP host routes (overlay host routes) are relayed from the hypervisor to the leaf nodes by the L3DL.
[0051] Architecture 400 deploys a small virtual CE router on the server that terminates Layer 2 from the host. The virtual CE router provides IRB services for local host virtual machines. The virtual CE router routes all traffic to external host virtual machines over ECMP Layer 3 links to leaf nodes via a default route. The virtual CE router learns the IP and MAC addresses of host virtual machine interfaces at host boot time. Local VM IP host routes (overlay host routes) are relayed from the hypervisor to the leaf nodes by L3DL. The leaf nodes advertise local host routes to remote leaf nodes via Border Gateway Protocol (BGP).
[0052] In one embodiment, subnet stretching is enabled through host routing of both intra-subnet and inter-subnet flows at the leaf nodes. Virtual CE routers are configured as proxy ARPs that host route intra-subnet flows through the leaf nodes. The virtual CE routers can be configured with the same anycast gateway IP address and MAC address everywhere. Architecture 400 provides EVPN host mobility procedures that are applied at the leaf nodes. Architecture 400 enables flexible workload placement and virtual machine mobility across stretched subnets.
[0053] In architecture 400, end-to-end host routing is configured at power-on. Inter-subnet and intra-subnet traffic flows are enabled across stretched subnets via end-to-end host routing, which does not rely on non-deterministic data plane and ARP-based learning.
[0054] Architecture 400 provides local host learning to virtual customer edge routers (virtual CE routers) over L3DL. EVPN host routing is performed across the overlay. EVPN has Layer-3 host mobility and Layer-3 mass withdraw. Architecture 400 provides private subnets that are not redistributed into Border Gateway Protocol (BGP). In architecture 400, the first-hop anycast gateway provides local IRB services.
[0055] The virtual CE router can be configured as an ARP proxy for all directly connected host virtual machines, allowing it to route inter-subnet and intra-subnet traffic flows. The virtual CE router can be configured with a default route pointing to the set of upstream leaf nodes to which the virtual CE router is multihomed. Configuring the virtual CE router with the same anycast gateway MAC on all bare metal servers enables host virtual machine mobility across the DC fabric. The virtual CE router does not need to redistribute server-side connected subnets into the DC-side routing protocol, thereby avoiding IP addressing overhead on the server links. The virtual CE router can reside in a hypervisor provided as a default gateway for host virtual machines in a VLAN. The virtual CE router can also be a separate router VM, providing the router VM as the default gateway for host virtual machines in a VLAN.
[0056] In one embodiment, leaf nodes must advertise host routes learned from locally attached virtual CE routers across the EVPN overlay as EVPN RT-5. EVPN mobility procedures may be extended to EVPN RT-5 to enable host virtual machine mobility. EVPN mass withdraw procedures may be extended to EVPN RT-5 to enable faster convergence.
[0057] The embodiments described herein eliminate the need to support overlay bridging functionality on leaf nodes. Furthermore, the embodiments described herein eliminate the need for Layer 2 MLAG connections and associated complex procedures between leaf nodes and hosts. Furthermore, the embodiments described herein eliminate the need for data plane and ARP-based host learning on leaf nodes. The embodiments disclosed herein achieve these benefits while also providing IP unicast inter- and intra-subnet VPN connectivity, virtual machine mobility, and flexible workload placement across stretched IP subnets.
[0058] Embodiments of the present disclosure separate local Layer 2 switching and IRB functions from leaf nodes and localize them to a small virtual CE router on a bare metal server. This is achieved by running a small virtual router VM on the bare metal server, which acts as a first-hop gateway for host virtual machines and provides local IRB switching between virtual machines local to the bare metal server. This virtual router operates as a traditional CE router that can be multihomed to multiple leaf nodes via Layer 3 routed interfaces on the leaf nodes. Leaf nodes in the fabric function as pure Layer 3 VPN PE routers without any Layer 2 bridging or IRB functionality. To enable flexible placement and mobility of Layer 3 endpoints across the DC overlay while providing optimal routing, traffic can be host-routed on the leaf nodes rather than subnet-routed. This is the case with EVPN-IRB.
[0059] Figure 5 is a schematic diagram of an architecture 400 showing host learning at bootup. Host virtual machine routes learned via L3DL are installed in the FIB and point to the virtual CE router as the next hop. In the absence of multitenancy (no VPN), host virtual machine routes are advertised to remote leaf nodes via BGP global routing. In the presence of multitenancy, host virtual machine routes are advertised to remote leaf nodes via BGP-EVPN RT-5 using VPN encapsulation such as VXLAN or MPLS. Therefore, any other routing protocol can also be deployed as an overlay routing protocol.
[0060] In one embodiment, subnet extension across the overlay is enabled by routing intra-subnet traffic at the virtual CE router and then at the leaf nodes. To terminate Layer 2 at the virtual CE router, the virtual CE router must be configured as an ARP proxy for the host virtual machine subnets, allowing both intra-subnet and inter-subnet traffic to be routed at the virtual CE router and then at the leaf nodes.
[0061] In one embodiment, to avoid IP addressing overhead, the IP subnets used for the Layer 3 links to the servers should be locally scoped, in other words, the server-side connected subnets should not be redistributed into the northbound routing protocol.
[0062] In one embodiment, to achieve multi-tenancy, the overlay Layer 3 VLAN / IRB interface on the virtual CE router's first-hop gateway must be attached to a tenant VRF. Furthermore, routed VXLAN / VNI encapsulation is used between the virtual CE router and the leaf node to separate multiple tenant traffic. Furthermore, to install the L3DL overlay host route sent to the leaf node in the correct VPN / VRF table on the leaf node, the L3DL overlay host must also include a Layer 3 VNI ID. This VNI ID is used by the leaf node to identify the route and install it in the correct VRF.
[0063] Figure 6 shows a protocol 600 for a PE distributed anycast router. Figure 6 also shows the forwarding tables of leaf nodes L1 and L2. In protocol 600, host virtual machine routes learned via L3DL are installed in the FIB, pointing to the virtual CE router next hop in the resulting FIB state. In the absence of multitenancy (no VPN), host virtual machine routes are advertised to the remote distributed anycast router via BGP global routing. In the presence of multitenancy, host virtual machine routes are advertised to the remote distributed anycast router via BGP-EVPN RT-5 using VPN encapsulation such as VXLAN or MPLS.
[0064] 7 is a schematic diagram of a protocol 700 for a virtual CE router as an ARP proxy. In protocol 700, subnet extension across the overlay is enabled by routing intra-subnet traffic at the virtual CE router and then at the leaf nodes. Layer 2 termination at the hypervisor virtual CE router requires that the virtual CE router be configured as an ARP proxy for the host virtual machine subnets, allowing both intra-subnet and inter-subnet traffic to be routed at the virtual CE router and then at the distributed anycast router.
[0065] Figures 8A and 8B show the protocols for server-local flows. Figure 8A shows the protocol for intra-subnet flows, and Figure 8B shows the protocol for inter-subnet flows. The virtual CE router is configured with a default route that points to a set of upstream leaf nodes that are multihomed.
[0066] In the protocol shown in Figure 8A, host-to-host flows are local to the bare metal server protocol, and once the virtual CE router learns all host VM adjacencies and is configured as an ARP proxy, both inter-subnet and intra-subnet flows across host VMs local to the bare metal server are terminated at Layer 2 at the virtual CE router and routed to the local destination host VM. In Figure 8A, the default gateway (GW) sending the object to 12.1.1.1 via the anycast gateway (AGW) is 12.1.1.2 → veth2, anycast gateway medium access control (AGW_MAC).
[0067] In the inter-subnet flow protocol shown in Figure 8B, host-to-host flows are local to the bare metal server protocol, and once the virtual CE router learns all host VM adjacencies and is configured as an ARP proxy, both inter-subnet and intra-subnet flows across host VMs local to the bare metal server are terminated at Layer 2 at the virtual CE router and routed to local destination host VMs.
[0068] Figures 9A and 9B show protocols for overlay flows: Figure 9A shows the protocol for an intra-subnet overlay flow from 12.1.1.4 to 12.1.1.2; Figure 9B shows the protocol for an inter-subnet overlay flow from 12.1.1.4 to 10.1.1.2.
[0069] In the protocol shown in Figure 9B, host-to-host overlay inter-subnet flows pass through the leaf node. In this protocol, the virtual CE router is configured with a default route pointing to a set of upstream distributed anycast routers that are multihomed. All outbound inter-subnet and intra-subnet traffic from the host VM is routed by this virtual CE to the upstream leaf node over L3 ECMP links instead of hashing through a Layer 2 LAG as shown in Figures 9A and 9B. The leaf node operates as a pure Layer 3 router without any Layer 2 bridging or IRB functionality. Horizontal (east-west) flows across multiple servers connected to the same leaf node are routed locally by the leaf node to the destination virtual CE next hop.
[0070] The protocol shown in Figures 9A and 9B can include host-to-host overlay flows via leaf nodes. In this protocol, horizontal flows (both inter-subnet and intra-subnet) through multiple servers connected to different distributed anycast routers are routed from the virtual CE router to the local leaf node via a default route, and then at the leaf node, routed to the destination / next-hop leaf node via the routed overlay based on host routes learned via EVPN RT-5. Leaf-to-leaf routing is based on aggregated or subnet routes instead of host routes only if the subnet is not stretched across the overlay. Vertical flows (destinations outside the DC) can be routed via per-VRF default routes on the leaf node towards the border leaf / DCI GW.
[0071] Another protocol, shown in Figure 10, identifies leaf node-server link failures. This protocol can be used as an alternative redundancy mechanism. Routed backup links are configured between leaf nodes and pre-programmed as backup failure paths for server-facing overlay host routes. Upon a leaf node-server link failure, the backup path is activated in a prefix-agnostic manner for a given VRF associated with the same VLAN (VNI) encapsulation.
[0072] In the protocol shown in Figure 10, outbound traffic from the host VM converges following a link failure as a result of the virtual CE router removing the failed path from its default route ECMP path set. Inbound traffic from the DC overlay converges as a result of L3DL-learned host routes being removed and withdrawn from the affected leaf nodes. However, this convergence depends on the host route scale. To achieve prefix-independent convergence, the EVPN mass withdraw mechanism must be extended to IP host routes. An ESI configuration is associated with a set of Layer 3 addresses from distributed anycast routers in a redundancy group. Local ESI reachability is advertised to remote distributed anycast routers via the EAD RT-1 per ESI. As shown in Figure 10, via this route, forwarding indirection is established at the remote distributed anycast router, allowing fast convergence with a single RT-1 withdrawal from the local distributed anycast router after an ESI failure.
[0073] The protocol shown in Figure 10 may be implemented in the case of a server link failure. Following a link failure, outbound traffic from the host virtual machine may converge as a result of the virtual CE router removing the failed path from its default route ECMP path set. Inbound traffic from the DC overlay converges as a result of L3DL-learned host routes being removed and withdrawn from the affected leaf nodes. However, this convergence depends on the host route scale. To achieve prefix-independent convergence, the EVPN mass-withdraw mechanism must be extended to IP host routes. An ESI configuration is associated with a set of Layer 3 links from a leaf node in a redundancy group. Local ESI reachability is advertised to remote leaf nodes via an EAD RT-1 per ESI. Through this route, indirect forwarding is established at the remote leaf nodes, enabling fast convergence with a single RT-1 withdrawal from the local leaf node after an ESI failure.
[0074] All outbound inter-subnet and intra-subnet traffic from the host virtual machines is routed by this virtual CE router over Layer 3 ECMP links to upstream leaf nodes instead of being hashed across a Layer 2 LAG. The leaf nodes operate as pure Layer 3 routers without any Layer 2 bridging or IRB functionality. Horizontal flows across multiple servers connected to the same leaf node are routed locally by the leaf node to the destination virtual CE router next hop.
[0075] Horizontal flows (both inter-subnet and intra-subnet) through multiple servers connected to different leaf nodes are routed from the virtual CE router to the local leaf node via a default route, and then at the leaf node, routed to the destination via the routed overlay. The next-hop leaf node is based on the host route learned via EVPN RT-5. Leaf-node to leaf-node routing is based on aggregated or subnet routes instead of host routes only if the subnet is not stretched across the overlay.
[0076] Vertical flows to destinations outside the DC may be routed towards the border leaf via a per-VRF default route on the leaf node.
[0077] Another protocol provides a simplification and scaling implementation. In this implementation, in response to a first-hop GW localized on the virtual CE, the leaf node does not install host MAC routes, saving forwarding resources on the distributed anycast router. Furthermore, with default routing on the leaf node, the virtual CE maintains only adjacencies to the local host VMs for each bare-metal server. This completely removes all bridging and MLAG functionality from the leaf node, simplifying operations and improving reliability. The use of deterministic protocol-based host route learning between the virtual CE and the distributed anycast router eliminates the need for EVPN aliasing procedures on the distributed anycast router, and the use of deterministic protocol-based host route learning between the virtual CE and the leaf node eliminates the need for ARP flooding across the overlay. Furthermore, the use of deterministic protocol-based host route learning between the virtual CE and the distributed anycast router eliminates the need for unknown unicast flooding. Finally, Layer 3 ECMP links between the virtual CE and the leaf node eliminate the need for EVPN DF election and split horizontal filtering procedures.
[0078] 11 illustrates a block diagram of an exemplary computing device 1100. The computing device 1100 can be used to perform various procedures as described herein. In one embodiment, the computing device 1100 can function to perform the functions of an asynchronous object manager and can execute one or more application programs. The computing device 1100 can be any of a wide variety of computing devices, such as a desktop computer, an in-dash computer, a vehicle control system, a notebook computer, a server computer, a handheld computer, a tablet computer, etc.
[0079] The computing device 1100 includes one or more processors 1102, one or more memory devices 1104, one or more interfaces 1106, one or more mass storage devices 1108, one or more input / output devices 1110, and a display device 1130, all connected to a bus 1112. The processor 1102 includes one or more processors or controllers that execute instructions stored on the memory devices 1104 and / or the mass storage device 1108. The processor 1102 may also include various types of computer-readable media, such as cache memory.
[0080] The memory device 1104 includes a variety of computer-readable media, such as volatile memory (e.g., random access memory (RAM) 1114) and / or non-volatile memory (e.g., read-only memory (ROM) 1116). The memory device 1104 may also include re-writable ROM, such as flash memory.
[0081] The mass storage device 1108 includes various computer-readable media such as magnetic tape, magnetic disks, optical disks, solid-state memory (e.g., flash memory), etc. As shown in Figure 11, an exemplary mass storage device is a hard disk drive 1124. The mass storage device 1108 may also include various drives to allow reading from and / or writing to various computer-readable media. The mass storage device 1108 includes removable media 1126 and / or non-removable media.
[0082] Input / output (I / O) devices 1110 include various devices that allow data and / or other information to be input to or retrieved from computing device 1100. I / O devices 1110 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, etc.
[0083] Display device 1130 includes any type of device capable of displaying information to one or more users of computing device 1100. Display device 1130 may include, for example, a monitor, a display terminal, a video projection device, etc.
[0084] Interface 1106 includes various interfaces that allow computing device 1100 to interact with other systems, devices, or computing environments. Interface 1106 may include any number of different network interfaces 1120, such as interfaces to a local area network (LAN), a wide area network (WAN), a wireless network, and the Internet. Other interfaces include a user interface 1118 and a peripheral device interface 1122. Interface 1106 may also include one or more user interface elements 1118. Additionally, interface 1106 may include one or more peripheral interfaces, such as an interface for a printer, a pointing device (such as a mouse, trackpad, or any suitable user interface now known to those skilled in the art or any suitable user interface hereafter developed), a keyboard, etc.
[0085] The bus 1112 allows the processor 1102, memory device 1104, interface 1106, mass storage device 1108, and I / O device 1110 to communicate with each other and with other devices or components connected to the bus 1112. The bus 1112 may represent one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE bus, a USB bus, etc.
[0086] Although programs and other executable program components are illustrated herein as separate blocks for purposes of illustration, such programs and components may reside at various times in different storage components of computing device 1100 and be executed by processor 1102. Alternatively, the systems and procedures described herein may be implemented in hardware or a combination of hardware, software, and / or firmware. For example, one or more application specific integrated circuits (ASICs) may be programmed to perform one or more of the systems and procedures described herein.
[0087] The foregoing description has been presented for purposes of illustration and description, and the present disclosure is not intended to be limited to the precise forms set forth herein. Many modifications and variations are possible in light of the above teachings. Additionally, any or all of the above variations can be combined in any manner to form further combinations of the present disclosure.
[0088] Moreover, although specific embodiments of the present disclosure have been described and illustrated, the present disclosure is not limited to the specific forms or arrangements of parts so described and illustrated. The scope of the present disclosure is defined by the claims of this application or any future claims, separate applications based on this application, and their equivalents, if any. [Example]
[0089] The following examples relate to further embodiments.
[0090] Example 1 is a system. The system includes a virtual customer edge router on a server and a host routed overlay including a plurality of host virtual machines. The system includes routed uplinks from the virtual customer edge router to one or more of a plurality of leaf nodes. The system is configured such that the virtual customer edge router provides localized integrated routing and bridging (IRB) services to the plurality of host virtual machines in the host routed overlay.
[0091] In a second embodiment, the host-routed overlay is an Ethernet virtual private network (EVPN) host.
[0092] A third embodiment is the system according to any one of the first and second embodiments, in which the host routed overlay has EVPN Layer 3 mobility.
[0093] Example 4 is the system according to any one of Examples 1 to 3, wherein the virtual customer edge router is a first-hop anycast gateway for one or more of the plurality of leaf nodes.
[0094] Example 5 is the system of any of Examples 1 to 4, wherein the virtual customer edge router routes traffic to the external leaf node via equal-cost multipath (ECMP) routing links to the leaf node.
[0095] A sixth embodiment is the system according to any one of the first to fifth embodiments, in which the virtual customer edge router is configured as a proxy address resolution protocol (ARP) for host routing of intra-subnet flows in a host-routed overlay.
[0096] A seventh embodiment is the system according to any one of the first to sixth embodiments, wherein the routed uplink from the virtual customer edge router to one or more of the plurality of leaf nodes is a Layer 3 interface.
[0097] An eighth embodiment is the system according to any one of the first to seventh embodiments, in which the virtual customer edge router stores addresses locally and does not redistribute addresses in Border Gateway Protocol (BGP) routing.
[0098] Example 9 is a system described in any of Examples 1 to 8, wherein the virtual customer edge router includes memory that stores one or more of local Internet Protocol (IP) entries for the host-routed overlay, Media Access Control (MAC) entries for the host-routed overlay, or a default ECMP route to the host-routed overlay.
[0099] Example 10 is the system of any of Examples 1 to 9, wherein the host routed overlay is configured to perform host routed over stretched subnets.
[0100] An eleventh embodiment is the system according to any one of the first to tenth embodiments, in which the virtual customer edge router is deployed on a single tenant physical server.
[0101] Example 12 is a system described in any of Examples 1 to 11, wherein the virtual customer edge router is a virtual router virtual machine running on a single tenant physical server and is configured to operate as a first-hop gateway for one or more of the multiple leaf nodes.
[0102] A thirteenth embodiment is the system according to any one of the first to twelfth embodiments, in which the plurality of host virtual machines are arranged on a single tenant physical server.
[0103] A fourteenth embodiment is the system according to any one of the first to thirteenth embodiments, in which the virtual customer edge router is multihomed to a plurality of distributed anycast routers via a layer 3 routed interface on the distributed anycast router.
[0104] Example 15 is the system of any of Examples 1 to 14, wherein the virtual customer edge router is configured to learn local host virtual machine routes without relying on a gleaning process and ARP-based learning.
[0105] Example 16 is the system according to any one of Examples 1 to 15, wherein the virtual customer edge router is further configured to advertise the local host virtual machine route to the directly connected distributed anycast router.
[0106] Example 17 is a system described in any of Examples 1 to 16, wherein the virtual customer edge router is configured to learn one or more IP bindings and MAC bindings of multiple host virtual machines via link state over ethernet (LSoE).
[0107] Example 18 is a system described in any of Examples 1 to 17, wherein the virtual customer edge router has a memory and is configured to store in the memory the adjacency relationships of one or more host virtual machines local to the same bare metal server on which the virtual customer edge router is located.
[0108] Example 19 is a system described in any of Examples 1 to 18, which includes a distributed anycast router and a virtual customer edge router configured to perform deterministic protocol-based host route learning between the virtual customer edge router and the distributed anycast router.
[0109] Example 20 is the system of any of Examples 1 to 19, wherein the routed uplink from the virtual customer edge router to one or more of the plurality of host machines is a Layer 3 Equal-Cost Multipath (ECMP) routed link.
[0110] Example 21 is the system according to any one of Examples 1 to 20, wherein one or more of the plurality of leaf nodes includes a virtual private network-virtual routing and forwarding (VPI-VRF) table.
[0111] Example 22 is a system described in any of Examples 1 to 21, wherein one or more of the plurality of leaf nodes further includes a Layer 3 virtual network identifier (VNI) that is used by one or more of the plurality of leaf nodes to install routes in the correct virtual routing and forwarding table.
[0112] It should be noted that any features of the above-described configurations, examples, and embodiments may be combined in a single embodiment, including any combination of features from any of the configurations, examples, and embodiments disclosed herein.
[0113] The various features disclosed herein provide important advantages and advances in the art, and the following claims are illustrative of some of these features.
[0114] In the foregoing detailed description of the present disclosure, for purposes of streamlining the disclosure, various features of the disclosure are grouped together in a single embodiment. This method of disclosure is not to be interpreted as reflecting an intention that the claimed disclosure requires more features than are expressly recited in each claim. That is, inventive aspects may feature fewer than all features of a single foregoing disclosed embodiment.
[0115] The above-described arrangements are merely illustrative of the application of the principles of the present disclosure. Many modifications and alternative arrangements may be devised by those skilled in the art without departing from the spirit and scope of the present disclosure, and the appended claims are intended to cover all such modifications and arrangements.
[0116] Thus, while the present disclosure has been illustrated in the drawings and described in detail above, it will be apparent to those skilled in the art that numerous modifications, including but not limited to variations in size, material, shape, form, function, operation, assembly, and use, may be made thereto without departing from the principles and concepts described herein.
[0117] Additionally, the functions described herein may be implemented in one or more of hardware, software, firmware, digital components, or analog components, where appropriate. For example, one or more application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) may be programmed to implement one or more of the systems and procedures described herein. Specific terms are used throughout the description and claims to refer to particular system components. Those skilled in the art will recognize that components may be referred to by different names. This document does not intend to distinguish between components that differ in name but function.
[0118] The foregoing description has been presented for purposes of illustration and description. The present disclosure is not limited to the precise forms set forth herein. Many modifications and variations are possible in light of the above teachings. Moreover, any or all of the foregoing variations can be combined in any manner to form further combinations of the present disclosure.
[0119] Moreover, although specific embodiments of the present disclosure have been described and illustrated, the present disclosure is not limited to the specific forms or arrangements of parts so described and illustrated. The scope of the present disclosure is defined by the claims of this application or any future claims, separate applications based on this application, and their equivalents, if any.
Claims
1. Multiple bare metal servers and a plurality of leaf nodes in a leaf-spine network topology; a host-routed overlay including at least one host virtual machine; a leaf server link established between a first bare metal server of the plurality of bare metal servers and a first leaf node of the plurality of leaf nodes; a routed backup link established between two leaf nodes of the plurality of leaf nodes; A system comprising: The routed backup link is pre-programmed as a backup failure path for an overlay host route on the first bare metal server side when a failure occurs in the leaf server link; The system, wherein the plurality of bare metal servers and corresponding plurality of virtual customer edge routers are configured with the same anycast gateway medium access control (MAC) that enables host virtual machine mobility across the host-routed overlay.
2. 10. The system of claim 1, wherein the routed backup link is activated on the leaf server link without utilizing a prefix associated with virtual local area network (VLAN) encapsulation.
3. 2. The system of claim 1, wherein the plurality of virtual customer edge routers remove a leaf server link from an equal-cost multipath routing path set when the leaf server link fails.
4. 10. The system of claim 1, wherein an Ethernet Virtual Private Network (EVP) withdraw mechanism is extended to Internet Protocol (IP) host routes to enable activation of the routed backup link on the leaf server link without utilizing a prefix.
5. 10. The system of claim 1, further comprising a plurality of distributed anycast routers in a redundancy group, wherein an edge-side include (ESI) configuration is associated with a set of Layer 3 links in the redundancy group.
6. 6. The system of claim 5, wherein local reachability of an ESI configuration is advertised to one or more remote leaf nodes outside the redundancy group via a per-ESI Ethernet Access Direct (EAD) that routes traffic.
7. 7. The system of claim 6, wherein indirect forwarding is established with the one or more remote leaf nodes via the routed backup link to enable convergence from the first leaf node after an ESI configuration failure.
8. 10. The system of claim 1, wherein the routed backup link is configured to route traffic from a Layer 3 equal-cost multipath (ECMP) link across to an upstream leaf node in the leaf-spine network topology instead of hashing across a Layer 2 link aggregation (LAG).
9. 10. The system of claim 8, wherein the upstream leaf nodes function as Layer 3 routers without any Layer 2 bridging or integrated routing and bridging (IRB) functionality.
10. 2. The system of claim 1, wherein the plurality of leaf nodes do not install any host MAC routes.
11. each of the plurality of bare metal servers includes a virtual customer edge router and a host virtual machine; 2. The system of claim 1, wherein the system further comprises a routed uplink from a first virtual customer edge router on the first bare metal server to one or more of the plurality of leaf nodes.
12. further including a distributed anycast router, wherein the first virtual customer edge router is configured to perform host route learning between the first virtual customer edge router and the distributed anycast router; 12. The system of claim 11, wherein the first virtual customer edge router provides localized integrated routing and bridging (IRB) for a first host virtual machine on the first bare metal server.
13. 13. The system of claim 12, wherein the first virtual customer edge router communicates with the first host virtual machine on the first bare metal server, and the first virtual customer edge router includes memory for storing adjacency relationships of the first host virtual machine local to the first bare metal server.
14. The leaf-spine network topology comprises: the plurality of leaf nodes; a plurality of spine nodes; Including, 2. The system of claim 1, wherein each of the plurality of spine nodes includes a link to each of the plurality of leaf nodes.
15. the first bare metal server includes a first virtual customer edge router; the first virtual customer edge router includes a plurality of subnets; 2. The system of claim 1, wherein the leaf server link between the first bare metal server and the first leaf node is established by a physical connection between the first leaf node and the first virtual customer edge router.
16. 16. The system of claim 15, wherein the leaf server link fails when the physical connection between the first leaf node and the first virtual customer edge router is deactivated.
17. the first bare metal server includes a first virtual customer edge router; the leaf server link is established between the first leaf node and the first virtual customer edge router; 2. The system of claim 1, wherein the first virtual customer edge router includes a memory for storing one or more of a local Internet Protocol (IP) entry for the host-routed overlay, a Media Access Control (MAC) entry for the host-routed overlay, or a default ECMP route for the host-routed overlay.
18. 20. The system of claim 17, wherein the first virtual customer edge router is a first-hop gateway for the first leaf node before the leaf-server link fails.
19. The host-routed overlay performs host routing for extended subnets; 10. The system of claim 1, wherein the host-routed overlay is an Ethernet Virtual Private Network (EVPn) host.
Citation Information
Cited By
Gas turbine engine and fuel cell assembly
US12584445B2