Cloud Native Software-Defined Network Architecture for Multiple Clusters

By running the SDN controller manager on a central cluster, merging configuration nodes and control nodes, the complexity and limitations of the existing SDN architecture in life cycle management, resource analysis, configuration management, and CLI interfaces are solved, and efficient and scalable management of cloud-native networks are achieved.

CN115941457BActive Publication Date: 2025-07-01HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210760750.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-31
Filing Date
2022-06-30
Publication Date
2025-07-01
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing software-defined network (SDN) architecture has complexity and limitations in life cycle management, resource analysis, configuration management, and command line interface (CLI) interfaces, making it difficult to achieve efficient and scalable network management on cloud-native.

Method used

Using a cloud-native SDN architecture deployed by multi-clusters, it can achieve better life management and security license management of SDN controller manager, configuration node and control node by running SDN controller manager on a central cluster and combining configuration nodes and control nodes.

Benefits of technology

It realizes the life cycle management of the SDN architecture, the efficiency of resource analysis components, the scalability of configuration management and the optimization of CLI-based interfaces, and improves the management efficiency and scalability of cloud-native networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115941457B_ABST
    Figure CN115941457B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a cloud-native software-defined network architecture for multiple clusters. In an example, a network controller for an SDN architecture system includes: processing circuitry for a central cluster of one or more computing nodes; a configuration node configured to be executed by the processing circuitry; and a control node configured to be executed by the processing circuitry. The configuration node includes a custom API server to process requests for operations on custom resources of the SDN architecture configuration. Each custom resource of the SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system. In response to detecting an event for an instance of a first custom resource among the custom resources, the control node obtains configuration data regarding the instance of the first custom resource and configures a corresponding instance of a configuration object in a workload cluster of a second one or more computing nodes. The first computing node may be different from the second computing node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 17 / 657,603, filed Mar. 31, 2022, which claims the benefit of Indian Provisional Patent Application No. 202141044924, filed Oct. 4, 2021, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to a virtualized computing infrastructure and, more particularly, to cloud native networking. Background Art

[0004] In a typical cloud data center environment, there is a large collection of interconnected servers that provide computing and / or storage capabilities for running various applications. For example, a data center can include a facility that hosts applications and services for users (i.e., customers of the data center). For example, a data center can host all infrastructure equipment such as networking and storage systems, redundant power supplies, and environmental controls. In a typical data center, storage systems are interconnected with clusters of application servers via a high - speed switching fabric provided by one or more layers of physical network switches and routers. More sophisticated data centers provide user - support equipment located in various physical hosting facilities for infrastructure spread around the world.

[0005] Virtualized data centers are becoming a core foundation of modern information technology (IT) infrastructure. Specifically, modern data centers have a widely - used virtualized environment in which virtual hosts such as virtual machines or containers, herein also referred to as virtual running elements, are deployed on and run on the underlying computing platform of physical computing devices.

[0006] Virtualization within a data center or any environment that includes one or more servers can provide several advantages. One advantage is that virtualization can provide significantly increased efficiency. Since underlying physical computing devices (i.e., servers) have become increasingly powerful with the emergence of multi - core microprocessor architectures with a large number of cores per physical CPU, virtualization has become easier and more efficient. A second advantage is that virtualization provides effective control of the computing infrastructure. Since physical computing resources in a cloud - based computing environment, for example, become replaceable resources, the provision and management of the computing infrastructure become easier. Thus, enterprise IT staff generally like the virtualized computing clusters in data centers due to their management advantages in addition to the efficiency and increased return on investment (ROI) provided by virtualization.

[0007] Containerization is a virtualization solution based on operating system-level virtualization. A container is a lightweight and portable runtime element for applications that are isolated from each other and from the host. Since containers are not tightly coupled to the host hardware computing environment, applications can be bundled into container images and run as a single lightweight package on any host or virtual host that supports the underlying container architecture. Thus, containers solve the problem of how to make software work in different computing environments. Containers provide the guarantee of continuous operation from one computing environment to another virtual or physical environment.

[0008] Due to the inherently lightweight nature of containers, a single host can typically support more container instances than traditional virtual machines (VMs). In general, short-lived containers can be created and moved more efficiently than VMs, and containers can also be managed as a group of logically related elements (sometimes called a "pod" for some orchestration platforms, such as Kubernetes). These container characteristics influence the requirements for container networking solutions: the network should be agile and scalable. In the same computing environment, VMs, containers, and bare-metal servers may need to coexist and enable communication between applications deployed in multiple ways. The container network should also be unaware of working with multiple types of orchestration platforms used to deploy containerized applications.

[0009] Managing the computing infrastructure for the deployment and infrastructure of application runtimes may involve two main roles: (1) Orchestration - to automate the deployment, scaling, and operation of applications across a cluster of hosts and provide the computing infrastructure, which may include container-centric computing infrastructure; and (2) Network management - to create virtual networks in the network infrastructure to enable packetized communication between applications running in virtual runtime environments such as containers or VMs and between applications running in legacy (e.g., physical) environments. Software-defined networking helps with network management. Summary of the Invention

[0010] Overall, a technology regarding a cloud-native SDN architecture deployed using multiple clusters is described. In some embodiments, the SDN architecture may include data plane elements implemented in compute nodes and network devices such as routers or switches, and the SDN architecture may further include a network controller for creating and managing virtual networks. The SDN architecture configuration and control plane are designed to have an extended cloud-native software with a container-based microservices architecture that supports in-service upgrades. Configuration nodes for the configuration plane may be implemented to expose custom resources. These custom resources for SDN architecture configuration may include configuration elements typically exposed by the network controller, however, the configuration elements may be merged with Kubernetes native / built-in resources together to support a unified intent model exposed by an aggregated API layer and implemented by Kubernetes controllers and custom resource controllers that work together to reconcile the actual state and the desired state of the SDN architecture.

[0011] In a multi-cluster deployment of the SDN architecture, configuration nodes and control nodes are deployed to a central cluster and centrally manage the configuration and control of one or more workload clusters. However, the data plane is distributed among the workload clusters. Each of the workload clusters and the central cluster may implement the data plane of the cluster using similar component microservices. A dedicated and different SDN controller manager running on the central cluster of each workload cluster creates custom resources regarding the SDN architecture configuration in the central cluster to be managed by the configuration nodes and configured by the control nodes in the corresponding workload clusters.

[0012] This technology may provide one or more technical advantages. For example, using an SDN controller manager running on the central cluster and merging the configuration nodes and control nodes into a single central cluster facilitates better lifecycle management (LCM) of the SDN controller manager, configuration nodes, and control nodes and facilitates better and more manageable handling of security and licensing by consolidating these tasks into a single central cluster.

[0013] As other embodiments, a cloud-native SDN architecture can address limitations in conventional SDN architectures related to the complexity of lifecycle management, the high enforceability of resource analysis components, the scalability limitations of configuration management, and the lack of a command-line interface (CLI)-based interface. For example, the network controller for an SDN architecture is a cloud-native, lightweight distributed application with simplified installation coverage. This also facilitates the easy and modular upgrade of the various component microservices of the configuration nodes and control nodes of the configuration and control planes. The technology can further support optional cloud-native monitoring (telemetry) and user interfaces, a high-performance data plane for containers that uses a DPDK-based virtual router to connect to a DPDK-enabled container pool (pod), and in some cases, a cloud-native configuration management that leverages the configuration frameworks of existing orchestration platforms such as Kubernetes or Openstack. As a cloud-native architecture, the network controller has scalability and elasticity to address and support multiple clusters. In some cases, the network controller can also support the scalability and performance requirements of key performance indicators (KPIs).

[0014] In an embodiment, a network controller for a software-defined network (SDN) architecture system is provided. The network controller includes: processing circuitry for a central cluster of one or more computing nodes; a configuration node configured to be executed by the processing circuitry; and a control node configured to be executed by the processing circuitry; wherein the configuration node includes a custom application programming interface (API) server to process requests for operations on custom resources regarding SDN architecture configuration; wherein each custom resource regarding SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system; and wherein, in response to detecting an event for an instance of a first custom resource among the custom resources, the control node is configured to obtain configuration data regarding the instance of the first custom resource and configure a corresponding instance of a configuration object in a workload cluster of a second one or more computing nodes, wherein the first one or more computing nodes in the central cluster are distinct from the second one or more computing nodes in the workload cluster.

[0015] In an embodiment, the method includes: a custom application programming interface (API) server implemented by a configuration node of a network controller for a software-defined network (SDN) architecture system processes a request for an operation on a custom resource regarding the SDN architecture configuration, wherein each custom resource regarding the SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system, wherein the network controller operates on a central cluster of one or more first computing nodes; a control node of the network controller detects an event regarding an instance of a first custom resource among the custom resources; and in response to the detection of the event regarding the instance of the first custom resource, the control node obtains configuration data regarding the instance of the first custom resource and configures a corresponding instance of the configuration object in a workload cluster of one or more second computing nodes, wherein the one or more first computing nodes of the central cluster are distinct from the one or more second computing nodes of the workload cluster.

[0016] In an embodiment, a non-transitory computer-readable medium includes instructions for causing a processing circuit: a custom application programming interface (API) server implemented by a configuration node of a network controller for a software-defined network (SDN) architecture system processes a request for an operation on a custom resource regarding the SDN architecture configuration, wherein each custom resource regarding the SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system, wherein the network controller operates on a central cluster of one or more first computing nodes; a control node of the network controller detects an event regarding an instance of a first custom resource among the custom resources; and in response to the detection of the event regarding the instance of the first custom resource, the control node obtains configuration data regarding the instance of the first custom resource and configures a corresponding instance of the configuration object in a workload cluster of one or more second computing nodes, wherein the one or more first computing nodes of the central cluster are distinct from the one or more second computing nodes of the workload cluster.

[0017] In the accompanying drawings and the following description, details of one or more embodiments of the present disclosure are set forth. Other features, objects, and advantages will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a block diagram illustrating an exemplary computing infrastructure in which embodiments of the techniques described herein may be implemented.

[0019] Figure 2 is a block diagram illustrating an embodiment of a cloud-native SDN architecture for a cloud-native network according to the techniques of the present disclosure.

[0020] Figure 3It is a block diagram showing another view of the components of the SDN architecture 200 according to the technology of the present disclosure in further detail.

[0021] Figure 4 It is a block diagram showing exemplary components of the SDN architecture according to the technology of the present disclosure.

[0022] Figure 5 It is a block diagram of an exemplary computing device according to the technology described in the present disclosure.

[0023] Figure 6 It is a block diagram of an exemplary computing device that operates as a computing node for one or more clusters of an SDN architecture system according to the technology of the present disclosure.

[0024] Figure 7A It is a block diagram showing the control / routing plane for the underlying network and overlay network configuration of an SDN architecture according to the technology of the present disclosure.

[0025] Figure 7B It is a block diagram showing a configured virtual network using a tunnel-connected container pool configured in the underlying network according to the technology of the present disclosure.

[0026] Figure 8 It is a block diagram showing an example of a custom controller for custom resources for an SDN architecture configuration according to the technology of the present disclosure.

[0027] Figure 9 It is a block diagram showing an exemplary flow of creating, monitoring, and coordinating between custom resource types that depend on different custom resource types.

[0028] Figure 10 It is a block diagram showing a multi-cluster deployment of a cloud-native SDN architecture according to the technology of the present disclosure.

[0029] Figure 11 It is a flowchart showing an exemplary mode of operation for a multi-cluster deployment of an SDN architecture.

[0030] Throughout the description and the drawings, like reference characters denote like elements. Detailed Description

[0031] Figure 1FIG. 8 is a block diagram of an exemplary computing infrastructure 8 that may implement the techniques described herein. Current implementations of software-defined network (SDN) architectures for virtual networks pose challenges to cloud-native adoption due to, for example, the complexity of lifecycle management, the high mandatory requirements for resource analysis components, the scalability limitations of configuration modules, and the lack of a command-line interface (CLI)-based interface such as kubectl. Computing infrastructure 8 includes the cloud-native SDN architecture system described herein, which addresses these challenges and modernizes the telecom cloud-native era. Exemplary use cases of the cloud-native SDN architecture include 5G mobile networks and cloud and enterprise cloud-native use cases. The SDN architecture may include data plane elements implemented in computing nodes (e.g., server 12) and network devices such as routers or switches, and the SDN architecture may also include an SDN controller (e.g., network controller 24) for creating and managing virtual networks. The SDN architecture configuration and control plane are designed as horizontally scalable cloud-native software with a container-based microservices architecture that supports hot upgrades.

[0032] Accordingly, SDN architecture components are microservices, and contrary to existing network controllers, the SDN architecture assumes a base container orchestration platform for managing the lifecycle of SDN architecture components. The container orchestration platform is used to build SDN architecture components; the SDN architecture uses cloud-native monitoring tools that can be integrated with cloud-native options provided by customers; the SDN architecture provides a declarative way of resources (i.e., custom resources) using an aggregated API of SDN architecture objects. SDN architecture upgrades can follow cloud-native patterns, and the SDN architecture can leverage Kubernetes constructs such as Multus, authentication & authorization, cluster API, KubeFederation, KubeVirt, and Kata Containers. The SDN architecture can support a data plane development kit (DPDK) container pool, and the SDN architecture is capable of scaling to support Kubernetes with virtual network policies and global security policies.

[0033] For service providers and enterprises, the SDN architecture automates network resource configuration and orchestration to dynamically create highly scalable virtual networks and link virtualized network functions (VNFs) with physical network functions (PNFs) to form differentiated service chains on demand. The SDN architecture can be integrated with orchestration platforms (e.g., orchestrator 23) such as Kubernetes, OpenShift, Mesos, OpenStack, VMware vSphere, and can be integrated with service provider operation support systems / business support systems (OSS / BSS).

[0034] Typically, one or more data centers 10 provide an operating environment for applications and services of customer sites 11 (shown as "Customer 11"), and the customer sites 11 have one or more customer networks coupled to the data centers via a service provider network 7. For example, each data center 10 may host infrastructure devices such as networking and storage systems, redundant power supplies, and environmental controls. The service provider network 7 is coupled to a public network 15 that may represent one or more networks managed by other providers, and thus, may form part of a large-scale public network infrastructure (e.g., the Internet). For example, the public network 15 may represent a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a Layer 3 virtual private network (VPN), an Internet Protocol (IP) intranet operated by a service provider that operates the service provider network 7, an enterprise IP network, or some combination thereof.

[0035] Although the customer sites 11 and the public network 15 are primarily shown and described as edge networks of the service provider network 7, however, in some embodiments, the customer sites 11 and one or more networks in the public network 15 may be tenant networks within any data center 10. For example, a data center 10 may host multiple tenants (customers) each associated with one or more virtual private networks (VPNs), where each tenant may implement one or more customer sites 11.

[0036] The service provider network 7 provides packet-based connectivity to the attached customer sites 11, data centers 10, and the public network 15. The service provider network 7 may represent a network held and operated by a service provider to interconnect multiple networks. The service provider network 7 may implement Multiprotocol Label Switching (MPLS) forwarding, and in this instance, may be referred to as an MPLS network or an MPLS backbone. In some instances, the service provider network 7 represents multiple interconnected autonomous systems that provide services from one or more service providers, such as the Internet.

[0037] In some embodiments, each data center 10 may represent one of a plurality of geographically distributed network data centers, and the plurality of geographically distributed network data centers may be connected to each other via the service provider network 7, dedicated network links, dark fiber, or other connections. As Figure 1As shown in the example, data center 10 may include facilities that provide network services to customers. Customers of the service provider may be general entities such as enterprises, governments, or individuals. For example, a network data center may host network services for several enterprises and end users. Other exemplary services may include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. Although shown as a discrete edge network of service provider network 7, elements in data center 10 such as one or more physical network functions (PNFs) or virtualized network functions (VNFs) may be included within the core of service provider network 7.

[0038] In this example, data center 10 includes storage and / or computing servers (or "nodes") interconnected via a switching fabric 14 provided by one or more layers of physical network switches and routers, and servers 12A - 12X (referred to herein as "servers 12") are depicted as being coupled to top-of-rack (TOR) switches 16A - 16N. Servers 12 are computing devices and may also be referred to herein as "compute nodes", "hosts", or "host devices". Although only server 12A coupled to TOR switch 16A is shown in detail in Figure 1 , data center 10 may include multiple additional servers coupled to other TOR switches 16 of data center 10.

[0039] In the example shown, switching fabric 14 includes interconnected top-of-rack (TOR) (or other "leaf") switches 16A - 16N (collectively "TOR switches 16") of the distribution layer coupled to chassis (or "spine" or "core") switches 18A - 18M (collectively "chassis switches 18"). Although not shown, for example, data center 10 may also include one or more non-edge switches, routers, hubs, gateways, security devices (such as firewalls, intrusion detection, and / or intrusion prevention devices), servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices. Data center 10 may also include one or more physical network functions (PNFs), such as physical firewalls, load balancers, routers, route reflectors, broadband network gateways (BNGs), mobile core network elements, and other PNFs.

[0040] In this example, the TOR switch 16 and the chassis switch 18 provide redundant (multi-homed) connections to the IP fabric 20 and the service provider network 7 for the servers 12. The chassis switch 18 aggregates traffic flows and provides connections between the TOR switches 16. The TOR switch 16 can be a network device that provides Layer 2 (MAC) and / or Layer 3 (e.g., IP) routing and / or switching functions. Each of the TOR switch 16 and the chassis switch 18 can include one or more processors and memories and be capable of running one or more software processes. The chassis switch 18 is coupled to the IP fabric 20, which can perform Layer 3 routing to route network traffic between the data center 10 and the customer site 11 through the service provider network 7. The switching fabric of the data center 10 is merely an example. For example, other switching fabrics can have more or fewer switching layers. The IP fabric 20 can include one or more gateway routers.

[0041] The term "packet flow", "traffic flow", or simply "flow" refers to a collection of packets that originate from a specific source device or endpoint and are sent to a specific destination device or endpoint. For example, a single flow of packets can be identified by a 5-tuple: <source network address, destination network address, source port, destination port, protocol>. Generally, this 5-tuple identifies the packet flow corresponding to the received packet. An n-tuple refers to any n items extracted from the 5-tuple. For example, a 2-tuple of a packet can refer to the combination of <source network address, destination network address> or <source network address, source port> of the packet.

[0042] Each of the servers 12 can represent a compute server or a storage server. For example, each server 12 can represent a computing device configured to operate according to the techniques described herein, such as an x86 processor-based server. The servers 12 can provide network function virtualization infrastructure (NFVI) for the NFV architecture.

[0043] Any one of the servers 12 can be configured with virtual execution elements such as a container pool or virtual machines by virtualizing the resources of the server to provide a certain degree of isolation between one or more processes (applications) running on the server. "Hypervisor-based", or "hardware-level", or "platform" virtualization refers to creating virtual machines each including a guest operating system for running one or more processes. Generally, the virtual machine provides a virtualized / guest operating system for executing applications in an isolated virtual environment. Since the virtual machines are virtualized from the physical hardware of the host server, the executed applications are isolated from the host's hardware and other virtual machines. Each virtual machine can be configured with one or more virtual network interfaces for communicating on a corresponding virtual network.

[0044] A virtual network refers to a logical construct implemented on top of a physical network. The virtual network can be used to replace VLAN-based isolation and provide multi-tenancy in a virtualized data center, such as Data Center 10. Each tenant or application can have one or more virtual networks. Unless explicitly permitted by the security policy, each virtual network must be isolated from all other virtual networks.

[0045] The virtual network can use the Data Center 10 gateway router ( Figure 1 not shown in ) to connect to a physical Multiprotocol Label Switching (MPLS) Layer 3 Virtual Private Network (L3VPN) and an Ethernet Virtual Private Network (EVPN) network and extend over the physical Multiprotocol Label Switching (MPLS) Layer 3 Virtual Private Network (L3VPN) and Ethernet Virtual Private Network (EVPN) network. The virtual network can also be used to implement Network Function Virtualization (NFV) and service chaining.

[0046] The virtual network can be implemented using various mechanisms. For example, each virtual network can be implemented as a Virtual Local Area Network (VLAN), a Virtual Private Network (VPN), etc. The virtual network can also be implemented using two networks - a physical underlying network composed of an IP fabric 20 and a switching fabric 14 and a virtual overlay network. The role of the physical underlying network is to provide an "IP fabric" that provides unicast IP connectivity from any physical device (server, storage device, router, or switch) to any other physical device. The underlying network can provide a unified low-latency, non-blocking, high-bandwidth connection from any point in the network to any other point in the network.

[0047] As further described below with respect to the virtual router 21 (shown as and also referred to herein as "vRouter21"), the virtual router running in the server 12 creates a virtual overlay network on top of the physical underlying network using its own dynamic "tunnel" mesh. For example, these overlay tunnels can be MPLS over GRE / UDP tunnels, or VXLAN tunnels, or NVGRE tunnels. For virtual machines or other virtual execution elements, the underlying physical routers and switches cannot store any state of each tenant, such as any Media Access Control (MAC) address, IP address, or policy. For example, the forwarding tables of the underlying physical routers and switches can only contain the IP prefixes or MAC addresses of the physical server 12. (The gateway router or switch that connects the virtual network to the physical network is an exception and can contain tenant MAC or IP addresses.)

[0048] The virtual router 21 of server 12 typically contains per-tenant state. For example, the virtual router may contain a separate forwarding table (routing instance) for each virtual network. This forwarding table contains IP prefixes (in the case of layer 3 overlay) or MAC addresses (in the case of layer 2 overlay) of virtual machines or other virtual execution elements (e.g., a pool of containers of a container). A single virtual router 21 does not need to contain all the IP prefixes or all the MAC addresses of all the virtual machines in the entire data center. A given virtual router 21 only needs to contain those routing instances that are locally present on server 12 (i.e., having at least one virtual execution element present on server 12).

[0049] "Container-based" or "operating system" virtualization refers to virtualization where the operating system runs multiple isolated systems on a single machine (virtual or physical). These isolated systems represent containers such as those provided by the open-source DOCKER container application or by CoreOS Rkt ("Rocket"). Like virtual machines, each container is virtualized and can be kept isolated from the host and other containers. However, unlike virtual machines, each container can dispense with a separate operating system and instead provide an application suite and application-specific libraries. Typically, containers are run by the host as isolated user-space instances and containers can share the operating system and common libraries with other containers running on the host. Thus, containers may require lower processing power, storage, and network resources than virtual machines. A group of one or more containers can be configured to share one or more virtual network interfaces for communicating on a corresponding virtual network.

[0050] In some examples, containers are managed by their host kernel to allow resource (CPU, memory, block I / O, network, etc.) limiting and prioritization in some cases using namespace isolation features without starting any virtual machines. The namespace isolation feature allows the application (e.g., a given container) view of the operating environment to be completely isolated, including process trees, networking, user identifiers, and mounted file systems. In some embodiments, containers can be deployed according to the Linux Containers (LXC), an operating system-level virtualization method that uses a single Linux kernel to run multiple isolated Linux systems (containers) on a control host.

[0051] Server 12 hosts virtual network endpoints for one or more virtual networks operating on the physical network represented here by IP fabric 20 and switching fabric 14. Although mainly described with respect to a data center-based switching network, other physical networks such as service provider network 7 can lie beneath one or more virtual networks.

[0052] Each server 12 may host one or more virtual execution elements, each of the one or more virtual execution elements having at least one virtual network endpoint of one or more virtual networks configured in the physical network. The virtual network endpoints of the virtual network may represent one or more virtual execution elements sharing a virtual network interface of the virtual network. For example, the virtual network endpoint may be a virtual machine, one or more sets of containers (e.g., a container pool), or another virtual execution element such as a Layer 3 endpoint of a virtual network. The term "virtual execution element" encompasses virtual machines, containers, and other virtualized computing resources that provide at least partially independent runtime environments for applications. The term "virtual execution element" may also encompass a container pool of one or more containers. A virtual execution element may represent an application workload. As Figure 1 shown, server 12A hosts one virtual network endpoint in the form of a container pool 22 having one or more containers. However, in practice, given the hardware resource limitations of a given server 12, the server 12 may execute for multiple virtual execution elements. Each virtual network endpoint may use one or more virtual network interfaces to perform packet I / O or other processing on packets. For example, the virtual network endpoint may use one virtual hardware component (e.g., an SR-IOV virtual function) enabled by NIC 13A to perform packet I / O and receive / send packets on one or more communication links with TOR switch 16A. Other examples of virtual network interfaces are described below.

[0053] Each of the servers 12 includes at least one network interface card (NIC) 13, and each of the NICs 13 includes at least one interface that exchanges packets with the TOR switch 16 via a communication link. For example, server 12A includes NIC 13A. Any of the NICs 13 can provide one or more virtual hardware components 21 for virtualized input / output (I / O). The virtual hardware components of the I / O can be virtualizations of physical NICs (“physical functions”). For example, in single root I / O virtualization (SR-IOV) described in the Peripheral Component Interconnect Special Interest Group SR-IOV specification, the PCIe physical function of a network interface card (or “network adapter”) is virtualized to provide one or more virtual network interfaces as “virtual functions” for use by the corresponding endpoints running on the server 12. Thus, the virtual network endpoints can share the same PCIe physical hardware resources and the virtual functions are examples of the virtual hardware components 21. As another example, one or more of the servers 12 can implement, for example, Virtio, a para-virtualization framework available for the Linux operating system, which provides emulated NIC functionality as a type of virtual hardware component to provide a virtual network interface to the virtual network endpoints. As another example, one or more of the servers 12 can implement Open vSwitch to perform distributed virtual multi-layer switching between one or more virtual network interface cards (vNICs) of the hosted virtual machines, where the vNIC can also represent a type of virtual hardware component that provides a virtual network interface to the virtual network endpoints. In some instances, the virtual hardware component is a virtual I / O (e.g., NIC) component. In some instances, the virtual hardware component is an SR-IOV virtual function. In some examples, any one of the servers 12 can implement a Linux bridge that emulates a hardware bridge and forwards packets between the virtual network interfaces of the server or between the virtual network interfaces of the server and the physical network interfaces of the server. For a Docker implementation of containers hosted by a server, a Linux bridge or other operating system bridge that runs on the server and switches packets between the containers can be referred to as a “Docker bridge”. As used herein, the term “virtual router” can encompass Contrail or Tungsten Fabric virtual routers, Open vSwitch (OVS), OVS bridges, Linux bridges, Docker bridges, or other devices and / or software located on a host device and performing switching, bridging, or routing packets between the virtual network endpoints of one or more virtual networks, where the virtual network endpoints are hosted by one or more of the servers 12.

[0054] Any NIC 13 may include an internal device switch that switches data between virtual hardware components associated with the NIC. For example, for an SR-IOV-enabled NIC, the internal device switch may be a virtual Ethernet bridge (VEB) that switches between SR-IOV virtual functions and correspondingly between endpoints configured to use SR-IOV virtual functions, where each endpoint may include a guest operating system. Alternatively, the internal device switch may be referred to as a NIC switch, or for SR-IOV implementations, as an SR-IOV NIC switch. The virtual hardware components associated with NIC 13A may be associated with Layer 2 destination addresses assigned by NIC 13A or a software process responsible for configuring NIC 13A. The physical hardware components (or "physical functions" in the case of SR-IOV implementations) are also associated with Layer 2 destination addresses.

[0055] Each of one or more servers 12 may include a virtual router 21 that executes one or more routing instances for a corresponding virtual network within data center 10 to provide a virtual network interface and route packets between virtual network endpoints. Each routing instance may be associated with a network forwarding table. Each routing instance may represent a virtual routing and forwarding instance (VRF) of an Internet Protocol - Virtual Private Network (IP-VPN). For example, packets received by the virtual router 21 of server 12A from the underlying physical network fabric of data center 10 (i.e., IP fabric 20 and switching fabric 14) may include an outer header to allow the physical network fabric to tunnel the payload or "inner packet" to the physical network address of the network interface card 13A of the server 12A on which the virtual router is executing. The outer header may include not only the physical network address of the network interface card 13A of the server, but also a virtual network identifier such as a VxLAN tag or a Multiprotocol Label Switching (MPLS) tag that identifies a virtual network and the corresponding routing instance executed by the virtual router 21. The inner packet includes an inner header having a destination network address that conforms to the virtual network addressing space of the virtual network identified by the virtual network identifier.

[0056] The virtual router 21 terminates the virtual network overlay tunnel and determines the virtual network of the received packet based on the tunnel encapsulation header of the packet, and forwards the packet to the appropriate destination virtual network endpoint of the packet. For example, for server 12A, for each packet outbound from a virtual network endpoint (e.g., container pool 22) hosted by server 12A, the virtual router 21 attaches a tunnel encapsulation header indicating the virtual network to the packet to generate an encapsulated or "tunneled" packet, and the virtual router 21 outputs the encapsulated packet via the overlay tunnel of the virtual network to a physical destination computing device such as another server 12. As used herein, the virtual router 21 may perform the operations of a tunnel endpoint to encapsulate internal packets originating from a virtual network endpoint to generate tunnel packets, and to decapsulate tunnel packets to obtain internal packets for routing to other virtual network endpoints.

[0057] In some examples, the virtual router 21 may be kernel-based and may run as part of the kernel of the operating system of server 12A.

[0058] In some examples, the virtual router 21 may be a virtual router that supports the Data Plane Development Kit (DPDK). In this example, the virtual router 21 uses DPDK as the data plane. In this mode, the virtual router 21 runs as a user space application linked to a DPDK library (not shown). This is the execution version of the virtual router and is typically used by telecommunications companies, where VNFs are typically DPDK-based applications. The execution of the virtual router 21 as a DPDK virtual router enables a throughput that is ten times higher than that of a virtual router operating as a kernel-based virtual router. The physical interface is used by the poll mode driver (PMD) of DPDK, rather than the interrupt-based driver of the Linux kernel.

[0059] A user-I / O (UIO) kernel module such as vfio or uio_pci_generic may be used to expose the registers of the physical network interface to user space so that the DPDK PMD can access the registers. When NIC 13A is bound to the UIO driver, NIC 13A is moved from the Linux kernel space to user space and thus the Linux OS no longer manages it and is no longer visible. Thus, the DPDK application (i.e., in this example, virtual router 21A) fully manages NIC 13. This includes packet polling, packet processing, and packet forwarding. The user packet processing steps may be performed by the DPDK data plane of the virtual router 21, and the kernel (in Figure 1(The kernel is not shown) participates to a limited extent or not at all. Compared with the interrupt mode, especially when the packet rate is high, the nature of this "polling mode" makes the packet processing / forwarding of the virtual router 21 DPDK data plane more efficient. There are limited interrupts or no interrupts during packet I / O, and there is context switching.

[0060] Additional details of an example of DPDK vRouter can be found in "DAY ONE: CONTRAIL DPDKvROUTER" by Kiran KN et al. of Juniper Networks, Inc., published in 2021, the entire content of which is incorporated herein by reference.

[0061] The computing infrastructure 8 implements an automation platform for automating the deployment, scaling, and operation of virtual execution elements across the servers 12 to provide a virtualized infrastructure for executing application workloads and services. In some examples, the platform can be a container orchestration system that provides a container-centric infrastructure to automate the deployment, scaling, and operation of containers, thereby providing a container-centric infrastructure. In the context of virtualized computing infrastructure, "orchestration" generally refers to the provisioning, scheduling, and management of host servers available to the orchestration platform, and the virtual execution elements and / or applications and services running on the virtual execution elements. Specifically, for example, container orchestration permits containers to be coordinated and refers to the deployment, management, scaling, and configuration of containers to host servers by a container orchestration platform. Exemplary instances of orchestration platforms include Kubernetes (container orchestration system), Docker swarm, Mesos / Marathon, OpenShift, OpenStack, VMware, and Amazon ECS.

[0062] The elements of the automation platform of the computing infrastructure 8 at least include the servers 12, the orchestrator 23, and the network controller 24. A cluster-based framework can be used to deploy containers to a virtualized environment, in which the cluster master node of the cluster manages the deployment and operation of containers to one or more cluster slave nodes of the cluster. The terms "master node" and "slave node" used herein cover different orchestration platform terms for simulating devices that distinguish the primary management element of the cluster from the primary container hosting device of the cluster. For example, the Kubernetes platform uses the terms "cluster master node" and "slave node", while the Docker Swarm platform refers to the cluster manager and cluster nodes.

[0063] The orchestrator 23 and the network controller 24 can run on separate computing devices or on the same computing device. Each of the orchestrator 23 and the network controller 24 can be a distributed application running on one or more computing devices. The orchestrator 23 and the network controller 24 can implement the respective master nodes of one or more clusters, each of the one or more clusters having one or more slave nodes (also referred to as "compute nodes") implemented by the respective servers 12.

[0064] Generally, for example, the network controller 24 controls the network configuration of the fabric of the data center 10 to establish one or more virtual networks for packetized communication between virtual network endpoints. The network controller 24 provides a controller that facilitates logical and in some cases physical centralization of the operation of one or more virtual networks within the data center 10. In some examples, the network controller 24 can operate in response to configuration inputs received from the orchestrator 23 and / or an administrator / operator. Additional information regarding the exemplary operation of the network controller 24 operating in conjunction with other devices or other software-defined networks of the data center 10 can be found in International Application No. PCT / US2013 / 044378, filed Jun. 5, 2013, and entitled "PHYSICAL PATH DETERMINATION FOR VIRTUAL NETWORK PACKET FLOWS", and in U.S. Patent Application No. 14 / 226,509, filed May 26, 2014, and entitled "Tunneled Packet Aggregation for Virtual Networks", the disclosures of which are hereby incorporated by reference in their entireties.

[0065] Generally, the orchestrator 23 controls the deployment, scaling, and operation of containers across a cluster of servers 12 and provides a computing infrastructure, which can include a container-centric computing infrastructure. The orchestrator 23, and in some cases, the network controller 24 can implement the respective cluster masters for one or more Kubernetes clusters. As an example, Kubernetes is a container management platform that provides portability across public and private clouds, where each cloud can provide virtualization infrastructure to the container management platform. Exemplary components of the Kubernetes orchestration system are described below with reference to Figure 3 FIG.

[0066] Kubernetes operates using various Kubernetes objects - entities that represent the state of a Kubernetes cluster. Kubernetes objects can include a name, namespace, labels, annotations, field selectors, and any combination of recommended labels. For example, a Kubernetes cluster can include one or more "namespace" objects. Each namespace of a Kubernetes cluster is isolated from other namespaces of the Kubernetes cluster. A namespace object can include at least one of organization, security, and performance of the Kubernetes cluster. As an example, a container pool can be associated with a namespace, and thus, associate the container pool with the characteristics (e.g., virtual network) of the namespace. This characteristic can organize multiple newly created container pools by associating them with a common set of characteristics. A namespace can be created based on namespace specification data that defines the characteristics of the namespace, including the namespace name. In one example, a namespace might be named "Namespace A", and each newly created container pool can be associated with the set of characteristics represented by "Namespace A". Additionally, Kubernetes includes a "default" namespace. If a newly created container pool does not specify a namespace, the newly created container pool can be associated with the characteristics of the "default" namespace.

[0067] Namespaces enable multiple users, user teams, or a single user with multiple applications to use a Kubernetes cluster. Additionally, each user, user team, or application can be isolated from every other user within the namespace. Thus, each user of the Kubernetes cluster within a namespace operates as if they were the only user of the Kubernetes cluster. Multiple virtual networks can be associated with a single namespace. Thus, a container pool belonging to a specific namespace has the ability to access each of the virtual networks associated with the namespace, including other container pools that serve as virtual network endpoints within the group of virtual networks.

[0068] In one example, Container Pool 22 is a Kubernetes container pool and an example of a virtual network endpoint. A container pool is a group of one or more logically related containers ( Figure 1(not shown), a shared memory for the container, and options on how to run the container. Alternatively, if instantiated for execution, the container pool may be referred to as a "container pool replica". Each container of the container pool 22 is an example of a virtual execution element. The containers of the container pool are always co-located on a single server, co-scheduled, and run in a shared context. The shared context of the container pool may be a collection of Linux namespaces, cgroups, and other isolation aspects. In the context of the container pool, individual applications may further apply sub-isolation. Typically, the containers within the container pool have a common IP address and port space and are able to detect each other via the localhost. Because they have a shared context, the containers within the container pool also communicate with each other using inter-process communication (IPC). Examples of IPC include SystemV semaphores or POSIX shared memory. Typically, containers that are members of different container pools have different IP addresses and, in the absence of a configuration to enable this feature, cannot communicate via IPC. Containers that are members of different container pools instead typically communicate with each other via the container pool IP address.

[0069] Server 12A includes a container platform 19 for running containerized applications such as those in container pool 22. The container platform 19 receives requests from the orchestrator 23 to obtain containers and host the containers in server 12A. The container platform 19 obtains and runs the containers.

[0070] A Container Network Interface (CNI) 17 configures virtual network interfaces for virtual network endpoints. An orchestrator 23 and a container platform 19 use the CNI 17 to manage the network of a container pool (including container pool 22). For example, the CNI 17 creates virtual network interfaces that connect the container pool to a virtual router 21 and enables the containers of the container pool to communicate via the virtual network interfaces to other virtual network endpoints on the virtual network. For example, the CNI 17 can insert virtual network interfaces for the virtual network into the network namespaces of the containers in the container pool 22 and configure (or request configuration of) virtual network interfaces for the virtual network in the virtual router 21 such that the virtual router 21 is configured to send packets received from the virtual network via the virtual network interfaces to the containers in the container pool 22 and send packets received via the virtual network interfaces on the virtual network from the containers in the container pool 22. The CNI 17 can assign network addresses (e.g., virtual IP addresses for the virtual network) and can set routes for the virtual network interfaces. In Kubernetes, by default all container pools can communicate with all other container pools without using Network Address Translation (NAT). In some cases, the orchestrator 23 and a network controller 24 create a service virtual network and a container pool virtual network shared by all namespaces, and allocate service and container pool network addresses from the namespaces respectively. In some cases, all container pools created in a Kubernetes cluster in all namespaces can communicate with each other, and network addresses for all container pools can be allocated from a container pool subnet specified by the orchestrator 23. When a user creates an isolated namespace for a container pool, the orchestrator 23 and the network controller 24 can create a new container pool virtual network and a new shared service virtual network for the new isolated namespace. Container pools created in the Kubernetes cluster in the isolated namespace obtain network addresses from the new container pool virtual network, and the corresponding services of the container pool obtain network addresses from the new service virtual network.

[0071] The CNI 17 can represent a library, plugin, module, runtime, or other executable code of the server 12A. The CNI 17 can at least partially conform to the Container Network Interface (CNI) specification or the rkt network proposal. The CNI 17 can represent Contrail, OpenContrail, Multus, Calico, cRPD, or other CNI. Alternatively, the CNI 17 can be referred to as a network plugin or a CNI plugin or a CNI instance. For example, a separate CNI can be invoked by the Multus CNI to establish different virtual network interfaces for the container pool 22.

[0072] The CNI 17 can be invoked by the orchestrator 23. For the purposes of the CNI specification, a container can be considered to be synchronized with a Linux network namespace. The corresponding unit depends on the specific container runtime implementation: for example, in an implementation of the application container specification such as rkt, each container pool runs in a unique network namespace. However, in Docker, there is typically a network namespace for each individual Docker container. For the purposes of the CNI specification, a network refers to a set of entities that are uniquely addressable and can communicate with each other. This can be a single container, a machine / server (physical or virtual), or some other network device (e.g., a router). A container can be conceptually added to one or more networks or conceptually removed from one or more networks. The CNI specification specifies many considerations for conformant plugins ("CNI plugins").

[0073] The container pool 22 includes one or more containers. In some examples, the container pool 22 includes a containerized DPDK workload that is designed to use DPDK to accelerate packet processing, e.g., by using the DPDK library to exchange data with other components to accelerate packet processing. In some examples, the virtual router 21 can run as a containerized DPDK workload.

[0074] The container pool 22 is configured with a virtual network interface 26 for sending and receiving packets via the virtual router 21. The virtual network interface 26 can be the default interface of the container pool 22. The container pool 22 can implement the virtual network interface 26 as an Ethernet interface (e.g., named "eth0"), while the virtual router 21 can implement the virtual network interface 26 as a tap interface, a virtio user interface, or some other type of interface.

[0075] The container pool 22 exchanges data packets with the virtual router 21 using the virtual network interface 26. The virtual network interface 26 can be a DPDK interface. The container pool 22 and the virtual router 21 can use vhost to set up the virtual network interface 26. The container pool 22 can operate according to an aggregation model. The container pool 22 can use virtual devices such as virtio devices with vhost-user adapters for inter-process communication of user space containers for the virtual network interface 26.

[0076] The CNI 17 can be combined with Figure 1 one or more other components shown in

[0077] The virtual network interface 26 can represent a virtual Ethernet (“veth”) pair, where each end of the pair is an independent device (e.g., a Linux / Unix device), and one end of the pair is assigned to the container pool 22 and one end of the pair is assigned to the virtual router 21. The veth pair or one end of the veth pair is sometimes referred to as a “port”. The virtual network interface can represent a macvlan network with media access control (MAC) addresses assigned to the container pool 22 and the virtual router 21 to communicate between the containers of the container pool 22 and the virtual router 21. Alternatively, for example, the virtual network interface can be referred to as a virtual machine interface (VMI), a container pool interface, a container network interface, a tap interface, a veth interface, or simply a network interface (in a specific context).

[0078] In Figure 1 the exemplary server 12A, the container pool 22 is a virtual network endpoint in one or more virtual networks. The orchestrator 23 can store or otherwise manage configuration data for application deployment that specifies the virtual network and specifies that the container pool 22 (or one or more containers therein) is a virtual network endpoint of the virtual network. For example, the orchestrator 23 can receive the configuration data from a user, an operator / administrator, or other machine systems.

[0079] As part of the process of creating the container pool 22, the orchestrator 23 requests the network controller 24 to create corresponding virtual network interfaces for one or more virtual networks (indicated in the configuration data). The container pool 22 can have different virtual network interfaces for each virtual network to which it belongs. For example, the virtual network interface 26 can be the virtual network interface for a specific virtual network. Additional virtual network interfaces (not shown) can be configured for other virtual networks. The network controller 24 processes the request to generate interface configuration data for the virtual network interfaces of the container pool 22. The interface configuration data can include a container or container pool unique identifier and a list or other data structure specifying the network configuration data for configuring the virtual network interfaces for each virtual network interface. The network configuration data for the virtual network interfaces can include a network name, an assigned virtual network address, a MAC address, and / or domain name server values. The following is an example of interface configuration data in JavaScript Object Notation (JSON) format.

[0080] The network controller 24 sends interface configuration data to the server 12A, and more specifically, in some cases, to the virtual router 21. To configure the virtual network interfaces for the container pool 22, the orchestrator 23 may invoke the CNI 17. The CNI 17 obtains the interface configuration data from the virtual router 21 and processes it. The CNI 17 creates the respective virtual network interfaces specified in the interface configuration data. For example, the CNI 17 may attach one end of a veth pair implementing the management interface 26 to the virtual router 21 and may attach the other end of the same veth pair using virtio user to implement the management interface 26 to the container pool 22.

[0081] The following is exemplary interface configuration data for the virtual network interface 26 of the container pool 22.

[0082]

[0083]

[0084] Conventional CNI plugins are invoked by the container platform / runtime, receive an Add command from the container platform to add a container to a single virtual network, and the plugin can then be invoked to receive a Del(ete) command from the container / runtime and remove the container from the virtual network. The term "invoke" may refer to the instantiation of a software component or module in memory as executable code to be executed by a processing circuit.

[0085] According to the techniques described in the present disclosure, the network controller 24 is a cloud-native distributed network controller for a software-defined network (SDN) implemented using one or more configuration nodes 30 and one or more control nodes 32. Each configuration node 30 itself may be implemented using one or more cloud-native component microservices. Each control node 32 itself may be implemented using one or more cloud-native component microservices.

[0086] In some examples, and as described in further detail below, the configuration node 30 may be implemented by extending the local orchestration platform to support custom resources for an orchestration platform for a software-defined network, and more specifically, for example, by configuring the virtual network interfaces of virtual execution elements, configuring the underlying network connection server 12, configuring an overlay routing function including an overlay tunnel for the virtual network and an overlay tree for multicast layer 2 and layer 3, to provide a northbound interface to the orchestration platform to support the purpose-driven / declarative creation and management of virtual networks.

[0087] As Figure 1As part of the SDN architecture shown, network controller 24 can be multi-tenant aware and support multi-tenancy for the orchestration platform. For example, network controller 24 can support Kubernetes role-based access control (RBAC) constructs, native identity access management (IAM), and external IAM integration. Network controller 24 can also support Kubernetes-defined network constructs and advanced network features such as virtual networks, BGPaaS, network policies, service chaining, and other telecommunications features. Network controller 24 can support network isolation using virtual network constructs and can support Layer 3 networks.

[0088] To interconnect multiple virtual networks, network controller 24 can use (and configure in the underlying and / or virtual router 21) a network policy referred to as a virtual network policy (VNP) and alternatively referred to herein as a virtual network router or virtual network topology. The VNP defines the connection policy between virtual networks. A single network controller 24 can support multiple Kubernetes clusters, and thus, the VNP allows connecting multiple virtual networks within namespaces, within Kubernetes clusters, and across Kubernetes clusters. The VNP can also be extended to support virtual network connections across multiple instances of network controller 24.

[0089] Network controller 24 is capable of supporting multi-layer security using network policies. The Kubernetes default behavior is for container pools to communicate with each other. To apply network security policies, the SDN architecture implemented by network controller 24 and virtual router 21 can operate as the CNI for Kubernetes via CNI 17. For Layer 3, network-level isolation occurs and virtual networks operate at L3. Virtual networks are connected via policies. Kubernetes native network policies provide Layer 4 security. The SDN architecture can support Kubernetes network policies. Kubernetes network policies operate at the boundaries of Kubernetes namespaces. The SDN architecture can add custom resources for enhanced network policies. The SDN architecture can support application-based security. (In some cases, these security policies can be based on meta-tags to apply granular security policies in a scalable manner). In some examples, for Layer 4+, the SDN architecture can support integration with containerized security devices and / or Istio and can provide encryption support.

[0090] As Figure 1As part of the SDN architecture shown, the network controller 24 can support multi-cluster deployments important for telecom cloud and high-end enterprise use cases. For example, the SDN architecture can support multiple Kubernetes clusters. The Cluster API can be used to support the lifecycle management of Kubernetes clusters. KubefedV2 can be used for cross-Kubernetes cluster configuration node 30 federation. The Cluster API and KubefedV2 are optional components for supporting a single instance of the network controller 24, which supports multiple Kubernetes clusters.

[0091] The SDN architecture can use network user interfaces and telemetry components to provide insights into the infrastructure, clusters, and applications. The telemetry nodes can be cloud-native and include microservices that support the insights.

[0092] Due to the above and other features described elsewhere herein, the computing infrastructure 8 implements a cloud-native SDN architecture and can provide one or more of the following technical advantages. For example, the network controller 24 is a cloud-native, lightweight distributed application with a simplified installation footprint. This also facilitates easier and modular upgrades of the various component microservices of the configuration node 30 and the control node 32 (and any other components of other examples of the network controller described in this disclosure). This technology can further enable optional cloud-native monitoring (telemetry) and user interfaces, a high-performance data plane of containers using DPDK-based virtual routers connected to a DPDK-supported container pool, and in some cases, cloud-native configuration management that takes advantage of the configuration frameworks of existing orchestration platforms such as Kubernetes or Openstack. As a cloud-native architecture, the network controller 24 is a scalable and resilient architecture that addresses and supports multiple clusters. In some cases, the network controller 24 can also support the scalability and performance requirements of key performance indicators (KPIs).

[0093] For example, an SDN architecture having the features and technical advantages described herein can be used to implement a cloud-native telecom cloud to support 5G mobile networks (and subsequent generations) and edge computing, as well as an enterprise Kubernetes platform including, for example, high-performance cloud-native application hosting. Telecom cloud applications are rapidly moving towards containerized cloud-native solutions. 5G fixed and mobile networks are driving the need to deploy workloads as microservices with significant disaggregation, specifically in 5G Next Generation RAN (5GNR). The 5G Next Generation Core Network (5GNC) may be deployed as a set of microservice-based applications corresponding to the various different components described by 3GPP. When considered as a group of microservices delivering applications, the 5GNC may be a highly complex combination of container pools with complex network, security, and policy requirements. For such use cases, a cloud-native SDN architecture having well-defined constructs for network, security, and policy described herein can be utilized. The network controller 24 can provide relevant APIs capable of creating these complex constructs.

[0094] Similarly, the user plane function (UPF) within the 5GNC is a super-high-performance application. It can be delivered as a set of highly distributed high-performance container pools. The SDN architecture described herein can provide a very high-throughput data plane (in terms of bits per section (bps) and packets per second (pps)). Integration with a DPDK virtual router having recent performance enhancements with eBPF and with a smart NIC will help achieve the required throughput. The DPDK-based virtual router is further described in U.S. Application No. 17 / 649,632, titled "CONTAINERIZED ROUTER WITH VIRTUAL NETWORKING", filed on February 1, 2022, the entire content of which is incorporated herein by reference.

[0095] High performance processing may also be relevant in GiLAN as workloads migrate from more traditional virtualized workloads to containerized microservices. In the data plane of the UPF and GiLAN services, such as GiLAN firewalls, intrusion detection and prevention, virtualized IP Multimedia Subsystem (vIMS) voice / video, etc., the throughput is high and remains constant in terms of bps and pps. For the control plane of 5G NC functions, such as Access and Mobility Management Function (AMF), Session Management Function (SMF), etc., and for some GiLAN services (e.g., IMS), when the absolute traffic can be moderate in terms of bps, the advantage of small packets is that the pps will remain high. In some examples, the SDN controller and the data plane provide millions of packets per second for each virtual router 21 implemented on server 12. In the 5G Radio Access Network (RAN), in order to move away from the proprietary vertically integrated RAN stack provided by traditional radio vendors, Open RAN decouples the RAN hardware and software in multiple components, including the non-real-time Radio Intelligent Controller (RIC), near real-time RIC, Central Unit (CU) control plane and user plane (CU-CP and CU-UP), Distributed Unit (DU), and Radio Unit (RU). If needed, the software components are deployed on a commercial server architecture implemented with programmable accelerators. The SDN architecture described herein can support the O-RAN specification.

[0096] Edge computing may be mainly targeted at two different use cases. The first scenario is as a support for containerized telecom infrastructure (e.g., 5G RAN, UPF, security functions), and the second scenario is for containerized service workloads for telecom and third parties such as providers or enterprise customers. In both cases, edge computing is actually a special case of GiLAN, where traffic is interrupted for special processing at highly distributed locations. In many cases, these locations will have limited resources (power, cooling, space). The SDN architecture described here may be very suitable for supporting the requirements of a very lightweight footprint, can support computing and storage resources in sites that are far from the associated control functions, and can sense location by means of deploying workloads and memories. Some sites may have as few as one or two computing nodes, which deliver a very specific set of services to highly localized users or other service sets. There may be a hierarchy of sites, where the central site is densely connected with multiple paths, the regional sites are multiply connected with two or four uplink paths, and the remote edge sites may have connections only to one or two upstream sites. This requires extreme flexibility in the way the SDN architecture can be deployed and the way (and location) in which the tunneling traffic within the overlay is terminated and bound to the core transport network (SRv6, MPLS, etc.). Similarly, in sites hosting telecom cloud infrastructure workloads, the SDN architecture described here can support the dedicated hardware (GPUs, smart NICs, etc.) required for high-performance workloads. There may also be workloads that require SR-IOV. Because, the SDN architecture can also support creating VTEPs at the ToR and making the VTEPs link back to the overlay as VXLANs.

[0097] A hybrid of fully distributed Kubernetes microclusters with each site running its own master cluster is expected, and the SDN architecture can support scenarios similar to remote computing.

[0098] For use cases involving enterprise Kubernetes platforms, high-performance cloud-native application power finance service platforms, online game services, and hosted application service providers, the cloud platforms delivering these applications must provide high performance, resilience to failures, high security, and visibility. Applications hosted on these platforms tend to be developed in-house. Application developers and platform owners work with infrastructure teams to deploy and operate instances of the organization's applications. These applications tend to require high throughput ( > 20 Gbps per server) and low latency. Some applications may also use multicast for signaling or payload traffic. Additional hardware and network infrastructure can be leveraged to ensure availability. Applications and microservices are partitioned using namespaces within the cluster. Isolation between namespaces is crucial in a highly secure environment. When the default deny policy is the standard posture in a zero-trust application deployment environment, additional network segmentation using virtual routing and forwarding instances (VRFs) adds an extra layer of security and allows the use of overlapping network ranges. Overlapping network ranges are a key requirement for managing the application hosting environment, which tends to standardize the set of reachable endpoints for all managed customers.

[0099] Complex microservice-based applications tend to utilize complex network filters. The SDN architecture described herein can deliver high-performance firewall filtering at scale. This filtering can exhibit consistent forwarding performance and lower latency degradation regardless of the rule set length or sequence. Some customers may also have some of the same regulatory pressures as the telecommunications industry for application partitioning not only at the network layer but also within the kernel. Specifically, when running on public clouds, finance and other services have a need for data plane encryption. In some examples, the SDN architecture described herein can include features to meet these needs.

[0100] In some examples, when the SDN architecture is automated through an application dev / test / stage / prod continuous integration / continuous deployment (CI / CD) pipeline, the SDN architecture can provide a GitOps-friendly UX for strict change management control, auditing, and reliability as the SDN architecture changes several times a day or even hundreds of times a day for the product.

[0101] Figure 2 is a block diagram showing an example of a cloud-native SDN architecture for a cloud-native network. The SDN architecture 200 is shown by abstracting the underlying connections between the various components. In this example, the network controller 24 of the SDN architecture 200 includes configuration nodes 230A - 230N (“configuration nodes” or “config nodes” and collectively “configuration nodes 230”) and control nodes 232A - 232K (collectively “control nodes 232”). The configuration nodes 230 and the control nodes 232 can each representFigure 1 Exemplary implementations of the configuration node 30 and the control node 32 in. Although shown as separate from the server 12, however, the configuration node 230 and the control node 232 may execute as one or more workloads on the server 12.

[0102] The configuration node 230 provides a northbound Representational State Transfer (REST) interface 248 to support the intent-driven configuration of the SDN architecture 200. Exemplary platforms and applications for pushing intents to the configuration node 230 include a virtual machine orchestrator 240 (e.g., Openstack), a container orchestrator 242 (e.g., Kubernetes), a user interface 244, or one or more other applications 246. In some examples, the SDN architecture 200 uses Kubernetes as its underlying platform.

[0103] The SDN architecture 200 is segmented into a configuration plane, a control plane, and a data plane, and an optional telemetry (or analytics) plane. The configuration plane is implemented using the horizontally scalable configuration node 230, the control plane is implemented using the horizontally scalable control node 232, and the data plane is implemented using compute nodes.

[0104] At a high level, the configuration node 230 uses the configuration store 224 to manage the configuration resource state of the SDN architecture 200. Generally, a configuration resource (or more simply "resource") is a named object schema that includes data and / or methods that describe a custom resource, and defines an application programming interface (API) for creating and manipulating data through an API server. A kind is the name of the object schema. Configuration resources can include Kubernetes native resources such as container pools, ingress, config maps, services, roles, namespaces, nodes, network policies, or load balancers. According to the techniques of the present disclosure, configuration resources also include custom resources that are used to extend the Kubernetes platform by defining application programming interfaces (APIs) that are not available in the default installation of the Kubernetes platform. In an example of the SDN architecture 200, the custom resources can describe the physical infrastructure, virtual infrastructure, configuration, and / or other resources of the SDN architecture 200. As part of configuring and operating the SDN architecture 200, various custom resources can be instantiated. The instantiated resources (whether native or custom) can be referred to as objects or instances of the resources, i.e., the persistent entities in the SDN architecture 200 that represent the intent (desired state) and condition (actual state) of the SDN architecture 200. The configuration node 230 provides an aggregated API for performing operations (i.e., create, read, update, and delete) on the configuration resources of the SDN architecture 200 in the configuration store 224. The load balancer 226 represents one or more load balancing objects that load balance configuration requests among the configuration nodes 230. The configuration store 224 can represent one or more etcd databases. The configuration node 230 can be implemented using Nginx.

[0105] The SDN architecture 200 can provide networking for Openstack and Kubernetes. Openstack uses a plugin architecture to support networking. With the virtual machine orchestrator 240, i.e., Openstack, the Openstack networking plugin driver converts Openstack configuration objects into configuration objects (resources) of the SDN architecture 200. The compute nodes run Openstack nova to provision virtual machines.

[0106] With the container orchestrator 242, i.e., Kubernetes, the SDN architecture 200 serves as the Kubernetes CNI. As described above, Kubernetes native resources (container pools, services, ingress, external load balancers, etc.) can be supported, and the SDN architecture 200 can support custom resources for Kubernetes for advanced networking and security of the SDN architecture 200.

[0107] The configuration node 230 provides REST monitoring to the control node 232 to monitor changes to configuration resources / objects, and the control node 232 within the computing infrastructure affects configuration resource changes. The control node 232 receives configuration resource data from the configuration node 230 by monitoring resources and constructs a complete configuration graph. A given control node in the control node 232 consumes the configuration resource data related to the control node and allocates the required configuration to the computing nodes (server 12) via the control interface 254 on the control plane side to the virtual router 21 (i.e., the virtual router agent - Figure 1 not shown in the figure) of the virtual router 21. As required by the processing, any computing node 232 can receive only a partial graph. The control interface 254 can be the Extensible Messaging and Presence Protocol (XMPP). The number of deployed configuration nodes 230 and control nodes 232 can be a function of the number of supported clusters. To support high availability, the configuration plane can include 2N + 1 configuration nodes 230 and 2N control nodes 232.

[0108] The control node 232 distributes routes among the computing nodes. The control node 232 uses the Interior Border Gateway Protocol (iBGP) to exchange routes among the control nodes 232, and the control node 232 can peer with any external BGP - supported gateway or other router. The control node 232 can use a route reflector.

[0109] The container pool 250 and the virtual machines 252 are examples of workloads that can be deployed to the computing nodes by the virtual machine orchestrator 240 or the container orchestrator 242 and interconnected by the SDN architecture 200 using one or more virtual networks.

[0110] Figure 3 is a block diagram showing another view of the components of the SDN architecture 200 according to the technology of the present disclosure and showing them in further detail. The configuration node 230, the control node 232, and the user interface 244 are shown as having their respective component microservices to implement the network controller 24 and the SDN architecture 200 as a cloud - native SDN architecture. Each component microservice can be deployed to the computing nodes.

[0111] Figure 3 A single cluster is shown divided into the network controller 24, the user interface 244, the computing nodes (server 12), and the telemetry 260 feature. The configuration node 230 and the control node 232 together constitute the network controller 24.

[0112] The configuration node 230 can include a component microservice API server 300 (or "Kubernetes API server 300" - the corresponding controller 406, Figure 3Not shown), a custom API server 301, a custom resource controller 302, and an SDN controller manager 303 (sometimes referred to as "kubemanager" or "SDN kubemanager", where the orchestration platform for the network controller 24 is Kubernetes). "contrail-k8s-kubemanager" is an example of the SDN controller manager 303. The SDN controller manager 303 is different from and has different responsibilities than the kube-controller-manager, which is a daemon that embeds core control loops into the controllers of Kubernetes, such as controllers including replication controllers, endpoint controllers, namespace controllers, and service account controllers. The configuration node 230 extends the interface of the API server 300 using the custom API server 301 to form an aggregation layer to support the data model of the SDN architecture 200. As described above, the configuration intent of the SDN architecture 200 can be a custom resource.

[0113] The control node 232 may include a component microservice control 320. As referred to above Figure 2 described, the control 320 performs configuration assignment and routing learning and assignment.

[0114] The server 12 represents a compute node. Each compute node includes a virtual router agent 316 and a virtual router forwarding component (vRouter) 318. Either the virtual router agent 316 or the vRouter 318, or both, can be component microservices. Generally, the virtual router agent 316 performs control-related functions. The virtual router agent 316 receives configuration data from the control node 232 and converts the configuration data into forwarding information for the vRouter 318. The virtual router agent 316 can also perform firewall rule processing, set flows for the vRouter 318, and interface with orchestration plugins (CNI for Kubernetes and Nova plugin for Openstack). When a workload (container pool or VM) is launched on the compute node, the virtual router agent 316 generates a route, and the virtual router 316 exchanges the route with the control node 232 for distribution to other compute nodes (the control node 232 uses BGP to distribute routes among the control nodes 232). When the workload is terminated, the virtual router agent 316 also withdraws the route. The vRouter 318 can support one or more forwarding modes, such as kernel mode, DPDK, Smart NIC offloading, etc. In some examples of container architectures or virtual machine workloads, the compute node can be a Kubernetes worker / minion node or an Openstack nova-compute node, depending on the specific orchestrator used.

[0115] One or more optional telemetry nodes 260 provide metrics, alerts, logging, and flow analysis. Telemetry for the SDN architecture 200 leverages cloud-native monitoring services such as Prometheus, Elastic, Fluentd, Kinaba stack (EFK), and Influx TSDB. SDN architecture component microservices of the configuration node 230, control node 232, compute nodes, user interface 244, and analytics nodes (not shown) can generate telemetry data. Services of the telemetry node 260 can consume this telemetry data. The telemetry node 260 can expose REST endpoints for users and can support insights and event correlation.

[0116] Optionally, the user interface 244 includes a network user interface (UI) 306 and a UI backend 308 service. Generally, the user interface 244 provides configuration, monitoring, visualization, security, and troubleshooting for SDN architecture components.

[0117] Each of the telemetry 260, user interface 244, configuration node 230, control node 232, and server 12 / compute node can be considered a node of the SDN architecture 200 in that each of these nodes is an entity that implements the functions of the configuration plane, control plane, or data plane, or the UI and telemetry nodes. The configuration node scale is configured during the "setup" process, and the SDN architecture 200 uses an orchestration system operator such as a Kubernetes operator to support the automatic scaling of the nodes of the SDN architecture 200.

[0118] Figure 4 is a block diagram showing exemplary components of an SDN architecture according to the techniques of the present disclosure. In this example, the SDN architecture 400 extends and uses the Kubernetes API server for network configuration objects to implement the user intent of the network configuration. In Kubernetes terminology, this network configuration object is called a custom resource and is abbreviated as an object when persisted in the SDN architecture. The configuration object is mainly the user intent (e.g., virtual network, BGPaaS, network policy, service chain, etc.). There may be similar objects for Kubernetes native resources.

[0119] The configuration node 230 of the SDN architecture 400 can use the Kubernetes API server for configuration objects. In kubernetes terminology, it is called a custom resource.

[0120] Kubernetes provides two ways to add custom resources to a cluster:

[0121] · Custom Resource Definition (CRD) is simple and can be created without any programming.

[0122] · API aggregation requires programming, but allows more control over the API behavior, such as controlling how data is stored and controlling the conversion between API versions.

[0123] The Aggregation API is a secondary API server that acts as a proxy behind the main API server. This setup is called API Aggregation (AA). To the user, it appears as simply an extension of the Kubernetes API. CRDs allow users to create new types of resources without adding another API server. Regardless of how they are installed, the new resources are called Custom Resources (CRs) to distinguish them from native Kubernetes resources (e.g., pod). CRDs are used in the initial Config prototype. This architecture can implement the Aggregation API using the API Server Builder Alpha library. The API Server Builder is a collective term for libraries and tools for building native Kubernetes aggregation extensions.

[0124] Typically, each resource in the Kubernetes API requires code to handle REST requests and manage persistent storage of the object. The main Kubernetes API server 300 (implemented via API server microservices 300A–300J) handles native resources and is also able to generally process custom resources via CRDs. The Aggregation API 402 represents an aggregation layer that extends the Kubernetes API server 300 to allow a dedicated implementation of custom resources by writing and deploying a custom API server 301 (using custom API server microservices 301A–301M). The main API server 300 delegates requests for custom resources to the custom API server 301, thereby making the resources available to all its clients.

[0125] Thus, the API server 300 (e.g., kube - api server) receives Kubernetes configuration objects, local objects (pods, services), and custom resources defined according to the techniques of the present disclosure. Custom resources for the SDN architecture 400 may include configuration objects that, when the desired state of the configuration object is implemented in the SDN architecture 400, implement the desired network configuration of the SDN architecture 400. The custom resources may correspond to configuration schemes that are conventionally defined for network configuration but conform to the techniques of the present disclosure, and the custom resources are extended to be manipulated via the aggregated API 402. Alternatively, the custom resources may be named and are referred to herein as "custom resources for SDN architecture configuration". Each custom resource for SDN architecture configuration may correspond to a type of configuration object that is typically exposed by an SDN controller but conforms to the techniques described herein, using the custom resources to expose the configuration objects and merge the configuration objects with Kubernetes local / built - in resources. These configuration objects may include virtual networks, bgp as a service (BGPaaS), subnets, virtual routers, service instances, projects, physical interfaces, logical interfaces, nodes, network ipam, floating ips, alarms, alias ips, access control lists, firewall policies, firewall rules, network policies, route targets, route instances, etc. Thus, the SDN architecture system 400 supports a unified intent model exposed by the aggregated API 402, i.e., the unified intent model is implemented via the Kubernetes controllers 406A - 406N and via the custom resource controller 302 (shown as component microservices 302A - 302L in Figure 4 ), and the Kubernetes controllers 406A - 406N and the custom resource controller 302 work to make the actual state of the computing infrastructure including network elements consistent with the desired state. The controllers 406 may represent kube - controller - managers.

[0126] The aggregation layer of the API server 300 sends the API custom resources to their corresponding registered custom API servers 301. There may be multiple custom API servers / custom resource controllers that support different kinds of custom resources. The custom API server 301 processes the custom resources for SDN architecture configuration and writes the custom resources to the configuration store 304 (which may be etcd). The custom API server 301 may be a host and expose the SDN controller identifier allocation service required by the custom resource controller 302.

[0127] The custom resource controller 302 begins to apply business logic to meet the user intent set with the user intent configuration. The business logic is implemented as a reconciliation loop. Figure 8FIG. is a block diagram illustrating an example of a custom controller for custom resources for SDN architecture configuration according to the techniques of the present disclosure. The custom controller 814 may represent an exemplary instance of the custom resource controller 301. In Figure 8 the example shown, the custom controller 814 may be associated with the custom resource 818. The custom resource 818 may be any custom resource for SDN architecture configuration. The custom controller 814 can include a coordinator 816 that includes logic to execute a coordination loop, where the custom controller 814 observes 834 (e.g., monitors) the current state 832 of the custom resource 818. In response to determining that the desired state 836 does not match the current state 832, the coordinator 816 can perform an action to adjust 838 the state of the custom resource so that the current state 832 matches the desired state 836. A request can be received by the API server 300 and relayed to the custom API server 301 to change the current state 832 of the custom resource 818 to the desired state 836.

[0128] In the case where the API request 301 is a create request for a custom resource, the coordinator 816 can act on a create event for instance data of the custom resource. The coordinator 816 can create instance data of custom resources that the requested custom resource depends on. As an example, an edge node custom resource can depend on a virtual network custom resource, a virtual interface custom resource, and an IP address custom resource. In this example, when the coordinator 816 receives a create event for an edge node custom resource, the coordinator 816 can also create the custom resources that the edge node custom resource depends on, such as a virtual network custom resource, a virtual interface custom resource, and an IP address custom resource.

[0129] By default, the custom resource controller 302 runs in an active - passive mode and uses leader election to achieve consistency. When the controller container pool starts, it attempts to create a ConfigMap resource in Kubernetes using a specified secret. If the creation is successful, the container pool becomes the leader and starts processing coordination requests; otherwise, the container pool blocks the attempt to create a ConfigMap in an infinite loop.

[0130] The custom resource controller 300 can track the status of the custom resources it creates. For example, a virtual network (VN) creates a routing instance (RI), and the routing instance (RI) creates a routing target (RT). If the creation of the routing target fails, the status of the routing instance degrades, and thus, the status of the virtual network also degrades. Therefore, the custom resource controller 300 can output custom messages indicating the status of these custom resources to troubleshoot. Figure 9Exemplary flows for creating, observing, and coordinating between custom resource types that depend on different custom resource types are shown.

[0131] The configuration plane implemented by the configuration node 230 has high availability. The configuration node 230 can be based on Kubernetes and includes a kube - api server service (e.g., API server 300) and a storage backend etcd (e.g., configuration store 304). In fact, the aggregated API 402 implemented by the configuration node 230 operates as the front - end of the control plane implemented by the control node 232. The main implementation of the API server 300 is a kube - api server designed to scale horizontally by deploying more instances. As shown, several instances of the API server 300 can be run to load - balance API requests and processing.

[0132] The configuration store 304 can be implemented as etcd. Etcd is a consistent and highly available key - value store used as the Kubernetes backup storage for cluster data.

[0133] In Figure 4 the example of, each of the servers 12 in the SDN architecture 400 includes an orchestration agent 420 and a containerized (or "cloud - native") routing protocol daemon 324. These components of the SDN architecture 400 are described in further detail below.

[0134] The SDN controller manager 303 can operate as an interface between Kubernetes core resources (services, namespaces, container pools, network policies, network attachment definitions) and the extended SDN architecture resources (virtual networks, routing instances, etc.). The SDN controller manager 303 monitors changes to the Kubernetes API on the custom resources configured for the Kubernetes core network and the SDN architecture and is thus able to perform CRUD operations on the relevant resources. As used herein, monitoring a resource can refer to monitoring an object or object instance of the resource type of the resource.

[0135] In some examples, the SDN controller manager 303 is a collective term for one or more Kubernetes custom controllers. In some examples, in a single or multiple cluster deployments, the SDN controller manager 303 can run on the Kubernetes cluster it manages.

[0136] The SDN controller manager 303 listens for the following Kubernetes objects for create, delete, and update events:

[0137] · Container pool

[0138] · Service

[0139] · Node port

[0140] · Inlet

[0141] · Endpoint

[0142] · Namespace

[0143] · Deployment

[0144] · Network policy.

[0145] When generating these events, the SDN controller manager 303 creates appropriate SDN architecture objects, which are then defined as custom resources of the SDN architecture configuration. In response to detecting an event on an instance of the custom resource, whether instantiated by the SDN controller manager 303 and / or through the custom API server 301, the control node 232 obtains configuration data about the instance of the custom resource and configures the corresponding instance of the configuration object in the SDN architecture 400.

[0146] For example, the SDN controller manager 303 monitors container pool creation events, and accordingly, can create the following SDN architecture objects: virtual machines (workloads / container pools), virtual machine interfaces (virtual network interfaces), and instance IPs (IP addresses). Then, in this case, the control node 232 can instantiate the SDN architecture objects in the selected compute nodes.

[0147] As an example, based on watch, the control node 232A can detect an event on an instance of a first custom resource exposed by the custom API server 301A, where the first custom resource is used to configure an aspect of the SDN architecture system 400 and corresponds to a configuration object type of the SDN architecture system 400. For example, the type of the configuration object can be a firewall rule corresponding to the first custom resource. In response to the event, the control node 232A can obtain configuration data about the firewall rule instance (e.g., firewall rule specification) and provide the firewall rule in the virtual router for server 12A. The configuration node 230 and the control node 232 can perform similar operations on other custom resources using the corresponding type of configuration objects of the SDN architecture, such as virtual networks, bgp as a service (BGPaaS), subnets, virtual routers, service instances, projects, physical interfaces, logical interfaces, nodes, network ipam, floating ips, alerts, alias ips, access control lists, firewall policies, firewall rules, network policies, route targets, route instances, etc.

[0148] Figure 5 is a block diagram of an exemplary computing device according to the techniques described in the present disclosure. Figure 5The computing device 500 therein may represent an actual server or a virtual server and may represent an exemplary example of any server 12 and may be referred to as a computing node, a master / slave node, or a host. In this example, the computing device 500 includes a bus 542 coupled to hardware components in the hardware environment of the computing device 500. The bus 542 is coupled to a network interface card (NIC) 530, a storage disk 546, and one or more microprocessors 510 (hereinafter referred to as "microprocessors 510"). The NIC 530 may support SR-IOV. In some cases, the front-side bus may be coupled to the microprocessors 510 and the memory device 544. In some examples, the bus 542 may be coupled to the memory device 544, the microprocessors 510, and the NIC 530. The bus 542 may represent a Peripheral Component Interconnect (PCI) Express (PCIe) bus. In some examples, a Direct Memory Access (DMA) controller may control DMA transfers between components coupled to the bus 542. In some examples, the components coupled to the bus 542 control DMA transfers between components coupled to the bus 542.

[0149] The microprocessors 510 may include one or more processors, each of which includes an independent execution unit that executes instructions conforming to an instruction set architecture and instructions stored to a storage medium. The execution units may be implemented as separate integrated circuits (ICs) or may be combined within one or more multi-core processors (or "multi-core" processors) each implemented using a single IC (i.e., a chip microprocessor).

[0150] The disk 546 represents a computer-readable storage medium, including volatile and / or non-volatile, removable and / or non-removable media implemented in any method or technology for storing information such as processor-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), EEPROM, flash memory, CD-ROM, digital versatile disk (DVD), or other optical memory, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the required information and that can be accessed by the microprocessors 510.

[0151] Main memory 544 includes one or more computer-readable storage media, which may include random access memory (RAM) such as various forms of dynamic RAM (DRAM) (e.g., DDR2 / DDR3 SDRAM), or static RAM (SRAM), flash memory, or any other form of fixed or removable storage media that can be used to carry or store instructions or data structures in the form of required program code and program data and that is accessible to a computer. Main memory 544 provides a physical address space consisting of addressable memory locations.

[0152] Network interface card (NIC) 530 includes one or more interfaces 532 configured to exchange packets using the links of the underlying physical network. Interface 532 may include a port interface card having one or more network ports. NIC 530 may also include, for example, on-card memory for storing packet data. Direct memory access transfers between NIC 530 and other devices coupled to bus 542 may be read from the NIC memory, or direct memory access transfers between NIC 530 and other devices coupled to bus 542 may be written to the NIC memory.

[0153] Memory 544, NIC 530, storage disk 546, and microprocessor 510 may provide an operating environment for a software stack that includes an operating system kernel 580 executing in kernel space. For example, kernel 580 may represent the kernel of Linux, Berkeley Software Distribution (BSD), another Unix variant, or a Windows Server operating system kernel commercially available from Microsoft Corporation. In some examples, the operating system may run a hypervisor and one or more virtual machines managed by the hypervisor. Exemplary hypervisors include the Kernel-based Virtual Machine (KVM) for the Linux kernel, Xen, ESXi commercially available from VMware, Windows Hyper-V commercially available from Microsoft, and other open source and proprietary hypervisors. The term hypervisor can encompass a virtual machine monitor (VMM). The operating system including kernel 580 provides an execution environment for one or more processes in user space 545.

[0154] Kernel 580 includes a physical driver 525 that uses network interface card 530. Network interface card 530 may also implement SR-IOV to enable it to operate in, for example, containers 529A or one or more virtual machines ( Figure 5Share physical network functions (I / O) among one or more virtual execution elements (not shown in the figure). Shared virtual devices such as virtual functions can provide dedicated resources so that each virtual execution element can access the dedicated resources of the NIC 530. Therefore, the NIC 530 appears to be a dedicated NIC to each virtual execution element. A virtual function can represent a lightweight PCIe function that uses the physical functions used by the physical drive 525 and shares physical resources with other virtual functions. For an NIC 530 that supports SR-IOV, the NIC 530 can have thousands of available virtual functions that comply with the SR-IOV standard. However, for I / O-intensive applications, the number of configured virtual functions is usually much less.

[0155] The computing device 500 can be coupled to a physical network fabric that includes an overlay network that extends the fabric from a physical switch to a software or "virtual" router (including the virtual router 506) of a physical server coupled to the fabric. The virtual router can be a process or thread executed by a physical server (e.g., Figure 1 Server 12 in the figure), or a combination thereof, and the virtual router dynamically creates and manages one or more virtual networks that can be used for communication between virtual network endpoints. In one example, each virtual router uses an overlay network to implement a virtual network, and the overlay network provides the ability to decouple the virtual addresses of the endpoints from the physical addresses (e.g., IP addresses) on the server where the endpoints are executed. Each virtual network can use its own addressing and security scheme and can be considered orthogonal to the physical network and its addressing scheme. Various techniques can be used to transmit packets within and across virtual networks over a physical network. As used herein, the term "virtual router" can cover Open vSwitch (OVS), OVS bridges, Linux bridges, Docker bridges, or other devices and / or software located on a host device and performing switching, bridging, or routing packets between virtual network endpoints of one or more virtual networks, where the virtual network endpoints are hosted by one or more servers 12. In Figure 5 the exemplary computing device 500, the virtual machine router 506 runs as a DPDK-based virtual router in user space. However, the virtual router 506 can run in a hypervisor, a host operating system, a host application, or a virtual machine through various implementations.

[0156] The virtual router 506 can replace and incorporate the virtual routing / bridging functions of the Linux bridge / OVS module that is typically used in the Kubernetes deployment of the container pool 502. The virtual router 506 can perform bridging (e.g., E-VPN) and routing (e.g., L3VPN, IP-VPN) for the virtual network. The virtual router 506 can perform network services such as applying security policies, NAT, multicast, mirroring, and load balancing.

[0157] The virtual router 506 can run as a kernel module or as a user-space DPDK process (here, the virtual router 506 located in the user space 545 is shown). The virtual router agent 514 can also run in the user space. In the exemplary computing device 500, the virtual router 506 runs as a DPDK-based virtual router within the user space. However, the virtual router 506 can run within the hypervisor, the host operating system, the host application, or the virtual machine through various implementations. The virtual router agent 514 uses a channel to connect to the network controller 24, and the channel is used to download configuration and forwarding information. The virtual router agent 514 programs the forwarding state into the virtual router data (or "forwarding") plane represented by the virtual router 506. The virtual router 506 and the virtual router agent 514 can be processes. The virtual router 506 and the virtual router agent 514 are containerized / cloud-native.

[0158] The virtual router 506 can be multi-threaded and run on one or more processor cores. The virtual router 506 can include multiple queues. The virtual router 506 can implement a packet processing pipeline. Depending on the operations applied to the packets, the virtual router agent 514 stitches the pipeline in the simplest way to the most complex way. The virtual router 506 can maintain multiple instances of the forwarding infrastructure. The virtual router 506 can use RCU (Read-Copy Update) locks to access and update the tables.

[0159] To send packets to other computing nodes or switches, the virtual router 506 uses one or more physical interfaces 532. Generally, the virtual router 506 exchanges overlay packets with workloads such as VMs or the container pool 502. The virtual router 506 has multiple virtual network interfaces (e.g., vifs). These interfaces can include the kernel interface vhost0 that exchanges packets with the host operating system, and the interface pkt0 that exchanges packets with the virtual router agent 514 to obtain the forwarding state from the network controller and send exception packets. There may be one or more virtual network interfaces corresponding to one or more physical network interfaces 532. Other virtual network interfaces of the virtual router 506 are used to exchange packets with the workloads.

[0160] In a kernel-based deployment of virtual router 506 (not shown), the virtual router 506 is installed as a kernel module within the operating system. The virtual router 506 registers itself through the TCP / IP stack to receive packets from any desired operating system interface. The interface can be bonded, physicalized, tapped (for VMs), veth (for containers), etc. In this mode, the virtual router 506 relies on the operating system to send packets to different interfaces and receive packets from different interfaces. For example, the operating system can expose a tapped interface supported by a vhost network driver to communicate with the VM. Once the virtual router 506 registers packets from this tapped interface, the TCP / IP stack sends all packets to the virtual router. The virtual router 506 sends packets via the operating system. In addition, the NIC queues (physical or virtual) are handled by the operating system. Packet processing can operate in interrupt mode, which generates interrupts and may cause frequent context switches. When there is a high packet rate, the total overhead associated with frequent interrupts and context switches may overwhelm the operating system and result in poor performance.

[0161] In a DPDK-based deployment of virtual router 506 ( Figure 5 not shown), the virtual router 506 is installed as a user space 545 application linked to the DPDK library. Specifically, this may result in faster performance than a kernel-based deployment when there is a high packet rate. The physical interface 532 is used by the poll mode driver (PMD) of DPDK, rather than the kernel's interrupt-based driver. The registers of the physical interface 532 can be exposed to the user space 545 to make them accessible to the PMD; the physical interface 532 bonded in this way is no longer managed by or visible to the host operating system, and the DPDK-based virtual router 506 manages the physical interface 532. This includes packet polling, packet processing, and packet forwarding. In other words, the user packet processing steps are performed by the DPDK data plane of the virtual router 506. Compared with the interrupt mode, the "poll mode" nature makes the packet processing / forwarding of the DPDK data plane of the virtual router 506 more efficient when the packet rate is high. Compared with the kernel-mode virtual router 506, there are relatively fewer interrupts and context switches during packet I / O, and in some cases, interrupts and context switches during packet I / O can be avoided altogether.

[0162] Generally, one or more virtual network addresses for use within a corresponding virtual network can be allocated to each of the container pools 502A–502B, where each virtual network can be associated with a different virtual subnet provided by the virtual router 506. For example, a third virtual layer (L3) IP address of its own can be allocated to the container pool 502B to send and receive communications, however, the container pool 502B of the computing device 500 may not know the IP address on which the container pool 502B of the computing device 500 runs. Thus, the virtual network address may be different from the logical address of the underlying physical computer system (e.g., the computing device 500).

[0163] The computing device 500 includes a virtual router agent 514 that controls the overlay of the virtual network of the computing device 500 and coordinates the routing of data packets within the computing device 500. Generally, the virtual router agent 514 communicates with a network controller 24 for the virtualization infrastructure, and the virtualization infrastructure generates commands to create a virtual network and configure network virtualization endpoints, such as the computing device 500, and more specifically, the virtual router 506, and virtual network interfaces. By configuring the virtual router 506 based on the information received from the network controller 24, the virtual router agent 514 can support the configuration of network isolation, policy-based security, gateways, source network address translation (SNAT), load balancers, and service chaining capabilities for orchestration.

[0164] In one example, a network packet, e.g., a Layer 3 (L3) IP packet or a Layer 2 (L2) Ethernet packet generated and used by containers 529A–529B located within a virtual network domain, can be encapsulated within another packet (e.g., another IP or Ethernet packet) transmitted by the physical network. The packet transmitted within the virtual network can be referred to herein as an “inner packet”, while the physical network packet can be referred to herein as an “outer packet” or a “tunnel packet”. Encapsulation and / or decapsulation of the virtual network packet within the physical network packet can be performed by the virtual router 506. Herein, this function is referred to as tunneling and can be used to create one or more overlay networks. Other exemplary tunneling protocols that can be used, in addition to IPinIP, include Generic Routing Encapsulation (GRE)-based Multiprotocol Label Switching (MPLS), UDP-based MPLS, etc. The virtual router 506 performs tunnel encapsulation / decapsulation on packets originating from / being destined for any container of the container pool 502, and the virtual router 506 exchanges packets with the container pool 502 via the bus 542 and / or the bridge of the NIC 530.

[0165] As described above, the network controller 24 can provide a logically centralized controller that facilitates the operation of one or more virtual networks. For example, the network controller 24 can maintain a routing information base, e.g., one or more routing tables that store routing information about the physical network and one or more overlay networks. The virtual router 506 enforces one or more virtual routing and forwarding instances (VRFs), such as VRF 222A, for the corresponding virtual networks for which the virtual router 506 operates as the corresponding tunnel endpoint. Generally, each VRF stores forwarding information about the corresponding virtual network and identifies where data packets are to be forwarded and whether the packets are encapsulated in a tunnel protocol, such as encapsulated with a tunnel header that can include one or more headers of different layers of the virtual network protocol stack. Each VRF can include a network forwarding table that stores routing and forwarding information about the virtual network.

[0166] The NIC 530 can receive tunnel packets. The virtual router 506 processes the tunnel packets to determine the virtual networks of the source and destination endpoints of the inner packet from the tunnel encapsulation header. The virtual router 506 can strip the Layer 2 header and the tunnel encapsulation header to forward only the inner packet internally. The tunnel encapsulation header can include a virtual network identifier that indicates the virtual network (e.g., the virtual network corresponding to VRF 222A), such as a VxLAN tag or an MPLS label. VRF 222A can include forwarding information about the inner packet. For example, VRF 222A can map the destination Layer 3 address of the inner packet to a virtual network interface. Accordingly, VRF 222A forwards the inner packet via the virtual network interface of the container pool 502A.

[0167] The container 529A can also initiate an inner packet as a source virtual network endpoint. For example, the container 529A can generate a Layer 3 inner packet that is destined for a destination virtual network endpoint executed by another computing device (i.e., non-computing device 500) or destined for another container. The container 529A can send the Layer 3 inner packet to the virtual router 506 via the virtual network interface attached to VRF 222A.

[0168] The virtual router 506 receives internal packets and Layer 2 headers and determines the virtual network of the internal packets. The virtual router 506 can determine the virtual network using any of the virtual network interface implementation techniques described above (e.g., macvlan, veth, etc.). The virtual router 506 uses the VRF 222A corresponding to the virtual network of the internal packet to generate an external header for the internal packet, an external header including an external IP header for overlay tunneling, and a tunnel encapsulation header identifying the virtual network. The virtual router 506 encapsulates the internal packet with the external header. The virtual router 506 can encapsulate the tunnel packet with a new Layer 2 header that has a destination Layer 2 address associated with a device external to the computing device 500 (e.g., the TOR switch 16 or a server 12). If external to the computing device 500, the virtual router 506 uses the physical function to output the tunnel packet with the new Layer 2 header to the NIC 530. The NIC 530 outputs the packet on the outbound interface. If the destination is another virtual network endpoint executing on the computing device 500, the virtual router 506 routes the packet to an appropriate virtual network interface among the virtual network interfaces.

[0169] In some examples, a controller for the computing device 500 (e.g., Figure 1 the network controller 24 therein) configures the default route in each container pool 502 to cause the virtual machines 224 to use the virtual router 506 as the initial next hop for outbound packets. In some examples, the NIC 530 is configured with one or more forwarding rules that cause all packets received from the virtual machines 224 to be switched to the virtual router 506.

[0170] The container pool 502A includes one or more application containers 529A. The container pool 502B includes an instance of the containerized routing protocol daemon (cRPD) 560. The container platform 588 includes a container engine 590, an orchestration agent 592, a service agent 593, and a CNI 570.

[0171] The container engine 590 includes code executed by the microprocessor 510. The container engine 590 can be one or more computer processes. The container engine 590 runs containerized applications in the form of containers 529A - 529B. The container engine 590 can represent Dockert, rkt, or other container engines for managing containers. Generally, the container engine 590 receives requests and manages objects such as images, containers, networks, and capacity. An image is a template with instructions for creating a container. A container is an executable instance of an image. Based on instructions from the controller agent 592, the container engine 590 can obtain an image and instantiate it as an executable container in the container pools 502A - 502B.

[0172] The service agent 593 includes code executable by the microprocessor 510. The service agent 593 can be one or more computer processes. The service agent 593 monitors the addition and removal of services and endpoint objects, and for example, it uses services to maintain the network configuration of the computing device 500 to ensure communication between the container pool and the containers. The service agent 593 can also manage the management ip table to capture the traffic of the virtual IP address and port of the service and redirect the traffic to the proxy port of the proxy backup container pool. The service agent 593 can represent the kube proxy of the slave node of the Kubernetes cluster. In some examples, the container platform 588 does not include the service agent 593, or disables the service agent 593 by the CNI 570 to facilitate the configuration of the virtual router 506 and the container pool 502.

[0173] The orchestration agent 592 includes code executable by the microprocessor 510. The orchestration agent 592 can be one or more computer processes. The orchestration agent 592 can represent the kubelet of the slave node of the Kubernetes cluster. The orchestration agent 592 is an agent that receives the container-specific data of the container and ensures the execution of the container by the computing device 500 by the orchestrator, for example, Figure 1 the orchestrator 23 in. The container-specific data can be in the form of a manifest file sent from the orchestrator 23 to the orchestration agent 592 or received indirectly via a command-line interface, an HTTP endpoint, or an HTTP server. The container-specific data can be the container pool specification of a container pool 502 of the container (for example, PodSpec - YAML (Yet Another Markup Language) or a JSON object describing the container pool). Based on the container-specific data, the orchestration agent 592 guides the container engine 590 to obtain the container image of the container 529 and instantiate the container image for the computing device 500 to execute the container 529.

[0174] The orchestration agent 592 instantiates the CNI 570 and calls the CNI 570 in other ways to configure one or more virtual network interfaces of each container pool 502. For example, the orchestration agent 592 receives the container-specific data of the container pool 502A and guides the container engine 590 to create the container pool 502A with the container 529A based on the container-specific data of the container pool 502A. The orchestration agent 592 also calls the CNI 570 to configure the virtual network interface of the virtual network corresponding to the VRF 222A for the container pool 502A. In this example, the container pool 502A is the virtual network endpoint of the virtual network corresponding to the VRF 222A.

[0175] The CNI 570 can obtain interface configuration data for configuring virtual network interfaces of the container pool 502. The virtual router agent 514 operates as a virtual network control plane module that enables the network controller 24 to configure the virtual router 506. Different from the orchestration control plane that manages the provisioning, scheduling, and management of virtual execution elements (including the container platform 588 for worker nodes and master nodes, e.g., the orchestrator 23), the virtual network control plane (including the network controller 24 and the virtual router agent 514 for worker nodes) manages the configuration of the virtual network partially implemented by the virtual router 506 of the worker node in the data plane. The virtual router agent 514 communicates the interface configuration data regarding the virtual network interface to the CNI 570 to enable the orchestration control plane element (i.e., the CNI 570) to configure the virtual network interface according to the configuration status determined by the network controller 24, thereby bridging the gap between the orchestration control plane and the virtual network control plane. Additionally, this enables the CNI 570 to obtain the interface configuration data for multiple virtual network interfaces of the container pool and configure the multiple virtual network interfaces, which can reduce the communication and resource overhead inherent in invoking separate CNI 570s for configuring each virtual network interface.

[0176] The containerized routing protocol daemon is described in U.S. Application No. 17 / 649,632, filed on February 1, 2022, the entire content of which is incorporated herein by reference.

[0177] Figure 6 It is a block diagram of an exemplary computing device of a computing node of one or more clusters operating as an SDN architecture system according to the technology of the present disclosure. The computing device 1300 may represent one or more physical or virtual servers. In some instances, the computing device 1300 may implement one or more master nodes of the corresponding cluster or multiple clusters.

[0178] Although shown and described as being run by a single computing device 1300, however, the scheduler 1322, the API server 300A, the controller 406A, the custom API server 301A, the custom resource controller 302A, the controller manager 1326, the SDN controller manager 1325, the control node 232A, and the configuration store 1328 may be distributed among multiple computing devices constituting a computing system or a hardware / server cluster. In other words, each of the multiple computing devices may provide a hardware operating environment for one or more instances of any one or more of the scheduler 1322, the API server 300A, the controller 406A, the custom API server 301A, the custom resource controller 302A, the network controller manager 1326, the network controller, the SDN controller manager 1325, the control node 232A, or the configuration store 1328.

[0179] In this example, computing device 1300 includes a bus 1342 coupled to hardware components in the hardware environment of computing device 1300. Bus 1342 is coupled to a network interface card (NIC) 1330, a storage disk 1346, and one or more microprocessors 1310 (hereinafter referred to as "microprocessors 1310"). In some cases, a front-side bus may couple microprocessors 1310 to a memory device 1344. In some examples, bus 1342 may couple memory device 1344, microprocessors 1310, and NIC 1330. Bus 1342 may represent a Peripheral Component Interconnect (PCI) Express (PCIe) bus. In some examples, a direct memory access (DMA) controller may control DMA transfers between components coupled to bus 1342. In some examples, components coupled to bus 1342 control DMA transfers between components coupled to bus 1342.

[0180] Microprocessors 1310 may include one or more processors, each of which includes an independent execution unit that executes instructions that conform to an instruction set architecture and instructions stored to a storage medium. The execution units may be implemented as separate integrated circuits (ICs) or may be combined within one or more multi-core processors (or "multi-core" processors) each implemented using a single IC (i.e., a chip multi-processor).

[0181] Disk 1346 represents a computer-readable storage medium, including volatile and / or non-volatile, removable and / or non-removable media implemented in any method or technology for storing information such as processor-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), EEPROM, flash memory, CD-ROM, digital versatile disk (DVD), or other optical memory, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by microprocessors 1310.

[0182] Main memory 1344 includes one or more computer-readable storage media, which may include random access memory (RAM) such as various forms of dynamic RAM (DRAM) (e.g., DDR2 / DDR3 SDRAM), or static RAM (SRAM), flash memory, or any other form of fixed or removable storage medium that can be used to carry or store instructions or data structures in the form of program code and program data required by a computer and that can be accessed by the computer. Main memory 1344 provides a physical address space composed of addressable memory locations.

[0183] Network interface card (NIC) 1330 includes one or more interfaces 1332 configured to exchange packets using a link of an underlying physical network. Interface 1332 may include a port interface card having one or more network ports. NIC 1330 may also include, for example, on-card memory for storing packet data. A direct memory access transfer between NIC 1330 and other devices coupled to bus 1342 may be read from or written to the NIC memory.

[0184] Memory 1344, NIC 1330, storage disk 1346, and microprocessor 1310 may provide an operating environment for a software stack including an operating system kernel 1314 executing in kernel space. For example, kernel 1314 may represent the kernel of Linux, Berkeley Software Distribution (BSD), another Unix variant, or a Windows Server operating system kernel commercially available from Microsoft Corporation. In some instances, the operating system may run a hypervisor and one or more virtual machines managed by the hypervisor. Exemplary hypervisors include the Kernel-based Virtual Machine (KVM) for the Linux kernel, Xen, ESXi commercially available from VMware, Windows Hyper-V commercially available from Microsoft, and other open source and proprietary hypervisors. The term hypervisor may encompass a Virtual Machine Manager (VMM). The operating system including kernel 1314 provides an execution environment for one or more processes in user space 1345. Kernel 1314 includes a physical driver 1327 for using network interface card 1330.

[0185] Computing device 1300 may be coupled to a physical network fabric that includes an overlay network of software or virtual routers that extends the fabric from a physical switch to physical servers coupled to the fabric (such as virtual router 21). Computing device 1300 may be configured with one or more dedicated virtual networks for the slave nodes of a cluster.

[0186] API server 300A, scheduler 1322, controller 406A, custom API server 301A, custom resource controller 302A, controller manager 1326, and configuration store 1328 may implement the master nodes of a cluster and may alternatively be referred to as "master components". The cluster may be a Kubernetes cluster, and the master nodes may be Kubernetes master nodes, in which case the master components are Kubernetes master components.

[0187] Each of the API server 300A, the controller 406A, the custom API server 301A, and the custom resource controller 302A includes code executable by the microprocessor 1310. The custom API server 301A validates and configures data for custom resources regarding the SDN architecture configuration. A service can be an abstraction that defines a logical collection of container pools and policies for accessing the container pools. A set of container pools implementing the service is selected based on the service definition. The service can be partially implemented as a load balancer or otherwise include a load balancer. The API server 300A and the custom API server 301A can implement a Representational State Transfer (REST) interface to handle REST operations and provide a front end as part of the configuration plane of the SDN architecture to the corresponding cluster shared state stored in the configuration store 1328. The API server 300A can represent a Kubernetes API server.

[0188] The configuration store 1328 is a backup store for all cluster data. The cluster data can include cluster state and configuration data. The configuration data can also provide a backend for service discovery and / or provide a locking service. The configuration store 1328 can be implemented as a key-value store. The configuration store 1328 can be a central database or a distributed database. The configuration store 1328 can represent an etcd store. The configuration store 1328 can represent a Kubernetes configuration store.

[0189] The scheduler 1322 includes code executable by the microprocessor 1310. The scheduler 1322 can be one or more computer processes. The scheduler 1322 monitors newly created or requested virtual execution elements (e.g., container pools in containers) and selects a slave node on which the virtual execution element runs. The scheduler 1322 can select the slave node based on resource requirements, hardware constraints, software constraints, policy constraints, locality, etc. The scheduler 1322 can represent a Kubernetes scheduler.

[0190] Typically, the API server 300A can call the scheduler 1322 to schedule the container pool. The scheduler 1322 can select a worker node and return the identifier of the selected worker node to the API server 300A. The API server 300A can write the identifier to the configuration store 1328 associated with the container pool. The API server 300A can call the orchestration agent for the selected worker node, so that the container engine 208 for the selected worker node obtains the container pool from the storage server and creates a virtual execution element on the worker node. The orchestration agent 310 for the selected worker node can update the status of the container pool to the API server 1320, and the API server 300A persists the new status to the configuration store 1328. In this way, the computing device 1300 instantiates a new container pool in the computing infrastructure 8.

[0191] The controller manager 1326 includes code executable by the microprocessor 1310. The controller manager 1326 can be one or more computer processes. The controller manager 1326 can be embedded in the kernel control loop to monitor the shared state of the cluster by obtaining notifications from the API server 300A. The controller manager 1326 can attempt to move the state of the cluster towards the desired state. The exemplary controller 406A and the custom resource controller 302A can be managed by the controller manager 1326. Other controllers can include the replication controller, the endpoint controller, the namespace controller, and the service account controller. The controller manager 1326 can perform lifecycle functions such as namespace creation and lifecycle, event garbage collection, terminated container pool garbage collection, cascading delete garbage collection, node garbage collection, etc. The controller manager 1326 can represent the Kubernetes controller manager for the Kubernetes cluster.

[0192] The network controller for the SDN architecture described herein can provide a cloud network for a computing architecture operating on top of a network infrastructure. The cloud network can include a private cloud for an enterprise or service provider, infrastructure as a service (IaaS), and a virtual private cloud (VPC) for a cloud service provider (CSP). Private cloud, VPC, and IaaS use cases may involve a multi-tenant virtualized data center such as described in the reference Figure 1 In this case, multiple tenants located within the data center share the same physical resources (physical servers, physical storage, physical network). Each tenant is allocated its own logical resources (virtual machines, containers, or other forms of virtual execution elements; virtual memory; virtual network). These logical resources are isolated from each other unless explicitly permitted by the security policy. The virtual network located within the data center can also be interconnected with a physical IP VPN or L2VPN.

[0193] A network controller (or "SDN controller") can provide network function virtualization (NFV) for a network, such as a service edge network, a broadband subscriber management edge network, and a mobile edge network. NFV involves the orchestration and management of network functions in virtual machines, containers, or other virtual execution elements, rather than in physical hardware devices, such as firewalls, intrusion detection or prevention systems (IDS / IPS), deep packet inspection (DPI), caching, wide area network (WAN) optimization, etc.

[0194] The SDN controller manager 1325 includes code executable by the microprocessor 1310. The SDN controller manager 1325 can be one or more computer processes. The SDN controller manager 1325 operates as an interface between the orchestration-oriented elements (e.g., the scheduler 1322, the API server 300A and the custom API server 301A, the controller manager 1326, and the configuration store 1328). Generally, the SDN controller manager 1325 monitors a cluster of new Kubernetes native objects (e.g., container pools and services). The SDN controller manager 1325 can isolate container pools in a virtual network and connect the container pools to services.

[0195] The SDN controller manager 1325 can run as a container on the master node of a cluster. As described herein, in some cases, using the SDN controller manager 1325 can enable the disabling of the service proxy (e.g., the Kubernetes kube-proxy) on the slave nodes so that all container pool connections are implemented using a virtual router.

[0196] The components of the network controller 24 can operate as a CNI for Kubernetes and can support multiple deployment modes. The CNI 17, CNI 750 are compute node interfaces for the entire CNI framework for managing the network of Kubernetes. The deployment modes can be divided into two categories: (1) an SDN architecture cluster as a CNI integrated into the workload Kubernetes cluster; and (2) an SDN architecture cluster as a CNI separate from the workload Kubernetes cluster.

[0197] Integration with the workload Kubernetes cluster

[0198] Run the components of the network controller 24 (e.g., the custom API server 301, the custom resource controller 302, the SDN controller manager 1325, and the control node 232) in the management Kubernetes cluster on the master node close to the Kubernetes controller components. In this mode, the components of the network controller 24 are actually part of the same Kubernetes cluster as the workload.

[0199] Detached from the workload Kubernetes cluster

[0200] The components of the network controller 24 are executed by a separate Kubernetes cluster from the workload Kubernetes cluster.

[0201] The SDN controller manager 1325 can use the controller framework for the orchestration platform to listen for (or otherwise monitor) changes to objects defined in the Kubernetes native API and add annotations to some of these objects. The annotations can be tags or other identifiers that indicate object characteristics (e.g., "virtual network green"). The SDN controller manager 1325 is a component of the SDN architecture, i.e., it listens for events of Kubernetes kernel resources (such as container pools, network policies, services, etc.) and converts these events into custom resources of the SDN architecture configuration as needed. The CNI plugins (e.g., CNI 17, 570) are SDN architecture components that support the Kubernetes network plugin standard: Container Network Interface.

[0202] The SDN controller manager 1325 can use the REST interface exposed by the aggregated API 402 to create network solutions for applications to define network objects such as virtual networks, virtual network interfaces, and access control policies. For example, the network controller 24 components can implement a network solution in the computing infrastructure by configuring one or more virtual networks and virtual network interfaces in the virtual router. (This is just an example of SDN configuration).

[0203] The following exemplary deployment configuration for the application consists of a container pool and the virtual network information of the container pool:

[0204]

[0205] This metadata information can be copied to each container pool replica created by the controller manager 1326. When these container pools are notified to the SDN controller manager 1325, the SDN controller manager 1325 can create the virtual networks listed in the annotations ("red-network", "blue-network", and "default / extns-network" in the above embodiments) and create virtual network interfaces for each container pool replica (e.g., container pool 202A) with unique private virtual network addresses from the cluster-wide address block of the virtual network (e.g., 10.0 / 16).

[0206] The following describes additional techniques according to the present disclosure. Contrail is an exemplary network controller architecture. Contrail CNI can be a CNI developed for Contrail. The cloud-native Contrail controller can be an example of the network controller described in the present disclosure, such as network controller 24.

[0207] Figure 7A is a block diagram showing a control / routing plane for an underlying network and an overlay network configuration using an SDN architecture according to the techniques of the present disclosure. Figure 7B is a block diagram showing a configured virtual network using a tunnel-connected container pool configured in an underlying network according to the techniques of the present disclosure.

[0208] The network controller 24 for the SDN architecture can use a distributed or centralized routing plane architecture. The SDN architecture can use a containerized routing protocol daemon (process).

[0209] In terms of network signaling, the routing plane can operate according to a distributed model, where the cRPD runs on each computing node in the cluster. Substantially, this means building intelligence into the computing nodes and the intelligence involves complex configurations at each node. In this model, the route reflector (RR) may not be able to make intelligent routing decisions, but is used as a repeater to reflect the routes between nodes. The distributed container routing protocol daemon (cRPD) is a routing protocol process that can be used, where each computing node runs an instance of its own routing daemon. At the same time, a centralized cRPD master instance can be used as an RR to relay routing information between computing nodes. The RR at a central location distributes routing and configuration intelligence across nodes.

[0210] Alternatively, the routing plane can operate according to a more centralized model, where components of the network controller run centrally and absorb the intelligence required to process configuration information, build a network topology, and program the forwarding plane into the virtual router. The virtual router agent is a local agent that processes the information programmed by the network controller. This design results in promoting more limited intelligence required at the computing nodes and tends to produce a simpler configuration state.

[0211] The centralized control plane provides the following:

[0212] · Allows the proxy routing framework to be simpler and lighter. Hides the complexity and limitations of BGP from the proxy. The proxy does not need to understand concepts such as route distinguishers, route targets, etc. The proxy simply exchanges prefixes and builds its corresponding forwarding information.

[0213] · There are indeed more control nodes than routers. The control nodes are built on the concept of a virtual network and can generate new routes using route replication and reorganization (e.g., supporting features such as service chaining and internal-VN routing for other use cases).

[0214] · Build a BUM tree for optimal broadcast and multicast forwarding.

[0215] It should be noted that the control plane has a distributed nature in certain aspects. As a control plane that supports distributed functions, it allows each local virtual router agent to announce its local routes and subscribe to configurations as needed.

[0216] The following functions can be provided by the cRPD of the network controller 24 or the control node.

[0217] Routing daemon / process

[0218] The control node and the cRPD can be used as routing daemons that implement different protocols and have the ability to program routing information in the forwarding plane.

[0219] The cRPD implements a routing protocol with a rich routing stack, which includes interior gateway protocols (IGPs) (e.g., Intermediate System to Intermediate System (IS-IS)), BGP-LU, BGP-CT, SR-MPLS / SRv6, Bidirectional Forwarding Detection (BFD), Path Computation Element Protocol (PCEP), etc. It can also be deployed to provide services such as route reflectors only for the control plane and is popular in Internet routing use cases due to these capabilities.

[0220] The control node 232 also implements a routing protocol, but mainly based on BGP. The control node 232 understands the overlay network. The control node 232 provides a rich feature set during overlay virtualization and caters to SDN use cases. Overlay features such as virtualization (using the abstraction of virtual networks) and service chaining are very popular among telecommunications and cloud providers. In some cases, the cRPD may not include support for this overlay function. However, the rich feature set of the cRPD provides strong support for the underlying network.

[0221] Network orchestration / automation

[0222] The routing function is only a part of the control node 232. The overall part of the overlay network is the orchestration. In addition to providing overlay routing, the control node 232 helps to model the orchestration function and provide network automation. The core of the orchestration ability of the control node 232 is the ability to model network virtualization using abstractions based on virtual networks (and related objects). The control node 232 interfaces with the configuration node 230 to relay configuration information to the control plane and the data plane. The control node 232 also helps to build an overlay tree for multicast layer 2 and layer 3. For example, the control node can build the virtual topology of the cluster it uses for this purpose. Generally, the cRPD does not include this orchestration ability.

[0223] High Availability and Horizontal Scalability

[0224] The control node design is more centralized, while the cRPD is more distributed. There are cRPD worker nodes running on each computing node. On the other hand, the control node 232 does not run computationally and can even run on a remote cluster (i.e., independent and in some cases geographically far from the workload cluster). The control node 232 also provides horizontal scalability for HA and runs in an active-active mode. The computing load is shared among the control nodes 232. On the other hand, the cRPD generally does not provide horizontal scalability. The control node 232 and the cRPD can provide a smooth restart for HA and can allow data plane operations in headless mode - where the virtual router can run even if the control plane restarts.

[0225] The control plane should be more than just a routing daemon. The control plane should support overlay routing and network orchestration / automation, while the cRPD is as good as a routing protocol in managing the underlying routing. However, the cRPD generally lacks network orchestration ability and does not provide strong support for overlay routing.

[0226] Accordingly, in some examples, the SDN architecture can have a cRPD on computing nodes as shown in Figures 7A to 7B the figure. Figure 7A An SDN architecture 700 is shown that can represent an exemplary implementation of the SDN architecture 200 or 400. In the SDN architecture 700, the cRPD 324 runs on the computing nodes and provides the underlying routing to the forwarding plane while running a set of centralized (and horizontally scalable) control nodes 232 that provide orchestration and overlay services. In some examples, instead of running the cRPD 324 on the computing nodes, a default gateway can be used.

[0227] The cRPD 324 on the compute node provides rich underlying routing to the forwarding plane by interacting with the virtual router agent 514 using interface 540, which can be a gRPC interface. The virtual router agent interface can permit programming of the routing, configuration of virtual network interfaces for overlays, and otherwise configuring the virtual router 506. This is described in further detail in U.S. Application No. 17 / 649,632. Meanwhile, one or more control nodes 232 operate as a pool of independent containers providing overlay services. Thus, the SDN architecture 700 can obtain the rich overlay and orchestration provided by the control nodes 232 and the modern underlying routing that supplements the control nodes 232 provided by the cRPD 324 on the compute node. An independent cRPD controller 720 can be used to configure the cRPD 324. The cRPD controller 720 can be a device / element management system, a network management system, an orchestrator, a user interface / CLI, or other controller. The cRPD 324 runs a routing protocol and exchanges routing protocol messages with routers, including other cRPD 324s. Each cRPD 324 can be a centralized routing protocol process and effectively operate as a software-only version of a router control plane.

[0228] The enhanced underlying routing provided by the cRPD 324 can replace the default gateway of the forwarding plane and provide a rich routing stack for supported use cases. In some examples where the cRPD 324 is not used, the virtual router 506 relies on the default gateway of the underlying routing. In some examples, the cRPD 324, which is an underlying routing process, is restricted to programming only the default inet(6).0 fabric using control plane routing information. In this example, non-default overlay VRFs can be programmed by the control node 232.

[0229] Figures 7A to 7B A dual routing / control plane solution as described above is shown. In Figure 7A In one aspect, similar to how a router control plane programs a router forwarding / data plane, the cRPD 324 provides underlying routing / forwarding information to the virtual router agent 514.

[0230] As Figure 7BAs shown, the cRPD 324 exchange can be used to create routing information for tunnels over the underlying network 702 of the VRF. Tunnel 710 is an example and connects the virtual routers 506 of servers 12A and 12X. Tunnel 710 can represent a Segment Routing (SR) or SRv6 tunnel, a Generic Routing Encapsulation (GRE) tunnel, and an IP-in-IP tunnel, an LSP, or other tunnels. The control node 232 uses tunnel 710 to create a virtual network 712 that connects the server 12A attached to the VRF of the virtual network and the container pool 22 of server 12X.

[0231] As described above, the cRPD 324 and the virtual router agent 514 can exchange routing information using the gRPC interface, and the virtual router agent 514 can program the virtual router 506 through configuration using the gRPC interface. It should also be noted that the control node 232 can be used for overlay and orchestration, while the cRPD 324 can be used to manage the underlying routing protocol. While communicating with the control node and the Domain Name Service (DNS) using XMPP, the virtual router agent 514 can use the gRPC interface with the cRPD 324.

[0232] Because there may be workers running on each computing node, the gRPC model performs well for the cRPD 324. And the virtual router agent 514 serves as a gRPC server that exposes the service of the client (cRPD 324) for programming the routing and configuration information (underlying). Thus, gRPC is an attractive solution when compared with XMPP. Specifically, gRPC transmits data as a binary stream and does not add overhead when encoding / decoding the data sent through gRPC.

[0233] In some examples, the control node 232 can use XMPP to interface with the virtual router agent 514. Since the virtual router agent 514 serves as a gRPC server and the cRPD 324 serves as a gRPC client. This means that the client (cRPD) needs to initiate a connection to the server (vRouter agent). In the SDN architecture 700, the virtual router agent 514 selects the set of control nodes 232 to which it subscribes (since there are multiple control nodes). In this regard, the control node 232 serves as a server and the virtual router agent 514 connects as a client and subscribes for updates.

[0234] The control node 232 needs to select the virtual router agent 514 it needs to connect to via gRPC and then subscribe as a client. Since the control node 232 does not run on each computing node, this requires implementing an algorithm to select the virtual router agent 514 it can subscribe to. Further, the control nodes 232 need to synchronize this information with each other. When a restart occurs, this also complicates the situation, and the control nodes 232 need to synchronize to select the agents they serve. Features such as graceful restart (GP) and fast convergence have been implemented on top of XMPP. XMPP is already lightweight and efficient. Therefore, XMPP may be superior to gRPC for the communication between the control node 232 and the virtual router agent 514.

[0235] The control node 232 and its additional uses are enhanced as follows. HA and horizontal scalability have three control nodes. As with any routing platform, having only two control nodes 232 to meet the HA requirements is sufficient. In many cases, this is advantageous. (However, one or more control nodes 232 can be used). For example, it provides a more deterministic infrastructure and conforms to the best practices of standard routing. Each virtual router agent 514 is attached to a unique pair of control nodes 232 to avoid randomness. With two control nodes 232, debugging may be simpler. In addition, with only two control nodes 232, edge replication for constructing multicast / broadcast trees can be simplified. Currently, since the vRouter agent 314 is only connected to two of the three control nodes, all control nodes may not have a complete picture of the tree at a certain time and rely on BGP to synchronize the state between control nodes. Since the virtual router agent 314 can randomly select two, the three control nodes 232 exacerbate this situation. If there are only two control nodes 232, each virtual router agent 314 is then connected to the same control node. This in turn means that the control nodes 232 do not need to rely on BGP to synchronize the state and have the same picture of the multicast tree.

[0236] The SDN architecture 200 can provide ingress replication as an alternative to edge replication and offer options to users. Ingress replication can be regarded as a special degenerate case of a general overlay multicast tree. However, in practice, the signaling of the ingress replication tree is simpler than that of the general overlay multicast tree. With ingress replication, each virtual router 21 ends up with a tree with itself as the root and each other vrouter as a leaf. A failed virtual router 21 should theoretically not cause the tree to be rebuilt. It should be noted that the performance of ingress replication deteriorates due to a large cluster. However, ingress replication performs well for a smaller cluster. Further, multicast is not a popular and common requirement for many customers. It is mainly limited to transporting only the initially occurring broadcast BUM traffic.

[0237] Enhanced Configuration Processing Module

[0238] In a conventional SDN architecture, the network controller processes the orchestration of all use cases. The configuration node converts the intent into a configuration object based on the data model and writes the configuration object to a database (e.g., Cassandra). In some cases, for example, notifications are sent via RabbitMQ to all clients waiting for configuration at the same time.

[0239] The control node not only acts as a BGP speaker but also has a configuration processing module that reads configuration objects from the database in the following ways. First, when the control node starts (or restarts), it connects to the database and directly reads all configurations from the database. Second, the control node can also be a message client. When there is an update to a configuration object, the control node receives a message notification listing the objects that have been updated. This again causes the configuration processing module to read the objects from the database.

[0240] The configuration processing module reads the configuration objects of the control plane (BGP-related configuration) and the vRouter forwarding plane. The configuration can be stored as a graph with objects as nodes and relationships as links. Then, this graph can be downloaded to the client (BGP / cRPD and / or vRouter agent).

[0241] According to the technology of the present disclosure, in some examples, the conventional configuration API server and message service are replaced by the kubeapi-server (API server 300 and custom API server 301), and the previous Cassandra database is replaced by etcd in Kubernetes. Due to this change, clients interested in configuration objects can directly establish monitoring on the etcd database to obtain updates instead of relying on RabbitMQ notifications.

[0242] Orchestration of cRPD Controller

[0243] BGP configuration can be provided to the cRPD 324. In some examples, the cRPD controller 720 can be a Kubernetes controller that caters to developing its own controller and implements the CRD required for orchestrating and provisioning the cRPD 324, and its own controller caters to the Kubernetes space.

[0244] Distributed Configuration Processing

[0245] As previously mentioned in this section, the configuration processing module can be part of the control node 232. The configuration processing module reads the configuration directly from the database, converts the data into JSON format, and stores it in its local IFMAP database as a graph with objects as nodes and relationships between them as links. Then, this graph is downloaded via XMPP to the interested virtual router agent 514 on the compute node. The virtual router agent 514 also locally constructs an IFMAP-based dependency graph to store these objects.

[0246] By having the virtual router agent 514 directly monitor the etcd server in the API server 300, the need for IFMAP as an intermediate module and for storing the dependency graph can be avoided. The cRPD 324 running on the compute node can use the same model. This will avoid the need for the IFMAP-XMPP configuration channel. The Kubernetes configuration client (for the control node 232) can be used as part of this configuration. The virtual router agent can also use this client.

[0247] However, this can increase the number of clients reading the configuration from the etcd server, especially in a cluster with hundreds of compute nodes. Adding more monitors will ultimately lead to a decrease in the write rate and the event rate dropping below the desired value. The gRPC proxy of etcd replays from one server monitor to multiple client monitors. The gRPC proxy merges multiple client monitors (c-monitors) for the same key or range into a single monitor (s-monitor) connected to the etcd server. The proxy broadcasts all events from the s-monitor to its c-monitors. Assuming N client monitors the same key, one gRPC proxy can reduce the monitoring load on the etcd server from N to 1. The user can deploy multiple gRPC proxies to further distribute the server load. These clients share a server monitor; the proxy effectively offloads resource pressure from the core cluster. By adding proxies, etcd can serve one million events per second.

[0248] DNS / Naming in the SDN Architecture

[0249] In the previous architecture, the DNS service was provided by the jointly working contrail-dns and contrail-naming processes to provide DNS services to VMs in the network. Naming serves as a DNS server providing an implementation of the BIND protocol, and contrail-dns receives updates from the vrouter-agent and pushes these records to Naming.

[0250] Four DNS modes are supported in the system, and the IPAM configuration can select the required DNS mode.

[0251] 1. No-VM-supported DNS.

[0252] 2. Default DNS server - DNS resolution for VMs is completed based on the name server configuration in the server infrastructure. When a VM obtains a DHCP response, the subnet default gateway is configured as the VM's DNS server. DNS requests sent by the VM to this default gateway are resolved by the (structured) name servers configured on the corresponding compute nodes, and the response is sent back to the VM.

[0253] 3. Tenant DNS server - Tenants can use their own DNS servers in this mode. A list of servers can be configured in IPAM and then sent to the VM as the DNS server in the DHCP response. DNS requests sent by the VM are routed as any other data packet based on the available routing information.

[0254] 4. Virtual DNS server - In this mode, the system supports virtual DNS servers to provide DNS servers for resolving DNS requests from VMs. Those skilled in the art can define multiple virtual domain name servers under each domain in the system. Each virtual domain name server is the authoritative server for the configured DNS domain.

[0255] The SDN architecture described here is effective in the DNS services it provides. Customers in the cloud-native world benefit from various DNS services. However, as moving to a next-generation Kubernetes-based architecture, the SDN architecture can be replaced by a core DNS for using any DNS service.

[0256] Data plane

[0257] The data plane consists of two components: the virtual router agent 514 (aka agent) and the virtual router forwarding plane 506 (also known as the DPDK vRouter / kernel vRouter). The agent 514 in the SDN architecture solution is responsible for managing the data plane components. The agent 514 establishes XMPP neighbor relationships with two control nodes 232 and then exchanges routing information with the two control nodes 232. The vRouter agent 514 also dynamically generates flow entries and injects the flow entries into the virtual router 506. This gives instructions to the virtual router 506 on how to forward packets.

[0258] The responsibilities of proxy 514 can include: interfacing with the control node 232 to obtain configuration. Converting the received configuration into a form that the data path can understand (e.g., converting the data model from IFMAP into the data model used by the data path). Interfacing with the control node 232 to manage routing. And collecting statistics from the data path and outputting the statistics to the monitoring solution.

[0259] The implementation of virtual router 506 can allow virtual network interfaces to have data plane functions associated with VRFs. Each VRF has its own forwarding and flow tables, while the MPLS and VXLAN tables are global within virtual router 506. The forwarding table can contain routes for the IP and MAC addresses of destinations and the IP-to-MAC association is used to provide proxy ARP capabilities. When a VM / container interface starts and is only important locally to the vRouter, the virtual router 506 selects the label value in the MPLS table. The VXLAN network identifier is global across all VRFs of the same virtual network in different virtual routers 506 within a domain.

[0260] In some examples, each virtual network has a default gateway address assigned to it, and each VM or container interface receives this address in the DHCP response when it is initialized. When a workload sends a packet to an address outside its subnet, the ARP for the MAC will correspond to the IP address of the gateway, and the virtual router 506 responds with its own MAC address. Thus, the virtual router 506 can support a fully distributed default gateway function for all virtual networks.

[0261] The following is an example of packet flow forwarding implemented by virtual router 506.

[0262] Packet flow between VM / container interfaces in the same subnet.

[0263] The worker node can be a VM or a container interface. In some examples, the packet processing is as follows:

[0264] · The VMI / container interface needs to send a packet to VM2. Therefore, the virtual router 506 first queries its own DNS cache to obtain the IP address. However, since this is the first packet, there is no entry.

[0265] · When its interface starts, VM1 sends a DNS request to the DNS server address supplied in the DHCP response.

[0266] · The virtual router 506 intercepts the DNS request and forwards it to the DNS server running in the SDN architecture controller.

[0267] · The DNS server in the controller responds with the IP address of VM2.

[0268] · The virtual router 506 sends the DNS response to VM1.

[0269] · VM1 needs to form an Ethernet frame and thus requires the MAC address of VM2. VM1 checks its own ARP cache, but since this is the first packet, there is no entry.

[0270] · VM1 issues an ARP request.

[0271] · The virtual router 506 intercepts the ARP request and looks up the MAC address of IP-VM2 in its own forwarding table and finds the association in the L2 / L3 route where the controller sends the MAC address for VM2.

[0272] · The virtual router 506 sends an ARP reply to VM1 using the MAC address of VM2.

[0273] · A TCP timeout occurs in the network stack of VM1.

[0274] · The network stack of VM1 reattempts to send the packet, and at this time, finds the MAC address of VM2 in the ARP cache and is able to form an Ethernet frame and send the Ethernet frame out.

[0275] · The virtual router 506 queries the MAC address of VM2 and finds the encapsulation route. The virtual router 506 constructs an outer header and sends the generated packet to the server S2.

[0276] · The virtual router 506 on the server S2 decapsulates the packet and looks up the MPLS label to identify the virtual interface to send the original Ethernet frame into. The Ethernet frame is sent into the interface and received by VM2.

[0277] Packet flow between VMs in different subnets

[0278] In some examples, except that the virtual router 506 responds as the default gateway, the order of sending packets to destinations in different subnets is the same. VM1 sends a packet with the MAC address of the default gateway in the Ethernet frame, and when VM1 starts, the IP address of the default gateway is provided in the DHCP response supplied by the virtual router 506. When VM1 makes an ARP request for the gateway IP address, the virtual router 506 responds with its own MAC address. When VM1 sends an Ethernet frame using the gateway MAC address, the virtual router 506 uses the destination IP address of the packet within the frame to look up the forwarding table in the VRF to find the route to the host running via the encapsulation tunnel to the destination.

[0279] Figure 10 It is a block diagram showing a multi-cluster deployment of a cloud-native SDN architecture according to the technology of the present disclosure. As used herein, the term "cluster" may refer to a Kubernetes cluster or other terms for similar structures in an orchestration platform used to implement the SDN architecture described in the present disclosure.

[0280] In a multi-cluster deployment with respect to the SDN architecture 1000, configuration nodes and control nodes are deployed to a central cluster 902 and the configuration and control of multiple distributed workload clusters 930-1 to 930-N (collectively "workload clusters 930") are centrally managed. However, the data plane is distributed among the workload clusters 930. As described with reference to other SDN architectures in the present disclosure, each of the workload clusters 930 and the central cluster 902 uses similar component microservices to implement the data plane of the cluster. For example, the workload cluster 930-1 includes a virtual router and a virtual router agent deployed to the computing nodes that make up the cluster. The computing nodes of the workload cluster 930-1 may also include a CNI, an orchestration agent, and / or other components described with reference to the server 12 and other computing nodes elsewhere in the present disclosure. The workload cluster 930-1 also includes an API server 300-1 for creating and managing local resources of the orchestration platform. The API server 300-1 may be similar to Figure 3 the API server 300 in. Among them, Kubernetes is an orchestration platform. In addition to the API server 300-1 (kube-api server), other orchestration platform components may include kube-scheduler, kube-controller-manager, and kubelet. The configuration store 920-1 may be similar to Figure 3 one or more of the configuration stores 304 in.

[0281] The workload clusters 930 may be geographically distributed to the network edge, such as an edge or a micro data center, or distributed towards the network edge of an edge or a micro data center, while the central cluster 902 may be geographically consolidated / centralized in, for example, a regional data center or a cloud provider. The computing nodes of each of the workload clusters 930 may be quite different from the computing nodes of the central cluster 902, and the computing nodes between the clusters do not overlap.

[0282] In the multi-cluster mode for multi-cluster deployment of the SDN architecture 1000, the central cluster 902 runs a complete installation of the orchestration platform including the API server 300-C and other orchestration platform components (possibly similar to those described above with reference to the workload cluster 930-1). The central cluster 902 also runs the configuration plane and control plane component microservices of the network controller for the SDN architecture. Specifically, the configuration node 230 includes one or more instances of each of the API server 300-C, the custom API server 301, the custom resource controller 302, and the SDN controller manager 303-C. These can be Figure 3 and Figure 4 exemplary instances in and similar similarly named components.

[0283] The API server 300-C and the custom API server 301 constitute an aggregated API server for SDN architecture resources (virtual networks, virtual machine interfaces, routing instances, etc.). The custom API server 301 can be registered as an APIService with the API server 300-C. As described above, requests for SDN architecture resources are received by the API server 300-C and forwarded to the custom API server 301, which performs operations on the custom resources of the SDN architecture configuration.

[0284] The custom resource controller 302 of the central cluster 902 implements the business logic of the custom resources of the SDN architecture configuration. The custom resource controller 302 converts the user intent into low-level resources consumed by the control node 232.

[0285] One or more instances of the control node 232 operate as SDN architecture control nodes for managing the data plane of the central cluster 902, the workload cluster 930-1, and the workload cluster 930-N. The central cluster 902 includes the virtual router 910-C deployed to the compute nodes of the central cluster 902, the workload cluster 930-1 includes the virtual router 910-1 deployed to the compute nodes of the workload cluster 930-1, and the workload cluster 930-N includes the virtual router 910-N deployed to the compute nodes of the workload cluster 930-N. For ease of illustration, the virtual routers located on separate compute nodes are not shown. As detailed elsewhere in this disclosure, different instances of the virtual router run on the compute nodes of the corresponding cluster. For example, each such compute node can be a server 12.

[0286] The virtual network 1012 represents one or more virtual networks that can be configured by the configuration node 230 and the control node 232. The virtual network 1012 can enable workloads running in different workload clusters 930 and in the same workload cluster to connect. Workloads of different virtual networks or the same virtual network can communicate across the workload clusters 930 and within the same workload cluster.

[0287] The central cluster 902 includes respective SDN controller managers 303-1–303-N (collectively "SDN controller manager 303") for the workload clusters 930. The SDN controller manager 303-1 is an interface between the local resources of the orchestration platform (e.g., services, namespaces, container pools, network policies, network attachment definitions) and the custom resources configured in the SDN architecture, and more specifically, an interface between the local resources of the orchestration platform and the custom resources configured in the workload cluster 930-1.

[0288] The SDN controller manager 303-1 monitors the API server 300-1 of the workload cluster 930-1 for changes to the local resources of the orchestration platform of the workload cluster 930-1. The SDN controller manager 303-1 also monitors the custom API server 301 of the central cluster 902. This is referred to as "dual monitoring". For example, to implement dual monitoring, the SDN controller manager 303 can use the admiralty multicluster-controller go library provided by the Kubernetes community or a multicluster manager implementation, each of which supports the function of monitoring resources in multiple clusters. Due to dual monitoring, the SDN controller manager 303-1 performs operations on custom resources initiated at the API server 300-1 of the workload cluster 930-1 or at the 300-C / custom API server 301. In other words, the custom resources of the central cluster 902 can be created in the following ways: (1) directly or interactively by the interaction of the user or agent with the configuration node 230; or (2) indirectly by an event caused by the operation of the local resources of the orchestration platform of a workload cluster 930 and detected by a responsible SDN controller manager 303, and the SDN controller manager 303 can use the custom API server 301 to responsively create custom resources in the configuration store 920-C to implement the local resources of the workload cluster.

[0289] The SDN controller manager 303-N operates similarly for monitoring and workload cluster 930-N configurations. The SDN controller manager 303-C running on the central cluster 902 operates similarly, but since it is only responsible for docking the local resources of the API server 300 with the custom resources configured in the central cluster 902 (e.g., virtual router 910-C), it does not monitor the individual API servers in the workload clusters.

[0290] In the multi-cluster mode of the multi-cluster deployment of the SDN architecture 1000, each distributed workload cluster 930 is associated with the central cluster 902 via a dedicated SDN controller manager 303. The SDN controller manager 303 runs on the central cluster 902 to facilitate better lifecycle management (LCM) of the SDN controller manager 303, configuration nodes 230, and control nodes 232 and to facilitate better and more manageable handling of security and permissions by consolidating these tasks into a single central processor 902.

[0291] As described elsewhere in this disclosure, the virtual router agent of the virtual router 910 communicates with the control node 232 to obtain routing and configuration information. For example, when creating an orchestration platform local resource such as a container pool or service in the workload cluster 930-1, the SDN controller manager 303-1 running in the central cluster 902 receives an indication of the create event and its coordinator can create / update / delete custom resources of the SDN architecture configuration such as virtual machines, virtual machine interfaces, instance IPs. In addition, the SDN controller manager 303-1 can associate these new custom resources with the virtual network of the container pool or service. This virtual network can be the default virtual network or the virtual network indicated in the manifest of the container pool or service (user annotation).

[0292] Custom resources can have namespace scopes of different clusters. The custom resources associated with a workload cluster 930 will have a corresponding namespace created in the central cluster 902 and a cluster identifier (e.g., cluster name or unique identifier), and the custom resources are created according to this namespace. Cluster-wide custom resources can be stored with a naming convention such as clustername-resourcename-unique identifier. The uniqueidentifier can be a hash of the cluster name, the namespace of the resource, and the resource name.

[0293] Using resources appropriately associated with the corresponding workload cluster 930, e.g., using the cluster identifiers described above, the SDN controller manager 303 authenticates the user or agent and allows the user or agent to use only custom resources for SDN architecture configuration that belong to the namespace associated with the workload cluster. For example, if a user attempts to create a container pool in the workload cluster 930-1 by sending a request to the API server 300-1, and the annotation of the container pool manifest specifies a specific virtual network "VN1", then the SDN controller manager 303-1 for the workload cluster 930-1 will authenticate the request by determining whether "VN1" belongs to the namespace associated with the workload cluster 930-1. If valid, the SDN controller manager 303-1 then uses the custom API server 301 to create custom resources for the SDN architecture configuration. The control node 232 configures the configuration object for the new custom resources in the workload cluster 930-1. If invalid, the SDN controller manager 303-1 may use the API server 300-1 to delete the resources.

[0294] The user or agent can create custom resources for the SDN architecture configuration in the namespace using NetworkAttachmentDefinition. In this case, the custom resource controller 302 for the NetworkAttachmentDefinition custom resource can create a virtual network in the namespace with a name prefixed with the relevant cluster identifier. For example, create a virtual network in the namespace clustername-namespace. As described above, this allows the SDN controller manager 303 to perform authentication when the container pool annotation lists this virtual network.

[0295] In some embodiments, predicates can be implemented to filter events generated due to changes in custom resources for the SDN architecture configuration of other workload clusters 930. The coordinator of the SDN controller manager 303 for this workload cluster only needs to handle events related to the resources (i.e., namespace and cluster scope) owned by the specific workload cluster.

[0296] In some examples, each workload cluster 930 can have its own default Kubernetes container pool and service subnet, which should not overlap with the container pools and service subnets of other workload clusters 930. In some cases, although it does not support Network Address Translation (NAT), however, it can support Kubernetes network policies with routing. In some examples, there is no container pool or service network across workload clusters. Each workload cluster can have its own pool. The Kubernetes administrator should set up the clusters accordingly.

[0297] Bind the virtual network to the workload cluster 930 using namespaces and stitch the virtual network with the network policy / network router as shown by the virtual network 1012. However, there is no common virtual network that extends across clusters. As long as the virtual network subnets are different and the routing targets are shared between the virtual networks, the container pools or container pool pings across the workload clusters 930 located on the default container pool network interface and on the secondary virtual network interface should work together.

[0298] The following is an exemplary workflow of the SDN controller manager 303-1 that starts the container pool for the workload cluster 930-1 (named "Cluster 1" in these examples). Other SDN controller managers 303 can operate similarly with respect to the central cluster 902 and its dedicated workload clusters 930. NetworkAttachmentDefinition is a CRD schema specified by the Kubernetes network plumbing working group to express the intention of attaching a container pool to one or more logical or physical networks. The NetworkAttachmentDefinition specification typically defines the required state of the network attachment for the secondary interface of the container pool. Then, the container pool can be started with a text annotation, and the Kubernetes CNI attaches the network interface to the container pool. As described elsewhere in this disclosure, the SDN architecture can operate as the CNI and attach the virtual network interface of the virtual network to the container pool based on the NAD and the container pool manifest annotation.

[0299] 1. The user creates a NetworkAttachmentDefinition (NAD) in the distributed cluster "Cluster 1" (e.g., workload cluster 930-1)

[0300]

[0301]

[0302] The SDN controller manager 303-1 monitors the creation event of the above NAD and then:

[0303] a. Checks whether the namespace "cluster1-ns1" exists. If not, the SDN controller manager 303-1 creates this namespace in the central cluster using the following:

[0304]

[0305] 2. The user creates a container pool in the distributed cluster "cluster1" with the default container pool network

[0306]

[0307]

[0308] The SDN controller manager 303-1 monitors the container pool creation event and then:

[0309] a. Checks whether the cluster1_default namespace exists, and if not, creates the cluster1_default namespace.

[0310] b. Creates the following resources in the central cluster:

[0311]

[0312]

[0313] 3. The user creates a container pool in the distributed cluster "cluster1" with VN ns1 / vn1

[0314]

[0315] The SDN controller manager 303-1 monitors the container pool creation event and then:

[0316] a. Checks whether a virtual network exists in the namespace = cluster1-ns1 with name = vn1. If it exists, continue.

[0317] b. Creates the following resources in the central cluster:

[0318]

[0319]

[0320] Lifecycle management

[0321] For the deployer of the central cluster 902, the user can manually create a Secret in the central cluster 902 using the kubeconfig file of the workload cluster. Pass the name of the created Secret as a value to the kubeconfigSecretName String defined in the CRD of the SDN controller manager 303-C. The deployer will mount the data available in the kubeconfigSecretName and pass it to the SDN controller manager 303 of the workload cluster. If the kubeconfigSecretName value is provided to the deployer, it is assumed that the SDN controller manager 303 is installed in multi-cluster mode. In multi-cluster mode, the deployer will obtain the PodSubnet and ServiceSubnet from the kubeadmConfigMap of the workload cluster. Each workload cluster 930 should have a unique name. To achieve this, the LCM will pick the custom resource name corresponding to the SDN controller manager 303 as the cluster name of the workload cluster 930. Since there are multiple deployment instances of the SDN controller manager 303 in the central cluster 902, the deployment name of each SDN controller manager 303 must be unique. For example, to achieve this, the deployment name of the SDN controller manager 303 will be created based on the custom resource name of the SDN controller manager 303, as the [Prefix]-k8s-Kubemanager-kubemanager CR name.

[0322] For example, for the deployer in the workload cluster 930-1, the user can manually create a Secret in the workload cluster 930-1 using the kubeconfig file of the central cluster 902. Then, during the deployment of the workload cluster 930-1, the above-created Secret is mounted and passed to the controller of the virtual router 910 component via the Deployment yaml. The controller uses the context of the central cluster 902 to access or modify SDN architecture resources and uses the in-cluster context to access Kubernetes resources. Assertions can be added to the controller of the virtual router 910 component to find the SDN controller manager 303 that is only related to this specific workload cluster, for example, related to the SDN controller manager 303-1 of the workload cluster 930-1. The user can configure the IP and XMPP ports of one or more control nodes 232 for the custom resource of the virtual router 910. When this custom resource is applied, the deployer will connect the virtual router 910 of the workload cluster (for example, an instance of the virtual router 910-1) to the appropriate control node 232 of the central cluster 902.

[0323] In other words, the core components of the SDN architecture 1000 are divided into three planes: the configuration plane, the control plane, and the data plane. The configuration plane includes the components of the configuration node 230, including the SDN controller manager 303-C, the custom API server 301, and the custom resource controller 302 (which are respectively referred to as contrail-k8s-kubemanager, contrail-k8s-apiserver, contrail-k8s-controller, where the SDN architecture 1000 is based on Contrail). The control plane includes the components of the control node 232 (referred to as contrail-control, where the SDN architecture 1000 is based on Contrail). The data plane includes the components of the virtual router 910 (referred to as contrail-vrouter, where the SDN architecture 1000 is based on Contrail).

[0324] The data plane components are installed in the workload cluster 930, and the control plane and configuration components are run in the central cluster 902. It may be necessary to synchronize the configuration data of the SDN controller manager 303, the control node 232 (such as the SMPP port and the control node IP, Pod CIDR), and the virtual router 910 components between different clusters:

[0325] 1. When there is an update to the configuration data;

[0326] 2. When adding a workload cluster to an existing Multicluster setup;

[0327] 3. When there is a change in the control node 232 of an existing Multicluster setup.

[0328] To solve this synchronization problem:

[0329] 1. Share the shared configuration data of the SDN controller manager 303, the control node 232, and the virtual router 910 components between clusters via the installed ConfigMap.

[0330] 2. The SDN controller manager 303 instance will collect data about the control node 232 and create a mapping of its associated workload cluster 930 and install a pool of containers for its virtual router 910 to use.

[0331] 3. If there is a newly added workload cluster, the relevant SDN controller manager 303 will synchronize the data of the control node 232 with the resources of its virtual router 910 via the ConfigMap.

[0332] 4. If the control node 232 changes, the affected SDN controller manager 303 will update the resources of its virtual router 910, and the affected virtual router container pool will restart to connect to the new control node 232.

[0333] Role-based access control

[0334] In the central cluster 902, the controllers running in the central cluster 902 can use a single SDN architecture ServiceAccount defined by the deployer. The ServiceAccount grants full access permissions to the resources in the central cluster 902. Since there are many SDN architecture custom resources within the cluster scope and the controllers running on the central cluster 902 can access most of the custom resources, this access cannot be restricted.

[0335] In the workload cluster 930, users can create cluster roles with restricted read / write access to the resources in the workload cluster. Users can create ServiceAccount and Cluster-role binding to connect the cluster role and the ServiceAccount. Users can manually generate a kubeconfig file using the token generated for the above ServiceAccount. As shown above, the central cluster 902 uses this kubeconfig file to communicate with the workload cluster.

[0336] Exemplary Cluster-Role, Service account, and Cluster-Role Binding are as follows:

[0337]

[0338]

[0339]

[0340] Test scenarios

[0341] · Verify whether the deployer can configure the SDN controller manager 303-C in single-cluster mode or multi-cluster mode.

[0342] · Verify whether users can create a container pool in the workload cluster 930 and whether they can create corresponding SDN architecture resources in the central cluster 902.

[0343] · Verify whether the user can create Services (ClusterIP and NodePort) in the workload cluster 930 and whether the user can create corresponding SDN architecture resources in the central cluster 902.

[0344] · Verify whether the user can delete Pods or Services in the workload cluster 930 and whether the user can delete references to the corresponding SDN architecture resources in the central cluster 902.

[0345] · Verify the SDN controller manager 303 coordinator function for automatic correction of lost SDN architecture resources.

[0346] · Verify the container pool network status annotation updates that occur in the container pool for each VirtualNetwork used.

[0347] · Verify the container pool connectivity within the distributed cluster using different Virtual Networks with a common RouteTarget.

[0348] · Verify the container pool connectivity between clusters using different Virtual Networks with a common RouteTarget.

[0349] · Verify whether the user can create Kubernetes network policies in the workload cluster 930 and whether the user can create corresponding SDN architecture configured firewall policy objects in the central cluster 902.

[0350] Each cluster in a multi-cluster deployment should have non-overlapping Service CIDRs. Due to the same service CIDR, the cluster will not have connectivity from the vRouter (data path) to an appropriate API server 300 and may not be able to provide services properly.

[0351] Each cluster in a multi-cluster deployment should have non-overlapping Pod CIDRs. Clusters with overlapping Pod CIDRs may have problems. For example, consider a case where one container pool in each workload cluster 930 has the same container pool IP. The two container pools will not be able to communicate with each other. When using the container pool as an endpoint to create a service in a workload cluster 930, there will be two routes in BGP: one route to each container pool in the workload cluster 930, and the traffic across the two container pools will be ECMP traffic, which is not desirable.

[0352] Figure 11A flowchart showing an exemplary pattern of operations for multi-cluster deployment of an SDN architecture. In the pattern of operation 1100, a network controller for a software-defined network (SDN) architecture system includes: processing circuitry for a central cluster of one or more first computing nodes; a configuration node configured to be executed by the processing circuitry; and a control node configured to be executed by the processing circuitry. The configuration node of the network controller includes a custom API server. The custom API server processes requests for operations on custom resources regarding SDN architecture configuration (1102). Each custom resource regarding SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system. In response to detecting an event for an instance of a first custom resource among the custom resources, the control node of the network controller obtains configuration data regarding the instance of the first custom resource and configures a corresponding instance of a configuration object in a workload cluster of one or more second computing nodes (1104). The one or more first computing nodes in the central cluster are distinct from the one or more second computing nodes in the workload cluster.

[0353] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. The various features described as modules, units, or components may be implemented together in an integrated logic device or separately as discrete, but interoperable, logic devices or other hardware devices. In some cases, the various features of an electronic circuit may be implemented as one or more integrated circuit devices, such as an integrated circuit chip or chip set.

[0354] If implemented in hardware, the present disclosure may relate to apparatus such as a processor or an integrated circuit device such as an integrated circuit chip or chip set. Alternatively or in addition, if implemented in software or firmware, the techniques may be implemented at least in part via a computer-readable data storage medium that includes instructions that, when executed, cause a processor to perform one or more of the methods described above. For example, the computer-readable data storage medium may store the instructions for execution by the processor.

[0355] The computer-readable medium may form part of a computer program product that includes packaging material. The computer-readable medium may include a computer data storage medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. In some examples, an article of manufacture may include one or more computer-readable storage media.

[0356] In some examples, a computer-readable storage medium may include a non-volatile medium. The term "non-volatile" may indicate that the storage medium is not included in a carrier wave or propagated signal. In a particular example, a non-volatile storage medium may store data that changes over time (e.g., stored to RAM or a cache).

[0357] The code or instructions may be software and / or firmware that are executed by processing circuitry including one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein, the term "processor" may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described in this disclosure may be provided within software modules or hardware modules.

Claims

1. A network controller for a software-defined network (SDN) architecture system, the network controller comprising: Processing circuitry for a central cluster of one or more computing nodes, wherein each of the one or more computing nodes is a physical computing device; A configuration node configured to be executed by the processing circuitry; and A control node configured to be executed by the processing circuitry; Wherein the configuration node includes a custom application programming interface (API) server to handle requests for operations on custom resources of the SDN architecture configuration, and includes an application programming interface (API) server to handle requests for operations on local resources of a container orchestration system, the API server being different from the custom API server; Wherein the API server is configured to receive a first request to operate on an instance of a first custom resource among the custom resources, and based on receiving the first request, delegate the first request to the custom API server for handling by the custom API server; Wherein each custom resource among the custom resources of the SDN architecture configuration corresponds to a type of configuration object in the SDN architecture system; and Wherein, in response to detecting an event of the instance of the first custom resource among the custom resources, the control node is configured to obtain configuration data of the instance of the first custom resource and configure a corresponding instance of a configuration object in a workload cluster of a second one or more computing nodes, wherein each of the second one or more computing nodes is a physical computing device, and wherein the first one or more computing nodes in the central cluster are different from the second one or more computing nodes in the workload cluster.

2. The network controller according to claim 1, Among them, The instance of the first custom resource is a first instance and the workload cluster is a first workload cluster; Wherein, in response to detecting an event of a second instance of the first custom resource, the control node is configured to obtain configuration data of the second instance of the first custom resource and configure a corresponding instance of a configuration object in a second workload cluster of a third one or more computing nodes, wherein the first one or more computing nodes, the second one or more computing nodes, and the third one or more computing nodes are different.

3. The network controller according to claim 1, further comprising: A first SDN controller manager configured to be executed by the processing circuitry; And A second SDN controller manager configured to be executed by the processing circuitry; Wherein the first SDN controller manager is configured to monitor events of instances of first resources in the central cluster and monitor events of instances of resources in a first workload cluster; and Wherein the second SDN controller manager is configured to monitor events of instances of second resources in the central cluster and monitor events of instances of resources in a second workload cluster.

4. The network controller according to claim 1, further comprising: an SDN controller manager configured to be executed by the processing circuit; wherein the SDN controller manager is configured to monitor events of instances of resources in the central cluster and monitor events of instances of resources in the workload cluster.

5. The network controller according to claim 1, an SDN controller manager configured to be executed by the processing circuit; Among them, in response to detecting a creation event of an instance of a local resource of the container orchestration system of the workload cluster, the SDN controller manager is configured to create a custom resource for the SDN architecture configuration in the central cluster using a custom API server.

6. The network controller according to claim 5, wherein, The event of the instance of the first custom resource includes the creation event of the instance of the first custom resource.

7. The network controller according to claim 1, an SDN controller manager configured to be executed by the processing circuit; Among them, in response to detecting a container pool, the SDN controller manager is configured to create an event for the container orchestration system in the workload cluster; create an instance of the first custom resource for the SDN architecture configuration in the central cluster using a custom API server, wherein the first custom resource represents a container pool object; and create an instance of a second custom resource for the SDN architecture configuration in the central cluster using the custom API server, wherein the second custom resource represents a virtual network interface object.

8. The network controller according to claim 1, further comprising: an SDN controller manager configured to be executed by the processing circuit, wherein the SDN controller manager is configured to verify a request for the custom resource of the SDN architecture configuration by determining whether the custom resource belongs to a namespace associated with the workload cluster.

9. The network controller according to claim 1, further comprising: an SDN controller manager configured to be executed by the processing circuit, wherein the SDN controller manager is configured to synchronize configuration data between the central cluster and the workload cluster using an installed ConfigMap.

10. The network controller according to claim 9, wherein, The synchronized configuration data includes an extensible message of the control node and one of a representation of the protocol XMPP, a port, and an Internet protocol address.

11. The network controller according to any one of claims 1 to 10, wherein, To configure the corresponding instance of the configuration object in the workload cluster of the second one or more computing nodes, the control node is configured to send the configuration data to a virtual router of the second one or more computing nodes to implement the corresponding instance of the configuration object in the virtual router.

12. A software-defined network method, comprising: processing, by an application programming interface API server implemented by a configuration node of a network controller for a software-defined network SDN architecture system, a request for an operation of a local resource of a container orchestration system The custom API server implemented by the configuration node processes requests for operations on custom resources of the SDN architecture configuration, where each type of custom resource in the custom resources of the SDN architecture configuration corresponds to the type of configuration object in the SDN architecture system. Wherein, the network controller operates on the central cluster of the first one or more computing nodes, and each of the first one or more computing nodes is a physical computing device. The API server receives a first request to operate on an instance of the first custom resource in the custom resources. The API server delegates the first request to the custom API server. The custom API server processes the first request. The control node of the network controller detects an event of an instance of the first custom resource in the custom resources; and In response to detecting the event of the instance of the first custom resource, the control node obtains the configuration data of the instance of the first custom resource and configures the corresponding instance of the configuration object in the workload cluster of the second one or more computing nodes, where each of the second one or more computing nodes is a physical computing device, and the first one or more computing nodes of the central cluster are different from the second one or more computing nodes of the workload cluster.

13. The method according to claim 12, Among them, The instance of the first custom resource is the first instance and the workload cluster is the first workload cluster, and the method further includes: In response to detecting an event of a second instance of the first custom resource, the control node obtains the configuration data of the second instance of the first custom resource and configures the corresponding instance of the configuration object in the second workload cluster of the third one or more computing nodes, where the first one or more computing nodes, the second one or more computing nodes, and the third one or more computing nodes are different.

14. The method according to any one of claims 12 to 13, further includes: The first SDN controller manager monitors events of instances of the first resource in the central cluster and monitors events of instances of resources in the first workload cluster; And The second SDN controller manager monitors events of instances of the second resource in the central cluster and monitors events of instances of resources in the second workload cluster.

15. The method according to any one of claims 12 to 13, further includes: In response to detecting a creation event of an instance of a local resource of the container orchestration system of the workload cluster, the SDN controller manager uses the custom API server to create the custom resource for the SDN architecture configuration in the central cluster.

16. The method according to claim 15, wherein, The event of the instance of the first custom resource includes the creation event of the instance of the first custom resource.

17. The method according to any one of claims 12 to 13, further comprising: validating, by the SDN controller manager, a request for the custom resource configured for the SDN architecture by determining whether the custom resource belongs to a namespace associated with the workload cluster.

18. The method according to any one of claims 12 to 13, further comprising: synchronizing, by the SDN controller manager, configuration data between the central cluster and the workload cluster using an installed ConfigMap.

19. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to perform the method according to any one of claims 12 to 18.

Citation Information

Patent Citations

  • Containerized router with virtual networking

    US12160811B2

  • Tunneled packet aggregation for virtual networks

    US9571394B1

  • Container-based network policy configuration in software-defined networking (SDN) environments

    US10944691B1

  • Container-based connectivity check in software-defined networking (SDN) environments

    US20210218652A1