Extending central cluster membership to additional computing resources

Intent-based networking mechanisms solve the compatibility and management challenges of adding new computing resources to the cluster in the campus network through automated configuration and authentication, thereby improving deployment efficiency and security.

CN113039520BActive Publication Date: 2026-01-20CISCO TECHNOLOGY INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980075111.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-09
Filing Date
2019-11-06
Publication Date
2026-01-20
Estimated Expiration
2040-07-14

AI Technical Summary

Technical Problem

Existing campus networks face device compatibility issues when deploying computing resource clusters, requiring manual configuration and management, which leads to a heavy workload for administrators and makes it difficult to quickly innovate and adopt new technologies.

Method used

Through an intent-based networking mechanism, existing cluster members notify new computing resources of the appropriate configuration and software version. The new resources automatically install the software and join the cluster, utilizing the management cloud, cluster members, package repositories, and credential authorities for automated configuration and authentication.

Benefits of technology

It reduces the involvement of administrators, improves the efficiency and accuracy of adding new resources to the cluster, and achieves automated configuration and security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113039520B_ABST
    Figure CN113039520B_ABST
Patent Text Reader

Abstract

The present technology addresses the need to automatically configure a new computing resource to join an existing cluster of computing resources. The present technology provides a mechanism to ensure that the new computing resource is executing the same kernel version, which mechanism also allows for the subsequent exchange of at least one configuration message that informs the new computing resource of the necessary configuration parameters and addresses for obtaining the required software packages.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of and priority to U.S. Nonprovisional Patent Application No. 16 / 379,526, entitled “EXTENDING CENTER CLUSTER MEMBERSHIP TO ADDITIONAL COMPUTE RESOURCES,” filed April 9, 2019, and claims the benefit of U.S. Provisional Patent Application No. 62 / 770,143, entitled “EXTENDING CENTER CLUSTER MEMBERSHIP TO ADDITIONAL COMPUTE RESOURCES,” filed November 20, 2018, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The subject matter of the present disclosure relates generally to technology for improving network operations, and more specifically, to improving the addition of compute resources to an existing cluster of compute resources on a network. BACKGROUND

[0004] Campus networks can provide connectivity to computing devices (e.g., servers, workstations, desktop computers, laptops, tablets, mobile phones, etc.) and things (e.g., desktop phones, security cameras, lighting, heating, ventilating, and air-conditioning (HVAC), windows, doors, locks, medical devices, industrial and manufacturing equipment, etc.) located in an environment (e.g., an office, a hospital, a college and university, an oil and gas facility, a factory, and similar locations). Some unique challenges that campus networks can face include: integrating wired and wireless devices; onboarding computing devices and things that can appear anywhere in the network and maintaining connectivity as these devices and things migrate from location to location in the network; supporting bring your own device (BYOD) capabilities; connecting and powering Internet-of-Things (IoT) devices; and securing the network despite vulnerabilities associated with Wi-Fi access, device mobility, BYOD, and IoT. Current methods for deploying networks that are capable of providing these functions often require a number of different systems (e.g., directory-based identity services; authentication, authorization, and accounting (AAA) services, wireless local area network (WLAN) controllers; command-line interfaces for each switch, router, or other network device in the network; etc.) to be operated by highly skilled network engineers and manually pieced together for continuous and extensive configuration and management. This can make network deployment difficult and time consuming and hinder the ability of many organizations to quickly innovate and adopt new technologies (e.g., video, collaboration, and connected workspaces). BRIEF DESCRIPTION OF DRAWINGS

[0005] In order to provide a complete understanding of the present disclosure and its features and advantages, reference is made to the following description, taken in conjunction with the accompanying drawings, in which:

[0006] Figure 1 An example of a physical topology of an enterprise network is shown in accordance with an embodiment;

[0007] Figure 2 An example of a logical architecture for an enterprise network is shown in accordance with an embodiment;

[0008] Figures 3A-3I An example of a graphical user interface for a network management system is shown in accordance with an embodiment;

[0009] Figure 4An example of a physical topology for a multi-site enterprise network is shown in accordance with an embodiment;

[0010] Figure 5 A ladder diagram showing an example method in accordance with some embodiments of the present technology is shown;

[0011] Figure 6 An example of a system in accordance with some embodiments is shown. DETAILED DESCRIPTION

[0012] The detailed description set forth below is intended as a description of various configurations of the subject technology and is not intended to represent the only configurations in which the subject technology can be practiced. The appended drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details for the purpose of providing a thorough understanding of the subject technology. However, it will be clear and apparent that the subject technology is not limited to the specific details set forth herein and can be practiced without these details. In some instances, structures and components are shown in block diagram form in order to avoid obscuring the concepts of the subject technology.

[0013] SUMMARY

[0014] Aspects of the application are set out in the independent claims and preferred features are set out in the dependent claims. Features of one aspect can be applied to each aspect, alone or in combination with features of other aspects.

[0015] The present technology provides for provisioning new resources to a computing cluster with minimal administrator involvement. While there can be various provisioning and orchestration technologies, they are prone to the following problem: new resources can come with configuration software that is not compatible or ideal for membership in the computing cluster, which results in the administrator needing to spend time troubleshooting and then ultimately installing new software and reconfiguring the cluster. The present technology provides a mechanism by which existing cluster members can inform new computing resources of the appropriate configuration and software versions needed to optimally participate in the cluster. Furthermore, existing cluster members can even provide the appropriate software packages (or references to obtain the appropriate software packages) in executable format, so that the new resources can automatically install the software and join the cluster and be fully and correctly configured.

[0016] The present technology allows for a request to join a computing cluster to be sent by a new computing resource to at least one member of the computing cluster. The new computing resource can receive a reply from a member of the computing cluster in one or more communications, the reply including metadata that describes requirements to join the computing cluster and that describes a software bundle used by devices in the computing cluster. The new computing resource can download and install the software bundle and then establish membership in the computing cluster.

[0017] The present technology can include various system components, including a management cloud, existing members of a cluster, new resources to join the cluster, a software package repository, and a credential authority (among other possible components). The existing members of the cluster are configured to receive a request from a new computing resource for information required to join the cluster of computing resources, and in response send information describing a protocol, an identification of a software package, information for obtaining the software package, configuration parameters, and authentication information to the new computing resource.

[0018] The software package repository is configured to receive the request from the new computing resource, and in response provide the software package to the new computing resource. Thereafter, the new computing resource can receive the response from the software package repository and execute the software package, thereby configuring the new computing resource with all the software and configuration necessary to join the cluster of computing resources.

[0019] The credential authority is configured to receive a request to authenticate the new computing resource and accept the new computing resource as a member of the cluster of computing resources, and in response accept the membership of the new computing resource.

[0020] Example Embodiments

[0021] Intent-based networking is a method for overcoming the deficiencies of traditional enterprise networks discussed above and elsewhere in this disclosure. The motivation for intent-based networking is to enable a user to describe in plain language what he or she wants to accomplish (e.g., the user’s intent), and for the network to translate the user’s goal into configuration and policy changes that are automatically propagated across a complex and heterogeneous computing environment. Thus, intent-based networking can abstract network complexity, automate much of the work of provisioning and managing networks that is typically handled by network administrators, and ensure the secure operation and optimal performance of the network. When an intent-based network is aware of users, devices, and things connecting in the network, it can automatically apply security permissions and service levels according to the privileges and quality of experience (QoE) assigned to the users, devices, and things. Table 1 lists examples of intents and workflows that can be automated by an intent-based network to achieve a desired result.

[0022] Table 1: Examples of intents and related workflows

[0023]

[0024] Figure 1An example of a physical topology of an enterprise network 100 for providing intent-based networking is shown. It should be appreciated that there can be additional or fewer nodes, devices, links, networks, or components in similar or alternative configurations for the enterprise network 100 and any network discussed herein. Example embodiments with different numbers and / or types of endpoints, nodes, cloud components, servers, software components, devices, virtual or physical resources, configurations, topologies, services, appliances, or deployments are also contemplated herein. In addition, the enterprise network 100 can include any number or type of resources that can be accessed and utilized by endpoints or network devices. The diagrams and examples provided herein are for clarity and simplicity.

[0025] In this example, the enterprise network 100 includes a management cloud 102 and a network fabric 120. While shown as a network or cloud external to the network fabric 120 in this example, the management cloud 102 can alternatively or additionally be located on-premises or in a colocation center (further, hosted by a cloud provider or similar environment) of the organization. The management cloud 102 can provide a central management plane for building and operating the network fabric 120. The management cloud 102 can be responsible for forwarding configuration and policy distribution as well as device management and analytics. The management cloud 102 can include one or more network controller devices 104, one or more authentication, authorization, and accounting (AAA) devices 106, one or more wireless local area network controllers (WLCs) 108, and one or more fabric control plane nodes 110. In other embodiments, one or more elements of the management cloud 102 can be co-located with the network fabric 120.

[0026] The network controller device(s) 104 can function as a command and control system for one or more network fabrics and can host automated workflows for deploying and managing the network fabric(s). The network controller device(s) 104 can include automation, design, policy, provisioning, and assurance capabilities, among others, as discussed further below with respect to the Cisco® Digital Network Architecture (Cisco DNA®) Center. Figure 2 In some embodiments, one or more Cisco DNA Center® devices can function as the network controller device(s) 104. TM

[0027] ​The AAA device(s) 106 can control access to computing resources, facilitate enforcement of network policies, audit usage, and provide information necessary for service billing. The AAA device(s) can interact with the network controller device(s) 104 and with databases and directories containing information for users, devices, things, policies, billing, and the like to provide authentication, authorization, and accounting services. In some embodiments, the AAA device(s) 106 can utilize Remote Authentication Dial-In User Service (RADIUS) or Diameter to communicate with devices and applications. In some embodiments, one or more An Identity Services Engine (ISE) device can be used as the AAA device(s) 106.

[0028] The WLC(s) 108 can support structure-enabled access points attached to the network fabric 120, handle traditional tasks associated with a WLC, and interact with the fabric control plane for wireless endpoint registration and roaming. In some embodiments, the network fabric 120 can implement a wireless deployment that moves the data plane termination (e.g., VXLAN) from a centralized location (e.g., with a previously overlying Control and Provisioning of Wireless Access Points (CAPWAP) deployment) to the access point / fabric edge node. This can enable distributed forwarding and distributed policy application for wireless traffic while preserving the advantages of centralized provisioning and management. In some embodiments, one or more wireless controllers, wireless LANs and / or other support Cisco DNA TM wireless controllers can be used as the WLC(s) 108.

[0029] The network fabric 120 can include fabric border nodes 122A and 122B (collectively, 122), fabric intermediate nodes 124A-D (collectively, 124), and fabric edge nodes 126A-F (collectively, 126). While in this example the fabric control plane node(s) 110 are shown as being external to the network fabric 120, in other embodiments the fabric control plane node(s) 110 can be co-located with the network fabric 120. In embodiments where the fabric control plane node(s) 110 are co-located with the network fabric 120, the fabric control plane node(s) 110 can comprise a dedicated node or set of nodes, or the functionality of the fabric control plane node(s) 110 can be implemented by the fabric border nodes 122.

[0030] The fabric control plane node(s) 110 can serve as a central database for tracking all users, devices, and things as they attach to the network fabric 120 and as they roam. The fabric control plane node(s) 110 can allow network infrastructure (e.g., switches, routers, WLCs, etc.) to query the database to determine the location of users, devices, and things attached to the fabric, rather than using flooding and learning mechanisms. In this way, the fabric control plane node(s) 110 can serve as a single source of truth as to where each endpoint attached to the network fabric 120 is located at any point in time. In addition to tracking specific endpoints (e.g., / 32 addresses for IPv4, / 128 addresses for IPv6, etc.), the fabric control plane node(s) 110 can also track larger summary routers (e.g., IP / mask). This flexibility can help with summarization across fabric sites and improve overall scalability.

[0031] The fabric border nodes 122 can connect the network fabric 120 to traditional layer 3 networks (e.g., non-fabric networks) or to different fabric sites. The fabric border nodes 122 can also translate context (e.g., user, device, or thing mapping and identity) from one fabric site to another or to a traditional network. When encapsulation is the same across different fabric sites, translation of fabric context is typically a 1 : 1 mapping. The fabric border nodes 122 can also exchange reachability and policy information with the fabric control plane nodes of different fabric sites. The fabric border nodes 122 also provide border functionality for internal and external networks. Internal borders can advertise a defined set of known subnets, such as those leading to a set of branch sites or to a data center. External borders, on the other hand, can advertise unknown destinations (e.g., to the Internet in a similar operation to the function of a default route).

[0032] Fabric intermediate nodes 124 can function as pure Layer 3 forwarders that connect Fabric border nodes 122 to Fabric edge nodes 126 and provide Layer 3 underlay for Fabric overlay traffic.

[0033] Fabric edge nodes 126 can connect endpoints to network fabric 120 and can encapsulate / decapsulate and forward traffic from / to these endpoints to / from network fabric. Fabric edge nodes 126 can operate on the periphery of network fabric 120 and can be the first point of attachment for users, devices, and things as well as implementation of policies. In some embodiments, network fabric 120 can also include Fabric expansion nodes (not shown) for attaching downstream non-fabric Layer 2 network devices to network fabric 120, thereby expanding the network fabric. For example, expansion nodes can be small switches (e.g., compact switches, industrial Ethernet switches, building automation switches, etc.) that are connected to Fabric edge nodes via Layer 2. Devices or things connected to Fabric expansion nodes can use Fabric edge nodes 126 to communicate with external subnets.

[0034] In this example, network fabric can represent a single fabric site deployment, which can be distinguished from a multi-site fabric deployment, as discussed further below with respect to Figure 4

[0035] In some embodiments, all subnets hosted in a fabric site can be provisioned on each Fabric edge node 126 in that fabric site. For example, if subnet 10.10.10.0 / 24 is provisioned in a given fabric site, the subnet can be defined on all Fabric edge nodes 126 in that fabric site, and endpoints located in that subnet can be placed on any Fabric edge node 126 in the fabric. This can simplify IP address management and allow for fewer but larger subnets to be deployed. In some embodiments, one or more Catalyst switches, Cisco Catalyst switches, Cisco MS switches, Integrated Services Router (ISR), Aggregation Services Router (ASR), Enterprise Network Compute System (ENCS), ​Cloud Service Virtual Router (CSRv), Cisco Integrated Services Virtual Router (ISRv), Cisco MX appliances, and / or other Cisco DNA-ready TM Devices can be used as fabric nodes 122, 124, and 126.

[0036] Enterprise network 100 can also include wired endpoints 130A, 130C, 130D, and 130F, and wireless endpoints 130B and 130E (collectively, 130). Wired endpoints 130A, 130C, 130D, and 130F can be connected by wires to fabric edge nodes 126A, 126C, 126D, and 126F, respectively, and wireless endpoints 130B and 130E can be wirelessly connected to wireless access points 128A and 128B (collectively, 128), which in turn are connected by wires to fabric edge nodes 126B and 126E, respectively. In some embodiments, Cisco Access Points, Cisco MR Access Points, and / or other Cisco DNA TM ready access points can be used as wireless access points 128.

[0037] Endpoints 130 can include general-purpose computing devices (e.g., servers, workstations, desktop computers, etc.), mobile computing devices (e.g., laptops, tablets, mobile phones, etc.), wearable devices (e.g., watches, glasses or other head-mounted displays (HMDs), earpieces, etc.), and so on. Endpoints 130 can also include Internet of Things (IoT) devices or appliances, such as agricultural devices (e.g., livestock tracking and management systems, watering devices, unmanned aerial vehicles (UAVs), etc.); connected cars and other vehicles; smart home sensors and devices (e.g., alarm systems, security cameras, lighting, appliances, media players, HVAC devices, electricity meters, windows, automatic doors, doorbells, locks, etc.); office devices (e.g., desktop phones, copiers, fax machines, etc.); healthcare devices (e.g., pacemakers, biometric sensors, medical equipment, etc.); industrial devices (e.g., robots, factory machinery, construction equipment, industrial sensors, etc.); retail devices (e.g., vending machines, point of sale (POS) devices, Radio Frequency Identification (RFID) tags, etc.); smart city devices (e.g., street lights, parking meters, waste management sensors, etc.); transportation and logistics devices (e.g., turnstiles, rental car trackers, navigation devices, inventory monitors, etc.); and so on.

[0038] In some embodiments, network fabric 120 can support both wired and wireless access as part of a single integrated infrastructure, such that connectivity, mobility, and policy enforcement behavior are similar or identical for both wired and wireless endpoints. This can bring a uniform experience independent of access media for users, devices, and things.

[0039] In an integrated wired and wireless deployment, control plane integration can be achieved by the WLC(s) 108 informing the fabric control plane node(s) 110 of the joining, roaming, and disconnection of wireless endpoints 130, so that the fabric control plane node(s) can have connectivity information about both wired and wireless endpoints in the network fabric 120 and can serve as a single source of truth for endpoints connected to the network fabric. For data plane integration, the WLC(s) 108 can instruct the fabric wireless access points 128 to form VXLAN overlay tunnels to their adjacent fabric edge nodes 126. The AP VXLAN tunnels can carry segmentation and policy information to and from the fabric edge nodes 126, allowing the same or similar connectivity and functionality as wired endpoints. When a wireless endpoint 130 joins the network fabric 120 via a fabric wireless access point 128, the WLC(s) 108 can load the endpoint into the network fabric 120 and inform the fabric control plane node(s) 110 of the endpoint’s media access control (MAC) address. The WLC(s) 108 can then instruct the fabric wireless access point 128 to form a VXLAN overlay tunnel to the adjacent fabric edge node 126. Next, the wireless endpoint 130 can obtain its own IP address via dynamic host configuration protocol (DHCP). Once this operation is complete, the fabric edge node 126 can register the IP address of the wireless endpoint 130 with the fabric control plane node(s) 110 to form a mapping between the endpoint’s MAC address and IP address, and traffic to and from the wireless endpoint 130 can begin to flow.

[0040] Figure 2 An example of a logical architecture 200 for an enterprise network (e.g., enterprise network 100) is shown. Those of ordinary skill in the art will appreciate that additional or fewer components can be present in similar or alternative configurations for the logical architecture 200 and any system discussed in this disclosure. The illustrations and examples provided in this disclosure are for brevity and clarity. Other embodiments can include different numbers and / or types of elements, but those of ordinary skill in the art will appreciate that such variations do not depart from the scope of this disclosure. In this example, the logical architecture 200 includes a management layer 202, a controller layer 220, a network layer 230 (e.g., embodied by the network fabric 120), a physical layer 240 (e.g., embodied by the various elements of the network fabric 120), and a shared services layer 250. Figure 1

[0041] ​The management layer 202 can abstract away the complexities and dependencies of the other layers and provide users with tools and workflows for managing enterprise networks (e.g., enterprise network 100). The management layer 202 can include a user interface 204, a design function 206, a policy function 208, a provisioning function 210, an assurance function 212, a platform function 214, and a base automation function 216. The user interface 204 can provide users with a single point for managing and automating networks. The user interface 204 can be implemented within a web application / web server accessible by a web browser and / or an application / application server accessible by a desktop application, a mobile application, a shell program or other command line interface (CLI), an application programming interface (e.g., restful state transfer (REST), Simple Object Access Protocol (SOAP), Service Oriented Architecture (SOA), etc.), and / or another suitable interface in which a user can configure network infrastructure, devices, and things managed by the cloud; provide user preferences; specify policies, input data; view statistics; configure interactions or operations; and so on. The user interface 204 can also provide visibility information, such as views of networks, network infrastructure, computing devices, and things. For example, the user interface 204 can provide views of the status or condition of a network, ongoing operations, services, performance, topology or layout, implemented protocols, running processes, errors, notifications, alerts, network structures, ongoing communications, data analytics, and so on.

[0042] The design function 206 can include tools and workflows for managing site profiles, maps and floor plans, network settings, and IP address management, among others. The policy function 208 can include tools and workflows for defining and managing network policies. The provisioning function 210 can include tools and workflows for deploying networks. The assurance function 212 can provide end-to-end visibility of networks using machine learning and analytics by learning from network infrastructure, endpoints, and other contextual information sources. The platform function 214 can include tools and workflows for integrating network management systems with other technologies. The base automation function 216 can include tools and workflows to support the policy function 208, the provisioning function 210, the assurance function 212, and the platform function 214.

[0043] In some embodiments, the design function 206, the policy function 208, the provisioning function 210, the assurance function 212, the platform function 214, and the base automation function 216 can be implemented as microservices, in which the respective software functions are implemented in multiple containers that communicate with each other, instead of consolidating all tools and workflows into a single software binary. Each of the design function 206, the policy function 208, the provisioning function 210, the assurance function 212, and the platform function 214 can be considered a set of related automation microservices for covering the design, policy making, provisioning, assurance, and cross-platform integration stages of the network lifecycle. The base automation function 214 can support the top-level functions by allowing users to perform certain network-wide tasks.

[0044] Figures 3A-3I An example of a graphical user interface for implementing the user interface 204 is shown. While the graphical user interface is shown as including web pages displayed in a browser executing on a large form factor general purpose computing device (e.g., a server, workstation, desktop, laptop, etc.), the principles disclosed in this disclosure are widely applicable to other form factor client devices, including tablet computers, smartphones, wearable devices, or other small form factor general purpose computing devices; televisions; set-top boxes; Internet of Things devices; and other electronic devices capable of connecting to a network and including input / output components to enable a user to interact with a network management system. Those of ordinary skill in the art will also understand that the graphical user interface shown in FIG. 3 is merely one example of a user interface for managing a network. Other embodiments can include a fewer number or a greater number of elements. Figures 3A-3I Figures 3A-3I

[0045] Figure 3A An example of a graphical user interface 300A, which is an example of a landing screen or home screen of the user interface 204, is shown. The graphical user interface 300A can include user interface elements for selecting the design function 206, the policy function 208, the provisioning function 210, the assurance function 212, and the platform function 214. The graphical user interface 300A also includes user interface elements for selecting the base automation function 216. In this example, the base automation function 216 includes:

[0046] • a network discovery tool 302 for automatically discovering existing network elements to populate into an inventory;

[0047] • an inventory management tool 304 for managing a set of physical and virtual network elements;

[0048] • a topology tool 306 for visualizing the physical topology of network elements;

[0049] ​​• image repository tool 308 to manage software images for network elements;

[0050] • command runner tool 310 to diagnose one or more network elements based on a CLI;

[0051] • license manager tool 312 to manage visual software license usage in a network;

[0052] • template editor tool 314 to create and author CLI templates associated with network elements in a design configuration file;

[0053] • network PnP tool 316 to support automatic configuration of network elements;

[0054] • telemetry tool 318 to design telemetry configuration files and apply telemetry configuration files to network elements; and

[0055] • dataset and reporting tool 320 to access various datasets, schedule data extractions, and generate reports in a variety of formats (e.g., Post Document Format (PDF), comma-separated value (CSV), Tableau, etc.), such as inventory data reports, software image management (SWIM) server reports, and client data reports, among others.

[0056] Figure 3B A graphical user interface 300B is shown, which is an example of a landing screen for the design function 206. The graphical user interface 300B can include user interface elements for various tools and workflows to logically define an enterprise network. In this example, the design tools and workflows include:

[0057] • network layering tool 322 to set up geographic locations, buildings, and floor plans details and associate them with a unique site id;

[0058] • network setup tool 324 to set up network servers (e.g., domain name system (DNS), DHCP, AAA, etc.), device credentials, IP address pools, service provider configuration files (e.g., QoS classes for WAN providers), and wireless settings;

[0059] • image management tool 326 to manage software images and / or maintenance updates, set up version compliance, and download and deploy images;

[0060] • a network profile tool 328 for defining LAN, WAN, and WLAN connection profiles (including Service Set Identifier (SSID)); and

[0061] • an authentication template tool 330 for defining authentication modes (e.g., closed authentication, easy connect, open authentication, etc.).

[0062] The output of the design workflow 206 can include a hierarchical set of unique site identifiers that define global and forwarding configuration parameters for various sites of the network. The provisioning function 210 can use the site identifiers to deploy the network.

[0063] Figure 3C A graphical user interface 300C is shown, which is an example of a landing screen for the policy function 208. The graphical user interface 300C can include various tools and workflows for defining network policies. In this example, the policy design tools and workflows include:

[0064] • a policy dashboard 332 for viewing virtual networks, group-based access control policies, IP-based access control policies, traffic replication policies, scalable groups, and IP network groups. The policy dashboard 332 can also display the number of policies that have failed to deploy. The policy dashboard 332 can provide a list of policies and the following information about each policy: policy name, policy type, policy version (e.g., an iteration of the policy that can be incremented each time the policy is changed), user that modified the policy, description, policy scope (e.g., groups of users and devices or applications that the policy affects), and timestamp;

[0065] • a group-based access control policy tool 334 for managing group-based access control or SGACL. The group-based access control policy can define scalable groups and access contracts (e.g., rules that make up the access control policy, such as to allow or deny when traffic matches the policy);

[0066] • an IP-based access control policy tool 336 for managing IP-based access control policies. The IP-based access control can define IP network groups (e.g., IP subnets that share the same access control requirements) and access contracts;

[0067] • an application policy tool 338 for configuring QoS for application traffic. The application policy can define application sets (e.g., sets of applications that have similar network traffic requirements) and site scopes (e.g., sites for which the application policy is defined);

[0068] • a traffic replication policy tool 340 for setting encapsulated remote switched port analyzer (ERSPAN) configuration such that network traffic between two entities is replicated to a specified destination for monitoring or troubleshooting. The traffic replication policy can define the source and destination of the traffic flow to be replicated, as well as a traffic replication contract that specifies the device and interface to send the traffic copy from; and

[0069] • a virtual network policy tool 343 for segmenting a physical network into multiple logical networks.

[0070] The output of the policy workflow 208 can include a set of virtual networks, security groups, and access and traffic policies that define the policy configuration parameters for various sites of the network. The provisioning function 210 can use the virtual networks, groups, and policies to deploy in the network.

[0071] Figure 3D A graphical user interface 300D is shown, which is an example of a landing screen for the provisioning function 210. The graphical user interface 300D can include various tools and workflows for deploying a network. In this example, the provisioning tools and workflows include:

[0072] • a device provisioning tool 344 for assigning devices to inventory and deploying the required settings and policies, and adding devices to sites; and

[0073] • a structure provisioning tool 346 for creating domains and adding devices to structures.

[0074] The output of the provisioning workflow 210 can include the deployment of the network underlay and structure overlay, and the policies (defined in the policy workflow 208).

[0075] Figure 3E A graphical user interface 300E is shown, which is an example of a landing screen for the assurance function 212. The graphical user interface 300E can include various tools and workflows for managing a network. In this example, the assurance tools and workflows include:

[0076] • a health overview tool 344 for providing a global view of the enterprise network, including network infrastructure devices and endpoints. User interface elements (e.g., drop-down menus, dialog boxes, etc.) associated with the health overview tool 344 can also be converted to toggle to additional or alternative views, such as a view of the health of only network infrastructure devices, a view of the health of all wired and wireless clients, and a view of the health of applications running in the network, as described below with respect to Figures 3F-3Hfurther discussed;

[0077] • Assurance dashboard tool 346 for managing and creating custom dashboards;

[0078] • Issue tool 348 for displaying and resolving network issues; and

[0079] • Sensor management tool 350 for managing sensor-driven tests.

[0080] Graphical user interface 300E can also include a location selection user interface element 352, a time period selection user interface element 354, and a view type user interface element 356. Location selection user interface element 354 can enable a user to view the overall health of a particular site (e.g., as defined via the network hierarchy tool 322) and / or network domain (e.g., LAN, WLAN, WAN, data center, etc.). Time period selection user interface element 356 can enable the display of the overall health of the network over a particular time period (e.g., last 3 hours, last 24 hours, last 7 days, custom time period, etc.). View type user interface element 355 can enable a user to toggle between a geographical map view (not shown) of the sites of the network or a hierarchical site / building view (as shown).

[0081] Within the hierarchical site / building view, rows can represent network hierarchies (e.g., sites and buildings defined by the network hierarchy tool 322); column 358 can indicate the number of healthy clients in percentages; column 360 can indicate the health of wireless clients by score (e.g., 1-10), color, and / or descriptor (e.g., red or critical, associated with a health score of 1-3 and indicating that the client has a critical issue; orange or warning, associated with a health score of 4-7 and indicating a warning for the client; green or no errors or warnings, associated with a health score of 8-10; gray or no data available, associated with a health score of null or 0), or other indicator; column 362 can indicate the health of wired clients by score, color, descriptor, etc.; column 364 can include user interface elements for drilling down into the health of clients associated with the hierarchical site / building; column 366 can indicate the number of healthy network infrastructure devices in percentages; column 368 can indicate the health of access switches by score, color, descriptor, etc.; column 370 can indicate the health of core switches by score, color, descriptor, etc.; column 372 can indicate the health of distribution switches by score, color, descriptor, etc.; column 374 can indicate the health of routers by score, color, descriptor, etc.; column 376 can indicate the health of WLCs by score, color, descriptor, etc.; column 378 can indicate the health of other network infrastructure devices by score, color, descriptor, etc.; and column 380 can include user interface elements for drilling down into the health of network infrastructure devices associated with the hierarchical site / building. In other embodiments, client devices can be grouped in other ways besides wired or wireless, such as by device type (e.g., desktop, laptop, mobile phone, IoT device, or more specific type of IoT device, etc.), manufacturer, model, operating system, etc. Likewise, in additional embodiments, network infrastructure devices can also be grouped in these ways, as well as others.

[0082] The graphical user interface 300E can also include an overall health summary user interface element (e.g., view, pane, tile, card, container, widget, dashlet, etc.) that includes: a client health summary user interface element 384 indicating the number of healthy clients in percentages; a color-coded trend chart 386 indicating the percentages over a particular time period (e.g., as selected by the time period selection user interface element 354); a user interface element 388 breaking down the number of healthy clients by client type (e.g., wireless, wired) in percentages; a network infrastructure health summary user interface element 390 indicating the number of healthy network infrastructure devices in percentages; a color-coded trend chart 392 indicating the percentages over a particular time period; and a user interface element 394 breaking down the number of network infrastructure devices by network infrastructure device type (e.g., core switch, access switch, distribution switch, etc.) in percentages.

[0083] The graphical user interface 300E can also include a problems user interface element 396 that lists the problems that must be addressed, if any. The problems can be ordered based on timestamp, severity, location, device type, etc. Each problem can be selected for in-depth exploration to view a more detailed view of the selected problem.

[0084] Figure 3F A graphical user interface 300F is shown that is an example of a screen for summarizing the health of network infrastructure devices only, which can be navigated to, for example, by toggling the health summary tool 344. The graphical user interface 300F can include a timeline slider 398 for selecting a more granular time range than the time period selection user interface element (e.g., the time period selection user interface element 354). The graphical user interface 300F can also include similar information as shown in the graphical user interface 300E, such as the following: a user interface element including a hierarchical site / building view and / or a geographic map view similar to the graphical user interface 300E (except providing information for network infrastructure devices only) (not shown here); the number of healthy network infrastructure devices in percentages 390; a color-coded trend chart indicating percentages by device type 392; the number of healthy network infrastructure devices broken down by device type 394; and so on. In addition, the graphical user interface 300F can display a view of the health of network infrastructure devices by network topology (not shown). The view can be interactive, such as by enabling the user to zoom in or out, pan left or right, or rotate the topology (e.g., 90 degrees).

[0085] In this example, the graphical user interface 300F also includes: a color-coded trend chart 3002 showing performance of network infrastructure devices over a particular time period; a network health by device type tab that includes a system health chart 3004 (which provides system monitoring metrics (e.g., CPU utilization, memory utilization, temperature, etc.)), a data plane connectivity chart 3006 (which provides data plane metrics (e.g., uplink availability and link errors)), and a control plane connectivity chart 3008 (which provides control plane metrics for each device type); AP analysis user interface elements including an up and down color-coded chart 3010 (which provides AP status information (e.g., the number of APs connected to the network, and the number of APs not connected to the network, etc.)), and a chart of top N APs sorted by client count 3012 (which provides information about APs with the highest number of clients); a network device table 3014 that enables a user to filter (e.g., by device type, health, or custom filters), view, and export network device information. A detailed view of the health of each network infrastructure device can also be provided by selecting the network infrastructure device in the network device table 3014.

[0086] Figure 3G A graphical user interface 300G is shown, which is an example of a screen for summarizing health of client devices, which can be navigated to, for example, by toggling the health summary tool 344. The graphical user interface 300G can include: an SSID user interface selection element 3016 for viewing health of wireless clients by all SSIDs or a particular SSID; a band user interface selection element 3018 for viewing health of wireless clients by all bands or a particular band (e.g., 2.4 GHz, 5 GHz, etc.); and a time slider 3020, which can operate similarly to the time slider 398.

[0087] The graphical user interface 300G can also include a client health summary user interface element that provides similar information as shown in the graphical user interface 300E, e.g., the number of healthy clients 384 expressed in percentage, and a color-coded trend chart 386 that indicates the percentage over a particular time period for each group of client devices (e.g., wired / wireless, device type, manufacturer, model, operating system, etc.). In addition, the client health summary user interface element can include a color-coded pie chart that provides counts of client devices that are poor (e.g., red and indicates a client health score of 1-3), average (e.g., orange and indicates a client health score of 4-7), good (e.g., green and indicates a health score of 8-10), and inactive (e.g., gray and indicates a health score of null or 0). The counts of client devices associated with each color, health score, health descriptor, etc. can be displayed by a selection gesture (e.g., tap, double tap, long press, hover, click, right click, etc.) directed to that color.

[0088] The graphical user interface 300G can also include a number of other client health metrics charts for all sites or selected sites over a particular time period, e.g.:

[0089] • client load time 3024;

[0090] • received signal strength indication (RSSI) 3026;

[0091] • connectivity signal-to-noise ratio (SNR) 3028;

[0092] • client count per SSID 3030;

[0093] • client count per band 3032;

[0094] • DNS request and response counters (not shown); and

[0095] • connectivity physical link state information 3034 indicating the distribution of wired client devices that are up, down, and have errors on their physical link.

[0096] In addition, the graphical user interface 300G can include a client device table 3036 that enables a user to filter (e.g., by device type, health, data (e.g., load time > threshold, association time > threshold, DHCP > threshold, AAA > threshold, RSSI > threshold, etc.), or custom filters), view, and export client device information (e.g., user identifier, hostname, MAC address, IP address, device type, last heard information, location, VLAN identifier, SSID, overall health score, load score, connection score, network infrastructure device to which the client device is connected, etc.). A detailed view of the health of each client device can also be provided by selecting the client device in the client device table 3036.

[0097] Figure 3H A graphical user interface 300H is shown, which is an example of a screen for summarizing the health of applications, which can be navigated to, for example, by toggling the health summary tool 344. The graphical user interface 300H can include an application health summary user interface element that includes a percentage of the number of healthy applications expressed as a percentage 3038, a health score for each application or each type of application (e.g., transaction related, non-transaction related, default; HTTP, VoIP, chat, email, bulk transfer, multimedia / streaming, etc.) running in the network 3040, a chart of the top N applications ordered by usage 3042. The health score 3040 can be calculated based on quality indicators for the application (e.g., packet loss, network latency, etc.).

[0098] In addition, the graphical user interface 300H can also include an application table 3044 that enables a user to filter (e.g., by application name, domain name, health, usage, average throughput, traffic class, packet loss, network latency, application latency, custom filters, etc.), view, and export application information. A detailed view of the health of each application can also be provided by selecting the application in the application table 3044.

[0099] Figure 3I An example of a graphical user interface 300I is shown, which is an example of a landing screen for the platform functionality 210. The graphical user interface 300I can include various tools and workflows for integration with other technical systems. In this example, the platform integration tools and workflows include:

[0100] • a bundle tool 3046 for managing the packaging of domain-specific APIs, workflows, and other features for network programming and platform integration;

[0101] • Developer Toolbox 3048 for accessing an API catalog that lists available APIs and methods (e.g., GET, PUT, POST, DELETE, etc.), descriptions, runtime parameters, return codes, model schemas, etc. In some embodiments, the Developer Toolbox 3048 can also include a "Try It" button for allowing developers to experiment with a particular API to better understand its behavior;

[0102] • Runtime Dashboard 3050 for viewing and analyzing underlying metrics or API and integration flow usage;

[0103] • Platform Settings Tool 3052 for viewing and setting global or bundle-specific settings that define integration target and event usage preferences; and

[0104] • Notification user interface elements 3054 for presenting notifications related to availability of software updates, security threats, etc.

[0105] Returning to Figure 2 , the controller layer 220 can include subsystems for the management layer 202 and can include a network control platform 222, a network data platform 224, and AAA services 226. These controller subsystems can form an abstraction layer for hiding the complexity and dependencies of managing many network elements and protocols.

[0106] The network control platform 222 can provide automation and orchestration services for the network layer 230 and the physical layer 240, and can include settings, protocols, and tables for automating the management of the network layer and the physical layer. For example, the network control platform 222 can provide the design function 206 and the provisioning function 210. In addition, the network control platform 222 can include tools and workflows for discovering switches, routers, wireless controllers, and other network infrastructure devices (e.g., the network discovery tool 302); maintaining network and endpoint details, configurations, and software versions (e.g., the inventory management tool 304); Plug-and-Play (PnP) for automatically deploying network infrastructure (e.g., the network PnP tool 316); path tracing for creating visual data paths to speed up troubleshooting connectivity issues; simple QoS for automating quality of service to prioritize applications on the network; and Enterprise Service Automation (ESA) for automatically deploying physical and virtual network services, among others. The network control platform 222 can communicate with network elements using Network Configuration (NETCONF) / Yet Another Next Generation (YANG), Simple Network Management Protocol (SNMP), Secure Shell (SSH) / Telnet, and the like. In some embodiments, The network control platform (NCP) can be used as the network control platform 222.

[0107] The network data platform 224 can provide network data collection, analysis, and assurance, and can include settings, protocols, and tables for monitoring and analyzing network infrastructure and endpoints connected to the network. The network data platform 224 can collect multiple types of information from network infrastructure devices, including Syslog, SNMP, NetFlow, Switched Port Analyzer (SPAN), and streaming telemetry, among others.

[0108] In some embodiments, one or more Cisco DNA TM The central appliance can provide the functionality of the management layer 202, the network control platform 222, and the network data platform 224. The Cisco DNA TM The central appliance can support horizontal scalability by adding additional Cisco DNA TMCentral node; high availability for both hardware components and software packages; backup and storage mechanisms to support disaster recovery scenarios; role-based access control mechanisms to distinguish access for users, devices, and things based on roles and scopes; and programmable interfaces to enable integration with third-party vendors. Cisco DNA TM The central device can also be cloud-bound to provide upgrades to existing functionality and additions of new packages and applications without the need to manually download and install them.

[0109] The AAA service 226 can provide identity and policy services for the network layer 230 and the physical layer 240, and can include settings, protocols, and tables to support endpoint identification and policy enforcement services. The AAA service 226 can provide tools and workflows to manage virtual networks and security groups, and to create group-based policies and contracts. The AAA service 226 can use AAA / RADIUS, 802. IX, MAC authentication bypass (MAB), web authentication, and EasyConnect, among others, to identify and profile network infrastructure devices and endpoints. The AAA service 226 can also collect and use contextual information from the network control platform 222, the network data platform 224, and the shared services 250, among others. In some embodiments, The ISE can provide the AAA service 226.

[0110] The network layer 230 can be conceptualized as a combination of two layers (a bottom layer 234 and an overlying layer 232), the bottom layer 234 including physical and virtual network infrastructure (e.g., routers, switches, WLCs, etc.) and layer 3 routing protocols for forwarding traffic, and the overlying layer 232 including virtual topologies for logically connecting wired and wireless users, devices, and things and applying services and policies to these entities. The network elements of the bottom layer 234 can establish connectivity between each other, for example, via Internet Protocol (IP). The bottom layer can use any topology and routing protocols.

[0111] In some embodiments, the network controller 104 can provide local area network (LAN) automation services (e.g., by Cisco DNA TMThe central LAN automation implementation) for automatically discovering, provisioning, and deploying network devices. Once discovered, the automation underlay provisioning service can utilize plug-and-play (PnP) to apply the required protocols and network address configurations to the physical network infrastructure. In some embodiments, the LAN automation service can implement the Intermediate System to Intermediate System (IS-IS) protocol. Some advantages of IS-IS include: neighbor establishment without IP protocol dependency; peering capability using loopback addresses; and agnostic handling of IPv4, IPv6, and non-IP traffic.

[0112] The overlay 232 can be a logical virtualized topology built on top of the physical underlay 234 and can include a fabric data plane, a fabric control plane, and a fabric policy plane. In some embodiments, the fabric data plane can be created using a Virtual Extensible LAN (VXLAN) with Group Policy Option (GPO) via packet encapsulation. Some advantages of VXLAN-GPO include: its support of both layer 2 and layer 3 virtual topologies (overlays); and its ability to run on any IP network with built-in network segmentation.

[0113] In some embodiments, the fabric control plane can implement Locator / ID Separation Protocol (LISP) for logically mapping and resolving users, devices, and things. LISP can simplify routing by eliminating the need for each router to process every possible IP destination address and route. LISP can accomplish this by moving remote destinations to a centralized map database that allows each router to only manage its local routes and query the map system to locate destination endpoints.

[0114] The fabric policy plane is where intents can be translated into network policies. That is, the policy plane is where network operators can instantiate logical network policies based on the services offered by the network fabric 120 (e.g., security segmentation services, Quality of Service (QoS), capture / replication services, application visibility services, etc.).

[0115] Segmentation is a method or technique used to separate a particular group of users or devices from other groups with the goal of reducing congestion, improving security, containing network problems, controlling access, and so on. As discussed, the fabric data plane can provide network segmentation by implementing VXLAN encapsulation using the virtual network identifier (VNI) and Scalable Group Tag (SGT) fields in the packet header. The network fabric 120 can support both macro-segmentation and micro-segmentation. Macro-segmentation logically divides the network topology into smaller virtual networks by using unique network identifiers and separate forwarding tables. This can be instantiated as a virtual routing and forwarding (VRF) instance and is referred to as a virtual network (VN). That is, a VN is a logical network instance within the network fabric 120 defined by a Layer 3 routing domain and can provide both Layer 2 and Layer 3 services (using the VXLAN VNI to provide both Layer 2 and Layer 3 segmentation). Micro-segmentation logically separates groups of users or devices in a VN by enforcing source-to-destination access control rights (e.g., by using access control lists (ACLs)). A scalable group is a logical object identifier assigned to a group of users, devices, or things in the network fabric 120. It can be used as a source and destination classifier in a scalable group ACL (SGACL). The SGT can be used to provide group-based policies that are independent of addresses.

[0116] In some embodiments, the fabric control plane nodes 110 can implement Locator / Identifier Separation Protocol (LISP) to communicate with each other and with the management cloud 102. Thus, the control plane nodes can operate a host tracking database, a map server, and a map resolver. The host tracking database can track endpoints 130 connected to the network fabric 120 and associate the endpoints with fabric edge nodes 126, decoupling the endpoints' identifiers (e.g., IP or MAC addresses) from their location in the network (e.g., the nearest router).

[0117] The physical layer 240 can include network infrastructure devices, such as switches and routers 110, 122, 124, and 126, and wireless elements 108 and 128, as well as network devices, such as network controller device(s) 104 and AAA device(s) 106.

[0118] The shared services layer 250 can provide interfaces to the following external network services: for example, cloud services 252; domain name system (DNS), DHCP, IP address management (IPAM), and other network address management services 254; firewall services 256; Network as a Sensor (Naas) / Encrypted Threat Analytic (ETA) services; and Virtual Network Function (VNF) 260; and so forth. The management layer 202 and / or the controller layer 220 can share identities, policies, forwarding information, and so forth using APIs via the shared services layer 250.

[0119] Figure 4 An example of a physical topology for a multi-site enterprise network 400 is shown. In this example, the network fabric includes fabric sites 420A and 420B. Fabric site 420A can include fabric control node 410A, fabric border nodes 422A and 422B, fabric intermediate nodes 424A and 424B (shown here in dashed lines and not connected to fabric border or fabric edge nodes for simplicity), and fabric edge nodes 426A-C. Fabric site 420B can include fabric control node 410B, fabric border nodes 422C-E, fabric intermediate nodes 424C and 424D, and fabric edge nodes 426D-F. Multiple fabric sites corresponding to a single fabric (e.g., fabric 400) can be interconnected by a transit network. The transit network can be part of the network fabric with its own control plane nodes and border nodes but no edge nodes. In addition, the transit network shares at least one border node with each fabric site it interconnects. Figure 4

[0120] Generally, the transit network connects the network fabric to the outside world. There are several methods for external connectivity, such as a traditional IP network 436, a traditional WAN 438A, a software-defined WAN (SD-WAN) (not shown), or a software-defined access (SD-Access) 438B. Traffic across fabric sites and to other types of sites can use the control plane and data plane of the transit network to provide connectivity between these sites. Local border nodes can be used as a handoff point from the fabric site, and the transit network can deliver traffic to other sites. The transit network can use other features. For example, if the transit network is a WAN, features such as performance routing can also be used. To provide end-to-end policies and segmentation, the transit network should be able to carry endpoint context information (e.g., VRF, SGT) across the network. Otherwise, traffic can need to be reclassified at the destination site border. ​

[0121] The local control plane in a fabric site can only hold state related to endpoints connected to edge nodes within the local fabric site. For a single fabric site (e.g., network fabric 120), the local control plane can register local endpoints via the local edge nodes. Endpoints that are not explicitly registered with the local control plane can be assumed to be reachable via a border node connected to a transit network. In some embodiments, the local control plane can not hold state for endpoints attached to other fabric sites, such that border nodes do not register information from the transit network. In this way, the local control plane can be independent of other fabric sites, thus enhancing the overall scalability of the network.

[0122] The control plane in a transit network can hold a summarized state of all fabric sites it interconnects. This information can be registered to the transit control plane by border nodes from different fabric sites. Border nodes can register EID information from local fabric sites to the transit network control plane for summarized EIDs only, and thus further improve scalability.

[0123] The multi-site enterprise network 400 can also include a shared services cloud 432. The shared services cloud 432 can include one or more network controller devices 404, one or more AAA devices 406, and other shared servers (e.g., DNS; DHCP; IPAM; SNMP and other monitoring tools; NetFlow, Syslog, and other data collectors, etc.). These shared services can typically reside outside of the network fabric, and in the global routing table (GRT) of the existing network. In this case, some method of inter-VRF routing can be needed. One option for inter-VRF routing is to use a converged router, which can be an external router that performs inter-VRF leaking (e.g., import / export of VRF routes) to converge VRFs together. Multiprotocol can be used for this route exchange, as it can inherently prevent routing loops (e.g., using the AS_PATH attribute). Other routing protocols can also be used, but can require complex distribution lists and prefix lists to prevent loops.

[0124] However, using a converged router to implement inter-VN communication can have some drawbacks, such as: route replication, because routes that leak from one VRF to another VRF are programmed in the hardware tables and can cause higher TCAM utilization; manual configuration at multiple touch points to implement route leaking; loss of SGT context, because SGTs can not be maintained across VRFs and SGTs must be reclassified once traffic enters another VRF; and traffic hairpinning, because traffic can need to be routed to the converged router and then back to the fabric border node.

[0125] SD-Access Extranet can provide a flexible and scalable approach to implementing inter-VN communication by avoiding route replication, because inter-VN lookups are performed in the fabric control plane (e.g., software) so that there is no need to replicate route entries in hardware; providing a single touch point, because the network management system (e.g., Cisco DNA Center) can automate inter-VN lookup policies, making it a single point of management; maintaining SGT context, because inter-VN lookups are performed in the control plane node(s) (e.g., software); and avoiding hairpinning, because inter-VN forwarding can be performed at the fabric edge (e.g., within the same VN), so traffic does not need to be hairpinned at the border node. Another advantage is that separate VNs can be created for each common resource that is needed (e.g., shared services VN, Internet VN, data center VN, etc.). TM Centralized cluster membership extension to additional compute resources

[0126] Centralized cluster membership extension to additional compute resources

[0127] The system described above in Figures 1-4 System for managing enterprise networks The system described above in

[0128] Entities in a cluster use pre-agreed protocols to coordinate with each other. However, an entity seeking to become a new member may have a different version of software or protocol, which may be incompatible with the software or protocol currently used in the cluster. When this occurs, the new member will be unable to join the cluster and gain membership. Such issues often require software or configuration modifications before joining the cluster. Currently, this process may require administrator intervention and delays cluster formation.

[0129] In some embodiments, this technology can communicate appropriate software and configuration requirements to entities attempting to become new members of a cluster (or update existing members within a cluster) by transmitting a manifest file containing necessary information about the appropriate software required to join the cluster, where to download the software, and the appropriate configuration.

[0130] Figure 5 An example ladder diagram is shown, illustrating sample communication and methods for adding computing resources to an existing cluster of computing resources. The computing resources can be any physical or virtual resource and can provide any functionality, including networking, indexing, and storage, or computational tasks.

[0131] This technology can solve the above problems in the following way: by using the configuration function 210 of the management cloud 102 to configure multiple levels of the network ( Figure 2 As shown in the figure, it can automatically configure and accept new computing resources into the cluster.

[0132] exist Figure 5 Prior to the first communication in the method shown, a first resource needs to be configured. Configuration function 210 can configure the first resource to provide functionality, and thereafter the first resource can be used (e.g., Figure 5 The configuration (shown) adds the first resource to each additional resource in the cluster to provide this functionality. The first resource can be configured using configuration feature 210 (or other features of management cloud 102) under the guidance of a network administrator. Configuring the first resource may include defining appropriate software packages, including runtime environment, version, API, network configuration, etc. Once the resource is configured, a bundle (which includes a Docker image (or another type of container) containing the software package) will be generated and stored, and an inventory will be created (which includes information about the software release, configuration settings, and other necessary data for the appropriate configuration). In some embodiments, the bundle may be stored directly on the first resource, or the bundle may be stored or managed by image repository 308.

[0133] Once the primary resource has been configured, it can be used by newly added resources, such as... Figure 5are shown to ensure that each member of the cluster is configured in the same way.

[0134] Figures 1-5 A method 500 is shown by which a new computing resource 502 can join the cluster by communicating with existing cluster members 501 (previously configured resources) of the cluster as well as other management cloud 102 resources.

[0135] The method 500 can begin as follows, the new computing resource 502 is initialized 510 and executes firmware that effectively runs a preboot execution environment protocol in which the new computing resource 502 sends a preboot execution environment (PXE) request 511 to a preboot execution environment server that receives the request. In response, the preboot execution environment server sends a boot image (that includes a software kernel that defines an operating environment used by the existing cluster member(s) 501) to the new resource 502 joining the cluster to initialize the new resource 502 joining the cluster.

[0136] The new computing resource 502 can receive the boot image and initialize 513 using the boot image. These steps ensure that the new computing resource 502 is running an appropriate kernel.

[0137] In some embodiments, steps 511, 512, and 513 are optional, as indicated by the dashed lines. In some embodiments, the new computing resource 502 is already running an appropriate kernel, and steps 511, 512, and 513 can be skipped. In some embodiments, the new resource 502 joining the cluster simply does not perform steps 511, 512, and 513. In some embodiments, the preboot execution environment boot fails, and therefore steps 511, 512, and 513 are not performed or have the same effect as if they were not performed. These steps are optional because all resources will already include a Linux or Windows kernel and any software necessary to use the manifest, as described below. The Linux kernel is itself capable of executing Docker containers, and a resource running a Windows kernel under the control of the management cloud 102 includes additional software necessary to execute Docker containers. As will be described below, the software in the Docker container is able to otherwise update the new resource 502 to be able to join the cluster.

[0138] Whether or not the new computing resource 502 has performed the pre-boot execution environment boot 513, the new computing resource 502 can request 517 a manifest from an existing cluster member 501. To know where to send the request 517, the new computing resource 502 can receive a communication from the management cloud 102 that directs the new computing resource 502 to join an existing cluster. The communication from the management cloud 102 can provide instructions to make the manifest request 517 and direct the request to one or more existing cluster members 501. In some embodiments, the new resource 502 can prompt a user to provide an IP address of a member of an existing cluster to which the new resource 502 is to join through the management cloud 102.

[0139] The existing member 501 of the cluster can receive the manifest request 517 and either re-direct the request to another existing member 501 of the cluster that has been designated to handle such requests or it can respond itself. In response to receiving the request, the existing cluster member 501 can send 518 the requested manifest, which includes information describing the protocol, APIs, versions, identification of software package(s), information (pointers) for obtaining the software package, configuration parameters, file format information, path information, data schema, etc. In addition to the information in the manifest, the existing member 501 of the cluster can also send a seed package (e.g., public and private keys) for a credential authority that will be used as a basis for mutual trust between the new computing resource 502 and the existing cluster member(s) 501 in the cluster to proceed forward.

[0140] In some embodiments, the manifest request 517 can include information about the existing software environment and configuration that the new resource 502 is currently running and the existing cluster member 501 can determine that the new computing resource is not running the same version of the software bundle as the existing cluster member 501. In some embodiments, the existing cluster member 501 can determine the differences in the existing software environment and configuration of the new member 502 as compared to the existing cluster member 501. In such embodiments, the existing cluster member 501 can prepare a manifest that identifies the differences in configuration.

[0141] In some embodiments, the manifest request can include instructions for the joining new entity 502 to request and receive a Docker image from a docket daemon. In some embodiments, the existing cluster member 501 can send a Docker image that contains the software package(s) required to participate in the cluster.

[0142] Upon receiving the manifest, the new computing resource 502 can process 519 the manifest data and install authentication information, such as software required for a distributed credential authority. In some embodiments, the authentication services such as a distributed credential authority can be provided by the AAA service 226 that manages the cloud 102.

[0143] As directed by the manifest, the new computing resource 502 can then request 522 a software package, and the request can be received 522 by a software package repository. In response, the software package repository can provide the software package to the new computing resource 502, which can run the software package 523. Examples of software packages include, but are not limited to, executable Docker containers, Java executables, and the like.

[0144] In some embodiments, the software package is a Docker image. The Docker image contains the following software to install: kernel loadable modules, Deb, apk, pip, or whl packages, golang dependent packages, java software (including jars), new libraries / binaries, new ansible orchestration playbooks, new configuration files, and the like. When this Docker image is run to completion, it ensures that the new cluster members are running all the same software versions, and that they are configured in a compatible manner.

[0145] To make such Docker images available on the cluster, whenever a cluster software version is updated, a bundle needs to be generated as part of the upgrade package, containing all the software that was updated into the Docker image and tagged with the release version. Whenever a software update for the cluster is initiated, the cluster members will download the image. In this way, if the cluster is constantly moving from one software version to another, the new members will receive directly whatever the latest version the cluster is running, and later will be upgraded together with the rest of the members of the cluster. In this way, it is ensured that all members of the cluster are running the same version of the cluster software.

[0146] Now that the new computing resource 502 has all the software and configuration necessary to be part of the cluster, the new computing resource 502 can attempt to join the cluster 526 by requesting its membership be authenticated 527 by the credential authority. The credential authority can receive the request 527 to authenticate the new computing resource and accept the new computing resource as a member of the cluster of computing resources, and in response 528, the credential authority can accept the membership of the new computing resource. In some embodiments, the credential authority can be distributed.

[0147] While the systems shown herein (e.g., the cloud 102) have been described as including a single AAA service 226, in some embodiments, the AAA service 226 can be distributed. Figure 6The present technology is primarily discussed in the context of a system designed to manage an enterprise network (shown in FIG. 1), but the present technology is applicable to any system where it is desirable to expand cluster membership. The present technology can be expected in any system where a current cluster member or orchestration service can send a manifest that includes all information necessary for a new resource to automatically download and install the necessary software versions, set up the appropriate configuration, and learn the proper credentials to join the cluster.

[0148] ​ An example of a computing system 600 is shown, which can be, for example, any computing device that makes up the network controller device 104, the authentication, authorization, and accounting (AAA) device 106, the wireless local area network controller (WLC) 108, the fabric control plane node 110, the border node 122, the intermediate node 124, the edge node 126, the access point 128, the endpoint 130, or the management cloud 102 or any component thereof, where the components of the system communicate with each other using connections 605. The connections 605 can be physical connections via a bus, or direct connections into the processor 610 (e.g., in a chipset architecture). The connections 605 can also be virtual connections, networking connections, or logical connections.

[0149] In some embodiments, the computing system 600 is a distributed system in which the functionality described in this disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent many such components, each component performing some or all of the functions attributed to it. In some embodiments, the components can be physical or virtual devices.

[0150] The example system 600 includes at least one processing unit (CPU or processor) 610 and connections 605 that couple various system components including the system memory 615 (e.g., read-only memory (ROM) 620 and random access memory (RAM) 625) to the processor 610. The computing system 600 can include a high-speed memory cache 612 that is directly connected to the processor 610, is in close proximity to the processor 610, or is integrated as part of the processor 610.

[0151] The processor 610 can include any general purpose processor and a hardware service or software service (e.g., services 632, 634, and 636 stored in storage device 630) configured to control the processor 610 and the processor 610 can include a special purpose processor in which software instructions are incorporated into the actual processor design. The processor 610 can essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, and cache, etc. Multi-core processors can be symmetric or asymmetric.

[0152] To enable user interaction, the computing system 600 includes an input device 645, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and the like. The computing system 600 can also include output device(s) 635, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multi modal systems can enable a user to provide multiple types of input to communicate with the computing system 600. The computing system 600 can include communication interface 640, which can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here can easily be substituted for improved hardware or firmware arrangements as they are developed.

[0153] Storage device 630 can be a non-transitory memory device and can be a hard disk or other types of computer readable media which can store data that is accessible by a computer, such as a magnetic cassette, flash memory cards, solid-state memory devices, digital versatile disks, magnetic cassettes, random access memories (RAMs), read only memory (ROM), and / or some combination of these, amongst others.

[0154] The storage device 630 can include software services, servers, services, and the like, that, when code defining such software is executed by the processor 610, cause the system to perform a function. In some embodiments, a hardware service that performs a particular function can include software components stored in a computer-readable medium which are connected to the necessary hardware components (e.g., processor 610, connections 605, output device 635, etc.) to carry out the function.

[0155] In summary, the present technology addresses the need to automatically configure a new computing resource to join an existing cluster of computing resources. The present technology provides a mechanism to ensure that the new computing resource is executing the same kernel version, which mechanism also allows for the subsequent exchange of at least one configuration message that informs the new computing resource of the necessary configuration parameters and an address for obtaining the required software packages.

[0156] For the sake of clarity, in certain situations the technology can be described as including separate functional blocks that include functions which can be implemented in software by a device, by a combination of software and hardware, or entirely by hardware.

[0157] Any of the steps, operations, functions, or processes described herein can be performed or implemented by a hardware service and a software service, or a combination of services, alone or in combination with other devices. In some embodiments, a service can be software that resides in memory of a client device and / or memory of one or more servers of a content management system, and the service can perform one or more functions when a processor executes software associated with the service. In some embodiments, a service is a program or set of programs that performs a particular function. In some embodiments, a service can be considered a server. The memory can be a non-transitory computer-readable medium.

[0158] In some embodiments, computer-readable storage devices, media, and memories can include cables or wireless signals that contain a bit stream, etc. However, when referred to, non-transitory computer-readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals per se.

[0159] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available at a computer-readable medium. Such instructions can include, for example, instructions and data used to program general purpose computer, special purpose computer, or special purpose processing devices to perform particular functions or groups of functions. Portions of computer resources used can be accessible via a network. The computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, firmware, or microcode, or even source code. Examples of computer-readable media that can be used to store instructions, information used by the instructions, and / or information created during execution of the instructions include magnetic or optical disks, solid state memory devices, flash memory, USB devices provided with non-volatile memory, networked storage devices, and the like.

[0160] Devices that implement methods according to these disclosures can include hardware, firmware, and / or software, and can take any of a variety of form factors. Typical examples of such form factors include servers, laptops, smart phones, small form factor personal computers, personal digital assistants, and other devices. Functionality described herein can also be embodied in peripherals or add-in cards that can be inserted into device. As further examples, such functionality can also be implemented by a circuit on a motherboard or by a combination of circuits and / or program instructions operating together to cause a device to perform a described function.

[0161] Instructions, media for conveying such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.

[0162] While aspects have been illustrated and described in connection with various examples and other information, it will be understood by those of ordinary skill in the art that other variations can be made of the aspects within the scope of the appended claims without departing from the scope of the claims. Additionally, while aspects can have been described in connection with particular examples thereof, it will be understood that the subject matter defined by the appended claims is not necessarily limited to those examples, but can include other examples that fall within the scope of the claims. For example, although specific features and / or steps were described in relation to specific examples, an aspect can include one, some, or all of those features and / or steps. For example, features described in relation to one aspect can be combined with features described in relation to another aspect. Furthermore, although features can be described as being part of an example, an aspect can include one, some, or all of those features. For example, an aspect can include one, some, or all of the features described in relation to one example.

Claims

1. A computer-implemented method for joining a new computing resource to a computing cluster, comprising: sending, by a new computing resource, a request to an existing member of a computing cluster comprising at least one member for metadata describing requirements for joining the computing cluster; receiving, by the new computing resource, a reply to the request for metadata in one or more communications, the reply comprising: the metadata describing requirements for joining the computing cluster, and a reference to a software bundle used by devices in the computing cluster; sending, by the new computing resource, a request for the software bundle to a software package repository based on the reply; receiving, by the new computing resource, the software bundle from the software package repository as a response to the request for the software bundle; installing, by the new computing resource, the software bundle to configure the new computing resource for joining the computing cluster; and establishing, by the new computing resource, membership in the computing cluster after installing the software bundle.

2. The computer- implemented method of claim 1, wherein, Receiving the reply comprising the reference to the software bundle is a result of the existing member of the computing cluster determining that the new computing resource is not running the same version of the software bundle as the existing member of the computing cluster.

3. The computer-implemented method of any of claims 1-2, comprising: sending a pre-boot execution environment request prior to sending a request to join the computing cluster; receiving a boot image in response to the pre-boot execution environment request; and initializing, by the new computing resource, using the received boot image. The metadata describing requirements for joining the computing cluster comprises membership authentication information, the method comprising:

4. The computer- implemented method of any one of claims 1-2, wherein, authenticating, by the new computing resource, using the membership authentication information prior to establishing membership in the cluster. The software bundle is contained in a Docker container executable by the new computing resource.

5. The computer- implemented method of any one of claims 1-2, wherein, 6. The computer-implemented method of claim 1, comprising: installing, by the at least one member of the computing cluster, a software or configuration update; and generating a software bundle containing software or configuration included in the update and labeling the software bundle with a release version.

7. A distributed computing management system for managing a computing cluster comprising at least one member, comprising: an existing member of the computing cluster to receive a request from a new computing resource for metadata describing requirements for joining the computing cluster, and in response, to send to the new computing resource the metadata describing requirements for joining the computing cluster, and a reference to a software bundle; and a software package repository to receive a request for the software bundle from the new computing resource, and in response, to provide the software bundle to the new computing resource for installation to configure the new computing resource for joining the computing cluster, wherein after installing the software bundle, the new computing resource is configured to establish membership in the computing cluster. ​ ​ 8. The distributed computing management system of claim 7, wherein, The metadata describing requirements for joining the compute cluster includes information describing a protocol, configuration parameters, and membership authentication information.

9. The distributed computing management system of claim 7 or 8, comprising: The new compute resource to receive a response from the software package repository and execute the software bundle, thereby configuring the new compute resource with all software and configuration necessary to join the compute cluster.

10. The distributed computing management system of claim 7 or 8, wherein, The software bundle is a Docker container.

11. The distributed computing management system of claim 7 or 8, comprising: A credential authority to receive a request regarding authenticating the new compute resource and accepting the new compute resource as a member of the compute cluster and in response accept the new compute resource as a member of the compute cluster.

12. The distributed computing management system of claim 7 or 8, comprising: A pre-boot execution environment server to receive a pre-boot execution environment request from the new compute resource that is not a member of the compute cluster and send a boot image used by the compute cluster to the new compute resource to initialize the new compute resource.

13. The distributed computing management system of claim 7 or 8, wherein, The existing members of the compute cluster are configured to install a software or configuration update and generate a software bundle containing the software or configuration included in the update and tag the software bundle with a release version.

14. A non-transitory computer readable medium comprising instructions stored thereon that when executed are capable of causing one or more processors of a management cloud system to perform operations comprising: sending, by a new compute resource, a request to existing members of a compute cluster comprising at least one member for metadata describing requirements for joining the compute cluster; receiving, by the new compute resource, a reply to the request for metadata in one or more communications, the reply comprising: the metadata describing requirements for joining the compute cluster, and a reference to a software bundle used by devices in the compute cluster; sending, by the new compute resource, a request for the software bundle to a software package repository based on the reply; receiving, by the new compute resource, the software bundle from the software package repository as a response to the request for the software bundle; installing, by the new compute resource, the software bundle to configure the new compute resource for joining the compute cluster; and establishing, by the new compute resource, membership in the compute cluster after installing the software bundle.

15. The non-transitory computer-readable medium of claim 14, wherein, Receiving the reply comprising the reference to the software bundle is a result of the existing members of the compute cluster determining that the new compute resource is not running the same version of the software bundle as the existing members of the compute cluster.

16. The non-transitory computer readable medium of any one of claims 14-15, wherein, The instructions are capable of causing one or more processors of the management cloud system to perform operations comprising: sending a pre-boot execution environment request prior to sending a request to join the compute cluster; receiving a boot image in response to the pre-boot execution environment request; and initializing, by the new computing resource, using the received bootstrap image.

17. The non-transitory computer readable medium of any one of claims 14-15, wherein, The metadata describing requirements for joining the computing cluster includes membership authentication information, wherein the instructions can cause the one or more processors of the management cloud system to perform the following: authenticating, by the new computing resource, using the membership authentication information prior to establishing membership in the existing cluster.

18. The non-transitory computer readable medium of any one of claims 14-15, wherein, The software bundle is contained in a Docker container executable by the new computing resource.

19. The non-transitory computer-readable medium of claim 14, wherein, The instructions can cause the one or more processors of the management cloud system to perform the following: installing, by the at least one member of the computing cluster, a software or configuration update; and generating a software bundle and tagging the software bundle with a release version, the software bundle containing software or configuration included in the update.

20. An apparatus for joining a new computing resource to a computing cluster, comprising: means for sending, by a new computing resource, a request to an existing member of a computing cluster comprising at least one member for metadata describing requirements for joining the computing cluster; means for receiving, by the new computing resource, a reply to the request for metadata in one or more communications, the reply including the metadata describing requirements for joining the computing cluster and a reference to a software bundle used by devices in the computing cluster; means for sending, by the new computing resource, a request for the software bundle to a software package repository based on the reply; means for receiving, by the new computing resource, the software bundle from the software package repository as a response to the request for the software bundle; means for installing, by the new computing resource, the software bundle to configure the new computing resource for joining the computing cluster; and means for establishing, by the new computing resource, membership in the computing cluster after installing the software bundle.

21. The apparatus of claim 20, further comprising: means for implementing the method of any of claims 2-6.

22. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the method of any of claims 1-6.

23. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the steps of the method of any of claims 1-6.

Citation Information

Patent Citations

  • Plug and play cluster deployment

    US20070041386A1