Identifying the nodes contained within a partitioned system
The method for identifying and managing white boxes in a network cloud using LLDP and connectivity matrices addresses operational complexity and cost issues by enabling effective synchronization of distributed hardware white boxes in a network cloud, enhancing network functionality and scalability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2026-03-10
AI Technical Summary
Network operators face challenges in identifying and provisioning multiple software modules across distributed hardware white boxes in a network cloud, leading to operational complexity and increased costs due to the need for synchronized operation of modular network elements.
A method for identifying and assigning unique identities to white boxes based on their functions using Link Layer Discovery Protocol (LLDP) and connectivity matrices, enabling effective operation and management of a partitioned system by distinguishing between element controllers, managed switches, fabric modules, and data path forwarders.
Facilitates efficient identification and management of network elements, reducing operational complexity and costs by ensuring synchronized operation of white boxes in a network cloud, thereby enhancing network functionality and scalability.
Smart Images

Figure 0007827356000001 
Figure 0007827356000002
Abstract
Description
[Technical Field]
[0001] This disclosure relates generally to the field of distributed computing, and more particularly to the operation of disaggregated systems within communication networks.
[0002] BOM: Bill of Materials CSP: Cloud Service Provider DDOS: Distributed Denial-of-Service EC: Elements' Controller FM: Fabric Module MS: Management Switch DPF: Data Path Forwarder LLDP: Link Layer Discovery Protocol NC: Network Cloud NCR: Network Cloud Router NOS: Network Operating System ODM: Original Design Manufacturer ONIE: Open Network Install Environment OS: Operating System QOS: Quality of Service S / N: Serial Number SDN: Software Defined Network VPN: Virtual Private Network WB: White-Box WB-UID: White-Box Unique Identifier White Box: A commodity that is open or industry standards-based hardware for switches and / or routers in the forwarding plane. White boxes provide users with the fundamental hardware elements of their network. [Background technology]
[0003] A partitioned (or distributed) system is a system in which components are located on different networked computers, and they communicate and coordinate their actions by sending messages to each other. The components interact with each other to achieve a common goal. Three important characteristics of distributed systems are concurrent execution of components, lack of a global clock, and independent failure of components.
[0004] Computer programs that run within a distributed system are called distributed programs. There are many different types of implementations for message passing mechanisms, including pure HTTP, RPC-like connectors, and message queues.
[0005] A computer cluster is a set of loosely or tightly connected computers that work together so that in many respects they are viewed as a single entity. Unlike grid computers, computer clusters each have a set of nodes that perform the same tasks, controlled and scheduled by software.
[0006] The components of a cluster are typically connected to each other through a high-speed local area network, and each node (e.g., a computer used as a server) runs its own instance of an operating system. In most situations, all nodes use the same hardware and the same operating system, but in some setups (e.g., when Open Source Cluster Application Resources (OSCAR) is implemented), each computer may use a different operating system or different hardware.
[0007] Clusters are typically deployed to improve performance and availability over a single computer, and are typically significantly more cost-effective than a single computer of similar speed or availability.
[0008] The term Network Cloud (NC) refers to a cloud used to perform network functions such as routing and switching. That is, the term refers to the concept of separating the hardware and software of a network entity. The control plan of a network entity in a Network Cloud is separated from the data path and installed on a local server or within the cloud network. An abstraction layer separates the control elements and makes them independent of the data path-related hardware components. The data path runs on distributed hardware resources such as servers, network interfaces, and white-box devices and can be directly programmed. The Network Cloud concept uses cloud methodologies to perform Software Deterministic Network (SDN) functions such as routing, switching, VPN, QOS, and DDOS mitigation in a more efficient, centrally managed, and easily programmable manner.
[0009] The separation between hardware and software in the network field that exists today has resulted in a new model of network cloud, where optimal utilization of hardware resources is implemented to enable the deployment of a distributed network operating system. Currently, network operators face financial challenges in that network element prices are relatively expensive per device and, consequently, even on a "per port" basis, while revenue per subscriber remains roughly constant and, in some cases, is declining. Clearly, the above challenges impact the profitability of network owners, prompting them to seek ways to implement cost-reduction approaches in their networks. Many network operators and large network owners, such as web-scale operators, have adopted a white-box implementation approach, where white boxes are hardware elements manufactured by silicon merchants (commodity chipsets) using original equipment manufacturing (OEM) techniques. This approach allows network operators to use white boxes manufactured by different manufacturers within the same distributed network cloud cluster, thereby reducing hardware prices to a model of bill-of-materials costs plus an agreed-upon profit margin. However, this approach differs from the traditional approach in that the network element is purchased as a monolithic device that combines hardware and software. As described above, the problems with the hardware portion (i.e., the hardware portion of the network element) are resolved by adopting a white-box approach. However, adopting this approach creates new challenges in the software portion of the solution. Because this approach requires multiple software modules and containers, using a distributed hardware node solution by using multiple hardware white boxes requires that the multiple software modules and containers run synchronously.
[0010] The deployment and provisioning process performed in a partitioned white-box based virtual cluster poses several operational challenges. One such challenge resides in the fact that several modular network elements (nodes) are bundled together to function as a single powerful network element. Each node is responsible for a specialized function within the router and therefore needs to be automatically identified, provisioned, and assigned to the associated software (SW) components.
[0011] The solution provided by the present disclosure provides an apparatus and method for identifying nodes in a dynamic and evolving environment that need to improve the operation of such a network. Summary of the Invention [Problem to be solved by the invention]
[0012] The present disclosure can be summarized by reference to the appended claims.
[0013] An object of the present disclosure is to provide a novel partitioned system that includes multiple white boxes that effectively operate as a single entity (router, switch, etc.), with the functionality associated with that single entity distributed across multiple physical white boxes.
[0014] It is an object of the present disclosure to provide a novel partitioning system and a method for identifying elements included in the system based on their functionality.
[0015] Another object of the present disclosure is to provide a novel approach for identifying nodes in a distributed cluster to enable improved control and operation of a network cloud.
[0016] Other objects of the present disclosure will become apparent from the following description. [Means for solving the problem]
[0017] According to a first embodiment of the present invention, there is provided a split routing system for use in a communications network including a plurality of white boxes, wherein at least four of the plurality of white boxes are each configured to perform a function that is different from a function that at least four other three of the plurality of white boxes are configured to perform, and wherein each of the at least four of the plurality of white boxes is associated with an identity based on its function.
[0018] As used herein and in the claims, the term "cluster" is used to denote a virtual entity that includes multiple nodes, which are one or more element controllers, management switches, fabric modules, and data path forwarders. These nodes operate as a set of loosely or tightly connected computing entities that work together so that in many respects the cluster appears as a single system.
[0019] According to another embodiment, at least some of the plurality of whiteboxes are further identified based on their respective positions within the split routing system.
[0020] According to another embodiment, the functions of each of the at least four of the plurality of white boxes are selected from the group consisting of a data path forwarding, a fabric module, an element controller, and a managed switch.
[0021] According to another aspect of the present disclosure, there is provided a method for use in a split routing system including a plurality of white boxes for identifying each of the plurality of white boxes based on a function that each of the plurality of white boxes is configured to perform within the split routing system, wherein at least four of the plurality of white boxes are each configured to perform a function that is different from a function that at least four other three of the plurality of white boxes are configured to perform, and each of the at least four of the plurality of white boxes is identified based on its function.
[0022] According to another embodiment, the method provided comprises: identifying at least one of the plurality of whiteboxes configured to act as a controller for an element; identifying at least one of the plurality of whiteboxes configured to function as a managed switch; identifying at least one of the plurality of whiteboxes configured to function as a fabric module; identifying at least one of the plurality of whiteboxes configured to function as a datapath forwarding element; Includes:
[0023] According to yet another embodiment, one or more of the identifying steps is based on information obtained by using the Link Layer Discovery Protocol (LLDP).
[0024] According to yet another embodiment, the step of identifying at least one of the plurality of whiteboxes as a node configured to function as a controller of an element is performed by a network orchestrator by the serial number of each whitebox.
[0025] According to another embodiment, the step of identifying at least one of the plurality of whiteboxes as a node configured to function as a managed switch is performed according to the connectivity existing between each whitebox and one or more adjacent whiteboxes functioning as controllers of the element, regardless of whether each whitebox is directly or remotely connected to the controller of the adjacent element. Preferably, said identification is based on at least one connection port according to a connectivity matrix stored in the controller of the adjacent element.
[0026] According to yet another embodiment, identifying at least one of the plurality of white boxes as a node configured to function as a fabric module is performed by generating a fabric module ID through an internal management connection with the management switch, preferably based on at least one connection port according to a connectivity matrix stored in the management switch.
[0027] According to another embodiment, the step of identifying at least one of the plurality of white boxes as a node configured to function as a data path forwarder is performed by generating a data path forwarder ID via an internal management connection with the managed switch.
[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the description, serve to explain the principles of these embodiments disclosed herein. [Brief explanation of the drawings]
[0029] [Figure 1]
[0023] Figure 1 shows an example of a network cloud (NC) that contains various elements within it, where each node in the cluster needs to be identified separately from the other nodes. [Figure 2] FIG. 1 illustrates a method according to an embodiment of the present invention for performing an identification process for multiple white boxes included in a partitioned system. DETAILED DESCRIPTION OF THE INVENTION
[0030] Some of the specific details and values in the following detailed description illustrate particular examples of the present disclosure. However, this description is illustrative and is not intended to limit the scope of the present invention. It will be apparent to those skilled in the art that the claimed methods and apparatus can be implemented by other techniques known in the art. Furthermore, the embodiments described herein include different steps, not all of which are required in all embodiments of the present invention. The scope of the present invention is summarized by reference to the appended claims.
[0031] The present invention relates to a partitioned routing system (e.g., a distributed cluster) that includes a large number of network elements, where most (or all) of these network elements are white boxes. Using such a configuration, on the one hand, it is possible to implement very large systems (i.e., systems with large capacity) at a relatively low cost, and on the other hand, it is possible to easily expand the system by adding new white boxes. However, one of the main drawbacks of implementing such a system is its complexity. For the network cloud to operate properly, the present disclosure proposes to identify the function of each white box, preferably along with its location within the network cloud.
[0032] This disclosure describes a novel approach for identifying multiple white boxes (nodes) within a partitioned routing system (e.g., a distributed cluster), each configured to perform a specific function, such as control, to operate and maintain a powerful network cloud, and which, in combination, can perform various network functions, such as routing and switching, within the network cloud.
[0033] We begin by describing the general architecture of a Network Cloud (NC) and its basic building blocks.
[0034] The network cloud is a relatively new network architecture built for extreme growth, rapid service innovation, and economic profitability. It typically consists of three levels of decomposition: 1. Hardware and software: Software running on white boxes sold directly by manufacturers to cloud service providers (CSPs) on a cost-plus model. This new economic model allows cloud service providers to increase their profitability as demand for their services increases. 2. Router Architecture: Implementing a network cloud architecture means splitting traditional monolithic routers into clusters built from standard servers and multiple white boxes that run routing services. This architecture allows scaling up small routers to edge, aggregation, and large core routers while still using the same hardware solution. This approach simplifies operations and, as a result, reduces operational costs. 3. Data Plane and Control Plane: The data plane of the network cloud runs on white boxes and is designed to scale linearly by simply adding more white boxes. The control plane should typically be based on containerized microservices running different routing services for the different network functions involved (core, edge, aggregation, etc.), with service chaining allowing all routing services to share the same infrastructure.
[0035] The Network Cloud includes the following software building blocks: 1) An operating system that transforms a non-operational (dead) box into a working (active) network element, scalable from supporting a single standalone white box to supporting hundreds of white boxes, thereby creating a true cloud environment within the service provider's network. 2) An orchestrator designed to solve the unique challenges of deploying, integrating, and managing a split network. This block centrally automates and maintains lifecycle management to ensure smooth operation of all network cloud elements.
[0036] Next, a Network Cloud Router (NCR) is formed by using various white boxes (hardware building blocks) that have the following functions: Data Path Forwarder (DPF): A single high-rate forwarding element that is responsible for the data path forwarding process and remembers all relevant data path characteristics such as access lists, QOS, BFD, and Netflow. Fabric Module (FM): Responsible for non-disruptive intra-cluster data path connectivity between multiple data path forwarders contained within a cluster. Element Controller (EC): An external controller for intra-cluster management and network element control and management. For example, a single x86 server can be used to serve different network cloud routers in a collocated configuration. Management Switch (MS): Responsible for intra-cluster management and connectivity control between data path forwarders, fabric modules, and element controllers. Management switches from all racks collect all management traffic information from all rack elements. The management traffic is then aggregated by the management switches in the x86 server racks.
[0037] One of the challenges of implementing such a configuration is how to provide access to each of the hardware (HW) building blocks to operate the cluster properly, i.e., how to identify each of the hardware (HW) building blocks to ensure the desired network functionality is achieved.
[0038] In the following example, a cluster information processing flow is illustrated, and each of the hardware (HW) building blocks of the network cloud is identified along with their interdependencies. The illustrated flow includes the following main steps: Assigning the ID of the control device of the element, Assigning names to managed switches, Fabric module ID assignment, Data path forwarder ID assignment.
[0039] A general assumption is that a managed switch must be connected to the node elements (data path forwarders / fabric modules) as illustrated in Figure 1, but the element's controller can operate without being connected to a managed switch for pre-provisioning purposes.
[0040] FIG. 1 illustrates a schematic diagram of a system according to the present invention, including a number of different white boxes operating as managed switches, element controllers, data path forwarders, or fabric modules, an operating system (OS) associated with each white box, and the services performed by the white boxes according to their functions.
[0041] FIG. 2 illustrates a method according to an embodiment of the present invention for performing an identification process for multiple white boxes contained within a partitioned system.
[0042] First, various white boxes acting as managed switches, element controllers, datapath forwarders, or fabric modules continuously run the Link Layer Discovery Protocol (LLDP) on the control ports to which they are connected (step 10).
[0043] Next, the identification of the element's controller (EC) begins (step 20); in this example, this identification is performed via the network orchestrator by the serial number of the white box configured to act as the element's controller. Assume that for high availability purposes, it is decided that two white boxes contained in a split system will be used as element controllers, then ID0 is assigned (e.g., by default) to the preferred element's controller and ID1 to the other element's controller. Preferably, IDs should not change during operation, but can only be changed during deployment, where they remain fixed thereafter.
[0044] Next, the step of identifying a managed switch (MS) (step 30) is described. Its host / system name is determined by its connectivity with its neighboring element controllers. The element controllers can configure the managed switch through their shell prompt via the REST-API, i.e. Representational State Transfer (REST), which is a set of rules that must be followed when creating the associated API. The managed switch can be connected to these neighboring element controllers directly or remotely. The identification criteria is based on at least one connection port according to the connectivity matrix defined by the element controller. Below is an example of a typical process flow: The element controller is provided with information about the serial numbers (S / N) of the managed switches and their assigned IP addresses (e.g., the element controller's DHCP server can be used to extract the IP addresses from the managed switch serial numbers), Repeatedly (e.g., every few minutes) querying the assigned DHCP IP address to request information about the managed switch's neighbors, Meanwhile, the managed switch runs a link layer discovery protocol to obtain information about its neighbors. The managed switch responds to queries from the element's control unit and provides the requested information obtained during execution of the link layer discovery protocol; The element's control unit compares the information obtained from the response of the managed switch with its connectivity matrix and sets, via the REST API, the name of the responding managed switch (for example, S / N ABCDEFGH), e.g. "MS-A0"; The element's control device then stores the configured name in its operational database together with information about its managed switch.
[0045] Next, a step (step 40) is performed to identify the fabric modules (FMs) of the split system (cluster). Fabric module IDs are automatically generated and assigned to the fabric modules through their internal management connection with the managed switch using a protocol such as the Link Layer Discovery Protocol. Identification criteria can be at least one connected port through the connectivity matrix of the managed switch. Other factors considered in making this decision are, for example, differentiation and potential collisions between start-up mode and operational mode, where no new IDs are assigned (only serial number information is available).
[0046] The following example illustrates the latter embodiment. The fabric module uses a link layer discovery protocol to indicate that it is connected to "MS-XXX" via port ID "YY". The element controller obtains the following data from the information provided by the fabric module: Node type: Fabric module Fabric module serial number - Connection port with management switch: "MS-A0" with port ID "ZZ" The element's controller then compares the retrieved data with the data stored in its connectivity matrix and assigns the ID "TT" to that fabric module node.
[0047] The last type of white box that needs to be identified is a node of type Data Path Forwarder (DPF) (step 50). The ID of the Data Path Forwarder is generated automatically through the internal management connection with the managed switch, and the identification is performed using a link layer discovery protocol.
[0048] The identification criterion used is the use of at least one connection port by the connectivity matrix of the managed switch. Other factors taken into account when making this decision are, for example, differentiation and potential collisions between start-up mode and operational mode, where no new IDs are assigned (only serial number information is available). The following example illustrates the latter embodiment. Node Type: Datapath Forwarder Data path forwarder serial number Connection port
[0049] Another important parameter that should preferably be considered is the number of ports that must be correctly connected to make the data path forwarder work.
[0050] The element's controller then compares the acquired data with the data stored in the connectivity matrix, determines compatibility with the number of connection ports based on a predefined threshold, and assigns a data path forwarder ID to the node.
[0051] The present invention has been described using detailed descriptions of embodiments that are provided by way of example only and are not intended to limit the scope of the invention. The described embodiments include different configurations, and not all configurations are required in all embodiments of the invention. Some embodiments of the invention utilize only some of the configurations or possible combinations of configurations. Variations of the described embodiments of the invention, as well as embodiments of the invention that include different combinations of the configurations shown in the described embodiments, will be apparent to those skilled in the art. The scope of the invention is limited only by the following claims.
Claims
1. 1. A split routing system for use in a communications network including a plurality of white boxes that are switching and / or routing hardware, the plurality of white boxes effectively operating as a single routing and / or switching entity, functionality associated with the single routing and / or switching entity being distributed across the plurality of white boxes in the communications network, at least four white boxes in the plurality of white boxes being configured to perform functions that are different from functions configured to be performed by other three white boxes in the at least four white boxes, and for each of the at least four white boxes in the plurality of white boxes, a network orchestrator determines an identity of each white box in the at least four white boxes based on information related to the functionality of each white box in the at least four white boxes.
2. 2. The split routing system of claim 1, wherein the function of each of the at least four white boxes is selected from the group consisting of a data path forwarding, a fabric module, an element controller, and a managed switch.
3. 1. A method for use in a split routing system including a plurality of white boxes that are switching and / or routing hardware, the plurality of white boxes effectively operating as a single routing and / or switching entity, and functionality associated with the single routing and / or switching entity being distributed across the plurality of white boxes within a communications network, for identifying each of the plurality of white boxes based on a function that each of the plurality of white boxes is configured to perform, the method comprising: at least four white boxes among the plurality of white boxes configured to perform a function that is different from the functions that other three white boxes among the at least four white boxes are configured to perform; and for each of the at least four white boxes among the plurality of white boxes, a network orchestrator determines the identity of each of the at least four white boxes based on information related to the function of each of the at least four white boxes.
4. identifying at least one of the plurality of whiteboxes as a node configured to act as a controller of an element; identifying at least one of the plurality of white boxes as a node configured to function as a managed switch; identifying at least one of the plurality of white boxes as a node configured to function as a fabric module; identifying at least one of the plurality of whiteboxes as a node configured to function as a data path forwarder; 4. The method of claim 3, comprising:
5. 5. The method of claim 4, wherein one or more of the identifying steps is based on information obtained by using a Link Layer Discovery Protocol (LLDP).
6. 5. The method of claim 4, wherein the step of identifying at least one of the plurality of whiteboxes as a node configured to act as a controller of the element is performed by a network orchestrator by a serial number of at least one whitebox configured to act as a controller of the element.
7. 7. The method of claim 6, wherein identifying at least one of the plurality of white boxes as a node configured to function as the managed switch is performed by connectivity existing between at least one of the plurality of white boxes functioning as the managed switch and one or more adjacent white boxes functioning as controllers of the elements, regardless of whether at least one of the plurality of white boxes functioning as the managed switch is directly or remotely connected to a controller of the adjacent element.
8. 8. The method of claim 7, wherein said identification is based on at least one connection port according to a connectivity matrix stored in a controller of an adjacent element.
9. 7. The method of claim 6, wherein identifying at least one of the plurality of white boxes as a node configured to function as the fabric module is performed by identifying an internal management connection of at least one of the plurality of white boxes with the management switch.
10. 10. The method of claim 9, wherein the identification is based on at least one connection port according to a connectivity matrix stored in the managed switch.
11. 7. The method of claim 6, wherein identifying at least one of the plurality of white boxes as a node configured to function as the data path forwarder is performed by determining an identity of the data path forwarder through an internal management connection with the managed switch.
Citation Information
Patent Citations
Cell site routing based on latency
US20200008125A1
Orchestration of activities of entities operating in a network cloud
WO2020121293A1
Secured deployment and provisioning of a white-box based cluster
WO2020121295A1