Data Center Resource Monitoring for Managing Message Load Balancing with Reordering Consideration

A distributed architecture with protocol-specific strategies for message handling addresses data center monitoring challenges, ensuring reliable and efficient network resource management by enforcing compliance with protocol requirements, thereby improving performance visibility and optimization.

CN113961303BActive Publication Date: 2025-07-15HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011080132.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-20
Filing Date
2020-10-10
Publication Date
2025-07-15
Estimated Expiration
2041-04-03

AI Technical Summary

Technical Problem

The existing data center monitoring mechanism cannot effectively handle out-of-order or lost message data, resulting in network infrastructure management failure and fail to provide near real-time and historical performance visibility and dynamic optimization.

Method used

By deploying an entry engine in the data center, using distributed architecture and load balancing technology, we ensure that message data meets the requirements of their respective protocol types, reorder and distribute to appropriate collector applications, and realize effective storage and analysis of message data.

Benefits of technology

It realizes effective management of message data in the data center, ensures message order and integrity, provides near-real-time performance visibility and dynamic optimization, and supports network infrastructure management of multi-tenant and dynamic enterprise clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961303B_ABST
    Figure CN113961303B_ABST
Patent Text Reader

Abstract

Describes techniques for resource monitoring and management message reordering in a data center. In one example, a computing system includes an ingress engine for receiving messages from network devices in a data center, the data center including a plurality of network devices and computing systems; and in response to receiving a message from a network device in the data center, transmitting the message to an appropriate collector application corresponding to the protocol type of the message, the protocol type of the message meeting at least one requirement for data stored in a message stream transmitted from one or more network devices to the computing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to monitoring and improving the performance of cloud data centers and networks. Background Art

[0002] Virtual data centers are becoming the core foundation of modern information technology (IT) infrastructure. In particular, modern data centers have widely utilized virtualized environments in which virtual hosts such as virtual machines or containers are deployed on and executed on the underlying computing platform of physical computing devices.

[0003] The virtualization of large-scale data centers can provide several advantages. One advantage is that virtualization can significantly improve efficiency. As the underlying physical computing devices (i.e., servers) become more and more powerful, with the emergence of multi-core microprocessor architectures with a large number of cores per physical CPU, virtualization becomes easier and more efficient. A second advantage is that virtualization can provide important control over the infrastructure. As physical computing resources in, for example, cloud-based computing environments become replaceable resources, the provisioning and management of the computing infrastructure become easier. Thus, in addition to the efficiency and increased return on investment (ROI) provided by virtualization, enterprise IT staff generally prefer virtualized computing clusters in data centers because of their management advantages.

[0004] Some data centers include mechanisms for monitoring resources within the data center, collecting various data including statistics measuring the performance of the data center, and then using the collected data to support the management of the virtualized networking infrastructure. As the number of objects on the network and the metrics they generate grow, data centers cannot rely on traditional mechanisms to monitor data center resources, such as the health of the network and any potential risks. These mechanisms impose limitations on the scale, availability, and efficiency of network elements and generally require additional processing, such as directly polling network elements periodically. Summary of the Invention

[0005] This disclosure describes techniques for monitoring, scheduling, and performance management for a computing environment, such as virtualized infrastructure deployed within a data center. These techniques provide visibility into operational performance and infrastructure resources. As described herein, these techniques can utilize analytics in a distributed architecture to provide near or seemingly near real-time and historical monitoring, performance visibility, and dynamic optimization to improve orchestration, security, billing, and planning within the computing environment. The techniques can provide advantages within, for example, hybrid, private, or public enterprise cloud environments. The techniques can accommodate various virtualization mechanisms, such as containers and virtual machines, to support multi-tenant, dynamic, and evolving enterprise clouds.

[0006] Aspects of the present disclosure generally describe techniques and systems for maintaining the viability of various captured data while monitoring network infrastructure elements and other resources in any example data center. These techniques and systems rely on policies corresponding to the protocol types of the captured data to determine how to manage (e.g., properly convey) messages that transmit the captured data. According to one example policy, certain protocol types require proper ordering of messages on a destination server; any deviation from the correct order may prevent the stored data of the messages from becoming a viable data set for performance visibility and dynamic optimization to improve orchestration, security, billing, and planning in a computing environment. Techniques and systems different from the present disclosure do not rely on such policies, and thus, any message that arrives at a destination server out of order or arrives at the destination server without a proper collector application may render the stored data of the message unavailable or invalid, thereby undermining the management of the network infrastructure.

[0007] Generally, policy enforcement as described herein refers to meeting one or more requirements for making stored message data (e.g., captured telemetry data) a viable data set according to the telemetry protocol type of the message. Examples of policy enforcement can involve any combination of the following: modifying the structure of the captured telemetry data or the telemetry data itself, providing scalability via load balancing, high availability, and / or auto-scaling, identifying the appropriate collector application type corresponding to the telemetry protocol, and / or the like.

[0008] In an example computing system, an example ingress engine performs policy enforcement on incoming messages to ensure that the stored message data complies with each corresponding policy based on the protocol type of the message. One example policy sets one or more requirements for the captured telemetry data described above, such that if each requirement is met, the captured telemetry data can be extracted from the message, assembled into a viable data set, and then processed for network infrastructure monitoring and management processes. Making these messages meet their corresponding policy requirements enables one or more applications running on a destination server to collect and successfully analyze the captured telemetry data to obtain meaningful information about data center resources (e.g., health or risk assessment). The situation where message chaos or loss causes data center outages is (mostly) completely mitigated or eliminated. For this reason and other reasons described herein, the present disclosure provides technical improvements in computer networking and the practical application of technical solutions to problems of messages of different computer networking protocol types.

[0009] In one example, a method includes: an ingress engine running on a processing circuit of a computing system receiving a message from a network device in a data center, the data center including a plurality of network devices and the computing system; and in response to receiving the message from the network device in the data center, the ingress engine delivering the message to an appropriate collector application corresponding to the message protocol type, the message protocol type meeting at least one requirement for data stored in messages transmitted from the plurality of network devices to the computing system.

[0010] In another example, a computing system includes: a memory; and a processing circuit communicatively coupled to the memory, the processing circuit configured to execute logic including an ingress engine, the ingress engine configured to: receive a message from a network device in a data center including a plurality of network devices and the computing system; and in response to receiving the message from the network device in the data center, deliver the message to an appropriate collector application corresponding to the protocol type of the message, the protocol type of the message meeting at least one requirement for data stored in a message flow transmitted from one or more network devices to the computing system.

[0011] In another example, a computer-readable medium includes instructions for causing a programmable processor to: receive a message from a network device in a data center including a plurality of network devices and a computing system via an ingress engine running on a processing circuit of the computing system; and in response to receiving the message from the network device in the data center, transmit the message to an appropriate collector application corresponding to the protocol type of the message, the protocol type of the message meeting at least one requirement for data stored in a message flow transmitted from one or more network devices to the computing system.

[0012] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a conceptual diagram illustrating an example network in accordance with one or more aspects of the present invention, the example network including an example data center in which resources are monitored using message load balancing with reordering considerations.

[0014] Figure 2 is a more detailed block diagram of a portion of an example data center in accordance with one or more aspects of the present disclosure Figure 1 and in which resources are monitored by an exemplary server using managed message load balancing with reordering considerations.

[0015] Figure 3 is a diagram illustrating in accordance with one or more aspects of the present disclosureFigure 1 Block diagram of an example analysis cluster of an example data center.

[0016] Figure 4A It shows an example configuration of an example analysis cluster in an example data center according to one or more aspects of the present disclosure Figure 1 Block diagram of an example configuration of an example analysis cluster in an example data center. Figure 4B It shows an example routing instance of an analysis cluster according to one or more aspects of the present disclosure Figure 4A Block diagram of an example routing instance of an analysis cluster. Figure 4C It shows an example failover of an example analysis cluster according to one or more aspects of the present disclosure Figure 4A Block diagram of an example failover of an example analysis cluster. Figure 4D It shows an example auto - scaling operation in an example analysis cluster according to one or more aspects of the present disclosure Figure 4A Block diagram of an example auto - scaling operation in an example analysis cluster.

[0017] Figure 5 Flowchart showing an example operation of a computing system in an example analysis cluster of FIG. 4 according to one or more aspects of the present disclosure. Detailed Description

[0018] As described herein, a data center may deploy various network infrastructure resource monitoring / management mechanisms, some of which operate differently from others. Some implement specific protocol types through which various data can be captured at the network device level and then transmitted to a dedicated set of computing systems (e.g., servers forming a cluster, which are referred to herein as an analysis cluster) for processing. To successfully analyze data from network devices, the data center may establish policies with different requirements for the computing systems to follow in order to ensure successful analysis of data of any protocol type. The data may be passed as messages and then extracted for analysis; however, non - compliance with the policies typically results in analysis failure.

[0019] Example reasons for failed analysis include lost or corrupted messages, messages arriving out of order, messages having incorrect metadata (e.g., in the message header), messages having data in incorrect formats, etc. The data model for each message protocol type may need to conform to a dedicated set of requirements. Conventionally, analysis clusters within a data center cannot ensure compliance with the requirements corresponding to any protocol type of the relevant message data.

[0020] To overcome these and other limitations, the present disclosure describes techniques that ensure compliance with a set of dedicated requirements for any given message protocol type. Protocol types based on the UDP transport protocol (such as telemetry protocols) do not use sequence numbers to maintain the order of the message flow, which can cause difficulties when dealing with out-of-order messages. Collector applications for some protocol types cannot use out-of-order messages (e.g., due to dependencies between messages in the same message flow), and thus require at least some messages belonging to the same message flow to arrive in sequence or chronological order. However, these collector applications cannot rely on any known network function virtualization infrastructure (NFVi) mechanisms for reordering. Because there is no reordering mechanism, or because some messages are dropped or split between multiple nodes, collector applications according to one or more protocol types may require all message data, and thus, corresponding policies may require the same collector application to receive each message of the message flow. Other protocol types can implement a data model that is not affected by out-of-order messages. In such a data model, the data stored in one message does not depend on the data stored in another message. Corresponding policies for other protocol types may require multiple collector applications to process different messages (e.g., in parallel). Thus, one requirement in any policy described herein is to identify appropriate applications to receive each message of the message flow based on the message protocol type / data model's tolerance for reordering.

[0021] Another policy may require retaining specific message attributes (such as source information) in the message header (e.g., to conform message metadata to the data model of the protocol type); otherwise, these attributes may be overwritten, resulting in analysis failure. Complying with this requirement enables compatible collector applications to correctly extract message metadata and then correctly analyze the message payload data. An example technique for correcting any non-compliance can modify the message metadata and / or the message protocol type policy. Once the non-compliance is resolved, the message can be transmitted to the appropriate collector application.

[0022] Some of the techniques and systems described herein enable policy enforcement via the verification of telemetry data, which is stored and transmitted in message form. Example policies can encode requirements for the telemetry data to make it feasible for analysis in the distributed architecture of a data center. Telemetry data generated by sensors, agents, devices, and / or the like in the data center can be transmitted in message form to one or more dedicated destination servers. In these servers, various applications analyze the telemetry data stored in each message and obtain measurements, statistics, or some other meaningful information to evaluate the operating state (e.g., health / risk) of resources in the data center. Further analysis of the telemetry data may yield meaningful information for other tasks related to managing the network infrastructure and other resources within the data center.

[0023] For a cluster of computing systems that operate as destination servers for messages that carry telemetry data from network devices in a data center, at least one dedicated server, such as a management computing system, is communicatively coupled to the cluster of computing systems and operates as an intermediate destination and access point for messages to be monitored. The dedicated server can include a software-defined network (SDN) controller to perform services for other servers in the cluster, such as operating a policy controller in the data center, configuring ingress engines in the destination servers to enforce requirements for telemetry data to be viable, operating a load-balancing mechanism to distribute compatible collector applications for the protocol types of messages that carry telemetry data, and distributing messages among the destination servers accordingly. Because the load-balancing mechanism does not consider the protocol type of a message when distributing the message to a destination server, it is possible to deliver a message to a destination server without a proper collector application corresponding to the message protocol type that meets the telemetry data requirements being viable. To deliver a message that meets the message protocol type requirements, the ingress engine identifies a proper collector application to receive the message corresponding to the message protocol type.

[0024] Another aspect of network infrastructure resource monitoring / managing in a data center is performance management, which includes load balancing. In a typical deployment, network elements or devices transmit data flows to at least one destination server that acts as a performance management system microservice provider. The data transmitted in the flows can include telemetry data, such as physical interface statistics, firewall filter counter statistics, label switched path (LSP) statistics. The relevant data is exported from source network devices, such as line cards or network processing units (NPUs), by sensors using UDP and arrives as messages via a data port on at least one destination server. Further, an appropriately configured ingress engine running on one destination server will perform policy enforcement as described herein.

[0025] Figure 1 is a conceptual diagram illustrating an example network 105 in accordance with one or more aspects of the present disclosure. The example network 105 includes an example data center 110 in which the performance and usage metrics of infrastructure elements of a cloud-based computing environment are monitored and optionally includes internal processor metrics related to resources shared by multiple processes. Figure 1FIG. 0 illustrates an example implementation of network 105 and data center 110, which hosts one or more cloud-based computing networks, domains, or projects, commonly referred to herein as cloud computing clusters. The cloud-based computing clusters can be co-located in a common overall computing environment, such as a single data center, or can be distributed throughout the environment, such as across different data centers. The cloud-based computing clusters can be, for example, different cloud environments, such as various combinations of OpenStack cloud environments, Kubernetes cloud environments, or other computing clusters, domains, networks, etc. In other cases, other implementations of network 105 and data center 110 may be suitable. Such implementations can include subsets of the components included in the example of Figure 1 and / or can include additional components not shown in Figure 1 FIG.

[0026] In the example of Figure 1 FIG., data center 110 provides an operating environment for applications and services for customer 104 coupled to data center 110 via service provider network 106. Although the functions and operations described in connection with network 105 of Figure 1 FIG. may be shown as being distributed among multiple devices in Figure 1 FIG., in other examples, features and techniques attributed to one or more devices in Figure 1 FIG. may be performed internally by local components of one or more such devices. Similarly, one or more such devices may include certain components and perform various techniques that may otherwise be attributed to one or more other devices in the description herein. Additionally, certain operations, techniques, features, and / or functions may be described in connection with Figure 1 FIG., or otherwise performed by specific components, devices, and / or modules. In other examples, such operations, techniques, features, and / or functions may be performed by other components, devices, or modules. Thus, even if not specifically described in this manner herein, some operations, techniques, features, and / or functions attributed to one or more components, devices, or modules may be attributed to other components, devices, and / or modules.

[0027] Data center 110 hosts infrastructure devices, such as networking and storage systems, redundant power supplies, and environmental controls. Service provider network 106 can be coupled to one or more networks managed by other providers and can thus form part of a large-scale public network infrastructure such as the Internet.

[0028] In some examples, data center 110 can represent one of many geographically distributed network data centers. As shown in Figure 1As shown in the example of, data center 110 is a facility that provides network services to customer 104. Customer 104 can be a collective entity, such as an enterprise, government, or individual. For example, a network data center can host web services for multiple enterprises and end users. Other exemplary services may include data storage, virtual private networks, business engineering, file services, data mining, scientific or supercomputing, etc. In some examples, data center 110 is a separate network server, network peer, or other.

[0029] In Figure 1 the example of, data center 110 includes a collection of storage systems and application servers, including servers 126A through 126N (collectively referred to as "servers 126"), which are interconnected via a high-speed switching fabric 121 provided by one or more tiers of physical network switches and routers. Servers 126 serve as the physical computing nodes of the data center. For example, each server 126 can provide an operating environment for executing one or more customer-specific virtual machines 148 ( Figure 1 "VM" in ) or other virtualized instances such as containers. Each of servers 126 can alternatively be referred to as a host computing device, or more simply as a host. Servers 126 can execute one or more virtual instances, such as virtual machines, containers, or other virtual execution environments for running one or more services, such as virtual network functions (VNFs).

[0030] Although not shown, switching fabric 121 can include top-of-rack (TOR) switches coupled to the distribution layer of a chassis switch, and data center 110 can include one or more non-edge switches, routers, hubs, gateways, security devices such as firewalls, intrusion detection, and / or intrusion prevention devices, servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices. Switching fabric 121 can perform layer 3 routing to route network traffic between data center 110 and customer 104 via service provider network 106. Gateway 108 is used to forward and receive packets between switching fabric 121 and service provider network 106.

[0031] In accordance with one or more examples of the present disclosure, a software-defined networking (“SDN”) controller 132 provides a logically and in some cases physically centralized controller to facilitate the operation of one or more virtual networks within a data center 110. Throughout this disclosure, the terms SDN controller and virtual network controller (“VNC”) may be used interchangeably. In some examples, the SDN controller 132 operates in response to configuration inputs received via a northbound API 131 from an orchestration engine 130, which in turn operates in response to configuration inputs received from an administrator 128 interacting with and / or operating a user interface device 129. Additional information regarding the operation of the SDN controller 132 in cooperation with other devices in the data center 110 or other software-defined networks can be found in International Application No. PCT / US 2013 / 044378, filed on June 5, 2013, entitled “PHYSICAL PATH DETERMINATION FOR VIRTUAL NETWORK PACKET FLOWS,” which is incorporated herein by reference as if fully set forth herein.

[0032] The user interface device 129 may be implemented as any suitable device for interactively presenting output and / or accepting user input. For example, the user interface device 129 may include a display. The user interface device 129 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or an administrator 128. The user interface device 129 may represent, for example, a workstation, a laptop or notebook computer, a desktop computer, a tablet computer, or any other computing device operable by a user and / or presenting a user interface in accordance with one or more aspects of the present disclosure. In some examples, the user interface device 129 may be physically separated from and / or located at a different location than the policy controller 201. In such examples, the user interface device 129 may communicate with the policy controller 201 via a network or other communication means. In other examples, the user interface device 129 may be a local peripheral of the policy controller 201 or may be integrated into the policy controller 201.

[0033] In some examples, the orchestration engine 130 manages the functions of the data center 110, such as compute, storage, networking, and application resources. For example, the orchestration engine 130 may create virtual networks for tenants within the data center 110 or across data centers. The orchestration engine 130 may attach virtual machines (VMs) to a tenant's virtual network. The orchestration engine 130 may connect a tenant's virtual network to an external network, such as the Internet or a VPN. The orchestration engine 130 may enforce security policies across a group of VMs or at the boundary of a tenant's network. The orchestration engine 130 may deploy network services (e.g., load balancers) within a tenant's virtual network.

[0034] In some examples, the SDN controller 132 manages the network and networking services, such as load balancing, security, and allocates resources from the server 126 to various applications via the southbound API 133. That is, the southbound API 133 represents a set of communication protocols utilized by the SDN controller 132 to make the actual state of the network equal to the desired state specified by the orchestration engine 130. For example, the SDN controller 132 implements high-level requests from the orchestration engine 130 by configuring the following: physical switches, such as TOR switches, chassis switches, and fabric 121; physical routers; physical service nodes, such as firewalls and load balancers; and virtual services, such as virtual firewalls in VMs, etc. The SDN controller 132 maintains routing, networking, and configuration information within a state database.

[0035] Typically, traffic between any two network devices, such as between network devices (not shown) in the fabric 121, or between the server 126 and the client 104, or between the servers 126, can traverse the physical network using many different paths. For example, there may be multiple different paths with equal costs between two network devices. In some cases, a routing policy called multipath routing can be used at each network switching node to distribute packets belonging to the network traffic from one network device to another to various possible paths. For example, Internet Engineering Task Force (IETF) RFC 2992, “Analysis of an Equal-Cost Multi-Path Algorithm” describes a routing technique for routing packets along multiple paths of equal cost. The technique of RFC 2992 analyzes a specific multipath routing policy that involves hashing packet header fields to assign flows to bins and sending all packets from a particular network flow over a deterministic path.

[0036] For example, a "flow" can be defined by the headers of the packets or the five values used in the header of a "five-tuple", namely the protocol used to route packets through the physical network, the source IP address, the destination IP address, the source port, and the destination port. For example, the protocol specifies the communication protocol, such as TCP or UDP, and the "source port" and "destination port" refer to the source port and destination port of the connection. A set of one or more packet data units (PDUs) that match a specific flow entry represents a flow. Any parameter of the PDU can be used to roughly classify the flow, such as the source and destination data link (e.g., MAC) and network (e.g., IP) addresses, virtual local area network (VLAN) tags, transport layer information, multi-protocol label switching (MPLS) or generalized MPLS (GMPLS) tags, and the ingress port of the network device that receives the flow. For example, a flow can be all the PDUs sent in a Transmission Control Protocol (TCP) connection, all the PDUs sourced from a specific MAC address or IP address, all the PDUs with the same VLAN tag, or all the PDUs received at the same switch port.

[0037] The virtual router 142 (virtual routers 142A to 142N, Figure 1 collectively referred to as "virtual router 142" herein) executes multiple routing instances for the corresponding virtual networks in the data center 110 and routes packets to the appropriate virtual machines executing in the operating environment provided by the servers 126. Each server 126 can include a virtual router. For example, a packet received by the virtual router 142A of the server 126A from the underlying physical network infrastructure can include an external header to allow the physical network infrastructure to tunnel the payload or "inner packet" to the physical network address of the network interface of the server 126A. The external header can include not only the physical network address of the server network interface, but also a virtual network identifier, such as a VxLAN tag or a multi-protocol label switching (MPLS) tag that identifies one of the virtual networks, and the corresponding routing instance executed by the virtual router. The inner packet includes an inner header that has a destination network address that conforms to the virtual network addressing space of the virtual network identified by the virtual network identifier.

[0038] In some aspects, a virtual router buffers and aggregates packets transmitted over multiple tunnels received from an underlying physical network fabric and then delivers them to an appropriate routing instance for the packets. That is, a virtual router executed on one of the servers 126 can receive inbound tunnel packets of a packet flow from one or more TOR switches in the fabric 121 and process the tunnel packets to build a single aggregated tunnel packet before routing the tunnel packets to locally-executed virtual machines for forwarding to the virtual machines. That is, the virtual router can buffer multiple inbound tunnel packets and construct a single tunnel packet where the payloads of the multiple tunnel packets are combined into a single payload and the outer / overlapping headers on the tunnel packets are removed and replaced with a single header virtual network identifier. In this way, the virtual router can forward the aggregated tunnel packet to the virtual machine as if a single inbound tunnel packet had been received from the virtual network. Additionally, to perform the aggregation operation, the virtual router can utilize a kernel-based offload engine that seamlessly and automatically orchestrates the aggregation of the tunnel packets. Other example techniques for the virtual router to forward traffic to customer-specific virtual machines executed on the servers 126 are described in U.S. Patent Application 14 / 228,844, titled "PACKET SEGMENTATION OFFLOAD FOR VIRTUAL NETWORKS", which is incorporated herein by reference.

[0039] In some example implementations, a virtual router 142 executed on a server 126 steers received inbound tunnel packets among multiple processor cores to facilitate packet processing load balancing among the cores when processing the packets for routing to one or more virtual and / or physical machines. As an example, server 126A includes multiple network interface cards and multiple processor cores to execute virtual router 142A and steers received packets among the multiple processor cores to facilitate packet processing load balancing among the cores. For example, a particular network interface card of server 126A can be associated with a designated processor core to which the network interface card directs all received packets. Depending on a hash function applied to at least one of the internal and external packet headers, instead of processing each received packet, various processor cores offload a flow to one or more other processor cores for processing to utilize the available work cycles of the other processor cores.

[0040] In Figure 1In the example, data center 110 further includes a policy controller 201 that provides monitoring, scheduling, and performance management for data center 110. The policy controller 201 interacts with monitoring agents 205 in at least some of the physical servers deployed in each physical server 216 to monitor the resource usage of physical computing nodes and any virtualized hosts such as VM 148 running on the physical host. In this way, the monitoring agents 205 provide a distributed mechanism for collecting a variety of usage metrics and for the local enforcement of policies installed by the policy controller 201. In an example implementation, the monitoring agents 205 run at the lowest level "computing nodes" of the infrastructure of data center 110, which provide computing resources to execute application workloads. The computing nodes can be, for example, bare-metal hosts of server 126, virtual machines 148, containers, etc.

[0041] The policy controller 201 obtains usage metrics from the monitoring agents 205 and constructs a dashboard 203 (e.g., a collection of user interfaces) to provide visibility into the operational performance and infrastructure resources of data center 110. The policy controller 201 can, for example, transmit the dashboard 203 to the UI device 129 for display to the administrator 128. Additionally, the policy controller 201 can apply analytics and machine learning to the collected metrics to provide near or seemingly near real-time and historical monitoring, performance visibility, and dynamic optimization to improve orchestration, security, billing, and planning in data center 110.

[0042] As Figure 1 shown in the example, the policy controller 201 can define and maintain a rule base as a policy set 202. The policy controller 201 can manage the control of each of the servers 126 based on the policy set 202. The policies 202 can be created or exported in response to an input from the administrator 128 or in response to an operation performed by the policy controller 201. The policy controller 201 can, for example, observe the operation of data center 110 over a period of time and apply machine learning techniques to generate one or more policies 202. The policy controller 201 can periodically, occasionally, or continuously refine the policies 202 when making further observations about data center 110.

[0043] The policy controller 201 (e.g., the analysis engine within the policy controller 201) can determine how to deploy, implement, and / or trigger policies on one or more servers 126. For example, the policy controller 201 can be configured to push one or more policies 202 to one or more policy agents 205 executing on the servers 126. The policy controller 201 can receive information about internal processor metrics from one or more policy agents 205 and determine whether the rule conditions for one or more metrics are met. The policy controller 201 can analyze the internal processor metrics received from the policy agents 205 and, based on this analysis, instruct or cause one or more policy agents 205 to perform one or more actions to modify the operation of the servers associated with the policy agents.

[0044] In some examples, the policy controller 201 can be configured to determine and / or identify elements in the form of virtual machines, containers, services, and / or applications executing on each server 126. As used herein, resources are generally referred to as consumable components of the virtualization infrastructure, i.e., the components used by the infrastructure, such as CPUs, memory, disks, disk I / O, network I / O, virtual CPUs, and Contrail vRouters. Resources can have one or more characteristics, each associated with a metric analyzed and optionally reported by the policy agent 205 (and / or the policy controller 201). A list of example raw metrics describing resources is provided below with reference to Figure 2 a list of example raw metrics describing resources.

[0045] Generally, infrastructure elements, also referred to herein as elements, are components of the infrastructure that include or consume consumable resources to facilitate operation. Example elements include hosts, physical or virtual network devices, instances (e.g., virtual machines, containers, or other virtual operating environment instances), collections, projects, and services. In some cases, an element may be a resource of another element. Virtual network devices can include, for example, virtual routers and switches, vRouters, vSwitches, Open vSwitches, and Virtual Tunnel Forwarders (VTFs). A metric is a value used to measure the amount of a resource with respect to a characteristic of the resource consumed by an element.

[0046] The policy controller 201 can also analyze the internal processor metrics received from the policy agents 205 and classify one or more virtual machines 148 based on the degree to which each virtual machine uses the shared resources of the server 126 (e.g., the classification may be CPU-bound, cache-bound, memory-bound). The policy controller 201 can interact with the orchestration engine 130 to cause the orchestration engine 130 to adjust the deployment of one or more virtual machines 148 on the server 126 based on the classification of the virtual machines 148 executing on the server 126.

[0047] The policy controller 201 may be further configured to report information about whether the conditions of a rule are met to a client interface associated with the user interface device 129. Alternatively, or additionally, the policy controller 201 may be further configured to report information about whether the rule conditions are met to one or more policy agents 205 and / or the orchestration engine 130.

[0048] The policy controller 201 may be implemented as any suitable computing device or implemented within any suitable computing device, or implemented across multiple computing devices. The policy controller 201 or components of the policy controller 201 may be implemented as one or more modules of a computing device. In some examples, the policy controller 201 may include multiple modules executing on a type of computing node (e.g., an "infrastructure node") included within the data center 110. Such a node may be an OpenStack infrastructure service node or a Kubernetes master node, and / or may be implemented as a virtual machine. In some examples, the policy controller 201 may have network connections to some or all of the other computing nodes within the data center 110, and may also have network connections to other infrastructure services that manage the data center 110.

[0049] One or more policies 202 may include instructions to cause one or more policy agents 205 to monitor one or more metrics associated with the server 126. One or more policies 202 may include instructions to cause one or more policy agents 205 to analyze one or more metrics associated with the server 126 to determine whether the conditions of a rule are met. One or more policies 202 may alternatively or additionally include instructions to cause the policy agent 205 to report one or more metrics to the policy controller 201, including whether those metrics meet the conditions of the rule associated with the one or more policies 202. The reported information may include raw data, summary data, and sampled data specified or required by one or more policies 202.

[0050] In some examples, the dashboard 203 can be considered a collection of collections of user interfaces that present information about metrics, alerts, notifications, reports, and other information about the data center 110. The dashboard 203 can include one or more user interfaces presented by the user interface device 129. The dashboard 203 can be created, updated, and / or maintained primarily by the policy controller 201 or by a dashboard module executing on the policy controller 201. In some examples, the dashboard 203 can be created, updated, and / or maintained primarily by a dashboard module executing on the policy controller 201. The dashboard 203 and the associated dashboard module can be implemented together by software objects instantiated in memory, the software objects having associated data and / or executable software instructions that provide output data for rendering on a display. Throughout the specification, reference can be made to the dashboard 203 performing one or more functions, and in such cases, the dashboard 203 refers both to the dashboard module and to the collection of dashboard user interfaces and associated data.

[0051] The user interface device 129 can detect an interaction with the user interface from the dashboard 203 as user input (e.g., from the administrator 128). The policy controller can cause aspects of the data center 110 or items executing on one or more virtual machines 148 in the data center 110 to be configured in response to the user's interaction with the dashboard 203, which relates to network resources, data transfer limits or costs, storage limits or costs, and / or accounting reports.

[0052] The dashboard 203 can include a graphical view that provides a quick, visual overview of resource utilization by using an instance of a histogram. The bins of such a histogram can represent the number of instances using a given percentage of a resource such as CPU utilization. By presenting data using a histogram, the dashboard 203 presents information in a manner that allows the administrator 128 to quickly identify instances indicating under-provisioning or over-provisioning if the dashboard 203 is presented at the user interface device 129. In some examples, the dashboard 203 can highlight the resource utilization of instances on a particular item or host, or the total resource utilization across all hosts or items, so that the administrator 128 can understand resource utilization in the context of the entire infrastructure.

[0053] The dashboard 203 may include information related to the costs of using computing, network, and / or storage resources, as well as the costs incurred by a project. The dashboard 203 may also present information regarding the health and risks of one or more virtual machines 148 or other resources within the data center 110. In some examples, "health" may correspond to an indicator reflecting the current state of one or more virtual machines 148. For example, a sample virtual machine exhibiting health issues may currently be operating outside of a user-specified performance policy. "Risk" may correspond to an indicator reflecting the predicted future state of one or more virtual machines 148, such that a sample virtual machine exhibiting risk issues may be unhealthy in the future. The health and risk indicators may be determined based on monitored metrics and / or alerts corresponding to those metrics. For example, if the policy agent 205 does not receive a heartbeat from a host, the policy agent 205 may characterize that host and all of its instances as unhealthy. The policy controller 201 may update the dashboard 203 to reflect the health of the relevant host and may indicate that the reason for the unhealthy state is one or more "missing heartbeats".

[0054] One or more policy agents 205 may execute on one or more servers 126 to monitor some or all of the performance metrics associated with the servers 126 and / or the virtual machines 148 executing on the servers 126. The policy agent 205 may analyze the monitored information and / or metrics and generate operational information and / or intelligence related to the operational state of the servers 126 and / or one or more virtual machines 148 executing on such servers 126. The policy agent 205 may interact with the kernel operating on one or more servers 126 to determine, extract, or receive internal processor metrics associated with the execution at the servers 126 and the use of shared resources by one or more processes and / or virtual machines 148. The policy agent 205 may perform the monitoring and analysis locally at each server 126. In some examples, the policy agent 205 may perform the monitoring and / or analysis in an approximate and / or seemingly real-time manner.

[0055] In Figure 1In the examples, and in accordance with one or more aspects of the present disclosure, the policy agent 205 can monitor the server 126. For example, the policy agent 205A of the server 126A can interact with the components, modules, or other elements of the server 126A and / or one or more virtual machines 148 executed on the server 126. As a result of such interaction, the policy agent 205A can collect information about one or more metrics associated with the server 126 and / or the virtual machine 148. Such metrics can be raw metrics that can be read directly based on or directly from the server 126, the virtual machine 148, and / or other components of the data center 110; such metrics can alternatively or additionally be SNMP metrics and / or telemetry-based metrics. In some examples, one or more of such metrics can be computed metrics that include those derived from the raw metrics. In some examples, the metrics can correspond to a percentage of the total capacity related to a particular resource, such as the CPU utilization or CPU consumption or the percentage of the level 3 cache usage. However, the metrics can correspond to other types of metrics, such as the frequency at which one or more virtual machines 148 are reading and writing to the memory.

[0056] The policy agent 205 can perform network infrastructure resource monitoring and analysis locally at each server 126. In some examples, the policy agent 205 accumulates relevant data (e.g., telemetry data) captured by various devices / applications (e.g., sensors, detectors, etc.), packages the captured relevant data into a data model according to the message protocol type, and transmits the packaged relevant data in the message stream. As described herein, some protocol types define dependencies in the data model between the components of two or more adjacent messages, and their corresponding policy or policies require that the entire message stream be streamed to the same collector application for processing (e.g., in sequential order). Other protocol types can define zero (or a trivial amount) of dependencies between the components of two or more adjacent messages, such that the corresponding policy may require multiple collector applications to run in a distributed architecture to process at least one subsequence of messages in the message stream in parallel.

[0057] The policy agent 205 can send the messages in the message stream to a management computing system that acts as a mediator for these messages and any other messages transmitted to the analysis cluster within the example data center 110. In some examples, the analysis cluster can operate substantially similarly to the cluster shown in Figure 1 where each computing system within the analysis cluster is substantially similar to Figure 1Operate on the server 126. Analyze the message transmissions received by the cluster and monitor relevant data related to network infrastructure elements. Based on the relevant data, the collector applications running in the cluster can be analyzed to infer meaningful insights. In response to receiving such a message from one or more network devices of the example data center 110, the management computing system can run a load balancing service to select a computing system of the analysis cluster as the destination server to receive the message, and then transmit the message to the selected computing system. The load balancing service running outside the analysis cluster generally does not consider the protocol type of any message when identifying the computing system to receive the message.

[0058] According to the policy implementation techniques described herein, an ingress engine running in one or more or each computing system of the analysis cluster can receive a message, and by considering the protocol type of the message, can ensure that the message data is processed by an appropriate collector application. The ingress engine can apply each policy requirement corresponding to the message protocol type to each message in the message flow until the message meets each requirement. Satisfaction of the policy may occur when the data of a given message conforms to the data model of the protocol type and is available for processing.

[0059] In one example, the ingress engine accesses various data (such as configuration data) based on the protocol type / data model of the message to identify an appropriate collector application for receiving the message. The appropriate collector application may be compatible with the protocol type / data model of the mail and is assigned to the message flow. The ingress engine that meets the corresponding policy requirements can transmit the message to the appropriate collector application corresponding to the protocol type of the message. This may also involve converting the message into a compatible message, for example, by modifying the message metadata or message payload data. In one example, the ingress engine inserts source information into the message header to ensure that the appropriate collector application can process the message data.

[0060] In addition to identifying appropriate collector applications and delivering messages, the ingress engine can be configured to provide scalability for one or more collector applications via load balancing, high availability, and auto-scaling. In one example, the ingress engine can respond to a failure at or in the computing system of an appropriate collector application by delivering messages in the same message flow to a collector application that runs in a failover node of the same analytics cluster as the computing system of the collector application. The collector application corresponds to the protocol type of the message and is a replica of the appropriate collector application. In another example, the ingress engine of the computing system creates one or more new instances of an appropriate collector application for receiving message data (e.g., telemetry data) of a protocol type. The ingress engine can deliver messages of the same message flow to the new instances running on the same computing system or pass the messages to new instances running on different computing systems of the analytics cluster. In yet another example, if the ingress engine detects an excessive load being processed by an appropriate collector application, the ingress engine can create new instances to handle the excessive load, provided it complies with the corresponding policy.

[0061] Figure 2 is a block diagram of a portion of an example data center 110 that more particularly illustrates one or more aspects of the present disclosure, and in which internal processor metrics related to resources shared by multiple processes executed on an example server 126 are monitored. Figure 1 Shown therein are a user interface device 129 (operated by an administrator 128), a policy controller 201, and a server 126. Figure 2

[0062] The policy controller 201 can represent a collection of tools, systems, devices, and modules that perform operations in accordance with one or more aspects of the present disclosure. The policy controller 201 can execute a cloud service optimization service, which can include advanced monitoring, scheduling, and performance management for software-defined infrastructure, where the lifecycle of containers and virtual machines (VMs) can be much shorter than in traditional development environments. The policy controller 201 can leverage big data analytics and machine learning in a distributed architecture (such as data center 110). The policy controller 201 can provide near or seemingly near real-time and historical monitoring, performance visibility, and dynamic optimization. It can be implemented in a manner consistent with the description of the policy controller 201 provided in conjunction with Figure 1 Figure 2 ​​The policy controller 201. The policy controller 201 may execute a dashboard module 233, which creates, maintains, and / or updates a dashboard 203. The dashboard 203 may include a user interface, which may include a hierarchical network or virtualization infrastructure heat map. A color or range indicator may be presented to infrastructure elements within such a user interface, and the color or range indicator identifies a value range into which one or more utilization metrics associated with each infrastructure element may be classified.

[0063] In Figure 2 as Figure 1 shown, the policy controller 201 includes a policy 202 and a dashboard module 233. The policy 202 and the dashboard 203 may also be implemented in a manner consistent with the description of the policy 202 and the dashboard 203 provided in Figure 1 In Figure 2 the dashboard 203 is created, updated, and / or maintained primarily by the dashboard module 233 executed on the controller 201. In some examples, as Figure 2 shown, the policy 202 may be implemented as a data store. In such an example, the policy 202 may represent any suitable data structure or storage medium for storing the policy 202 and / or information related to the policy 202. The policy 202 may be maintained primarily by a policy control engine 211, and in some examples, the policy 202 may be implemented by a NoSQL database.

[0064] In Figure 2 the example of Figure 2 the policy controller 201 further includes a policy control engine 211, an adapter 207, a message bus 215, reports and notifications 212, an analysis engine 214, a usage metric data store 216, and a data manager 218.

[0065] According to one or more aspects of the present disclosure, the policy control engine 211 may be configured to control the interaction between one or more components of the policy controller 201. For example, the policy control engine 211 may manage the policy 202 and control the adapter 207. The policy control engine 211 may also cause the analysis engine 214 to generate reports and notifications 212 based on data from the usage metric data store 216, and may transmit one or more reports and notifications 212 to a user interface device 129 and / or other systems or components of the data center 110.

[0066] In one example, the policy control engine 211 invokes one or more adapters 207 to discover platform-specific resources and interact with platform-specific resources and / or other cloud computing platforms. For example, one or more adapters 207 may include an OpenStack adapter configured to communicate with an OpenStack cloud operating system running on server 126. One or more adapters 207 may include a Kubernetes adapter configured to communicate with a Kubernetes platform running on server 126. The adapter 207 may also include an Amazon Web Services adapter, a Microsoft Azure adapter, and / or a Google Compute Engine adapter. Such adapters enable the policy controller 201 to learn and map the infrastructure utilized by server 126. The policy controller 201 may use multiple adapters 207 simultaneously.

[0067] One or more reports may be generated over a specified time period, with the reports organized by different scopes: project, host, or department. In some examples, such reports may show the resource utilization of each instance scheduled within a project or on a host. The dashboard 203 may include information presenting the reports in graphical or tabular format. The dashboard 203 may further enable the report data to be downloaded as a report in HTML format, a raw comma-separated value (CSV) file, or data in JSON format for further analysis.

[0068] The data manager 218 and the message bus 215 provide a messaging mechanism for communicating with the policy agent 205 deployed in server 126. The data manager 218 may, for example, publish messages to configure and program the policy agent 205, and may manage metrics and other data received from the policy agent 205, and store some or all of such data in the usage metric data store 216. The data manager 218 may communicate with the policy engine 211 via the message bus 215. The policy engine 211 may subscribe to information (e.g., metric information via a publish / subscribe messaging pattern) by interacting with the data manager 218. In some cases, the policy engine 211 subscribes to information by passing an identifier to the data manager 218 and / or when invoking an API exposed by the data manager 218. In response, the data manager 218 may place the data on the message bus 215 for use by the data manager 218 and / or other components. The policy engine 211 may unsubscribe from receiving data from the data manager via the message bus 215 by interacting with the data manager 218 (e.g., passing an identifier and / or making an API unsubscribe call).

[0069] The data manager 218 can receive, for example, raw metrics from one or more policy agents 205. The data manager 218 can alternatively or additionally receive the results of the analysis performed by the policy agents 205 on the raw metrics. The data manager 218 can alternatively or additionally receive information related to the usage patterns of one or more input / output devices 248, which can be used to classify the one or more input / output devices 248. The data manager 218 can store some or all of such information in the usage metric data store 216.

[0070] In Figure 2 the example of, the server 126 represents a physical computing node that provides an execution environment for virtual hosts such as VM 148. That is, the server 126 includes underlying physical computing hardware 244, and the underlying physical computing hardware 244 includes one or more physical microprocessors 240, memory 249, such as DRAM, power supply 241, one or more input / output devices 248, and one or more storage devices 250. As Figure 2 shown, the physical computing hardware 244 provides an execution environment for the hypervisor 210, which is a software and / or firmware layer that provides a lightweight kernel 209 and operates to provide a virtualized operating environment for virtual machines 148, containers, and / or other types of virtual hosts. The server 126 can represent Figure 1 one of the servers 126 shown (e.g., server 126A to server 126N).

[0071] In the example shown, the processor 240 is an integrated circuit having one or more internal processor cores 243 for executing instructions, one or more internal caches or cache devices 245, a memory controller 246, and an input / output controller 247. Although Figure 2 the example of the server 126 only shows one processor 240, in other examples, the server 126 can include multiple processors 240, and each processor can include multiple processor cores.

[0072] One or more of the devices, modules, storage areas, or other components of server 126 may be interconnected to enable communication (physically, communicatively, and / or operatively) among the components. For example, core 243 may read data from or write data to memory 249 via memory controller 246, which provides a shared interface to memory bus 242. Input / output controller 247 may communicate with one or more input / output devices 248 and / or one or more storage devices 250 via input / output bus 251. In some examples, some aspects of such connections may be provided via communication channels, which include system buses, network connections, interprocess communication data structures, or any other means for conveying data or control signals.

[0073] Within processor 240, each processor core 243A - 243N (collectively referred to as "processor cores 243") provides an independent execution unit to execute instructions that conform to the instruction set architecture of the processor core. Server 126 may include any number of physical processors and any number of internal processor cores 243. Typically, each processor core 243 is combined into a multi-core processor (or "multi-core" processor) using a single IC (i.e., a chip multi-processor).

[0074] In some cases, the physical address space for computer-readable storage media may be shared among one or more processor cores 243 (i.e., shared memory). For example, processor cores 243 may be connected via memory bus 242 to one or more DRAM packages, modules, and / or chips (not shown either), which present a physical address space accessible to processor cores 243. Although this physical address space may provide the lowest memory access time for processor cores 243 in any part of memory 249, processor cores 243 may directly access at least some of the remainder of memory 249.

[0075] Memory controller 246 may include hardware and / or firmware to enable processor cores 243 to communicate with memory 249 via memory bus 242. In the example shown, memory controller 246 is an integrated memory controller and may be physically implemented (e.g., as hardware) on processor 240. However, in other examples, memory controller 246 may be implemented separately or differently and may not be integrated into processor 240.

[0076] The input / output controller 247 can include hardware, software, and / or firmware that enables the processor core 243 to communicate with and / or interact with one or more components connected to the input / output bus 251. In the illustrated example, the input / output controller 247 is an integrated input / output controller and can be physically implemented on the processor 240 (e.g., as hardware). However, in other examples, the memory controller 246 can also be implemented separately and / or in a different manner and may not be integrated into the processor 240.

[0077] The cache 245 represents a memory resource within the processor 240 that is shared among the processor cores 243. In some examples, the cache 245 can include a level 1, level 2, or level 3 cache or a combination thereof and can provide the lowest latency memory access to any storage medium accessible by the processor cores 243. However, in most of the examples described herein, the cache 245 represents a level 3 cache, which, unlike a level 1 cache and / or a level 2 cache, is typically shared among multiple processor cores in a modern multi-core processor chip. However, in accordance with one or more aspects of the present disclosure, in some examples, at least some of the techniques described herein can be applied to other shared resources, including other shared memory spaces other than a three-level cache.

[0078] The power supply 241 provides power to one or more components of the server 126. The power supply 241 typically receives power from a primary alternating current (AC) power supply in a data center, building, or other location. The power supply 241 can be shared among many servers 126 and / or other network devices or infrastructure systems within the data center 110. The power supply 241 can have intelligent power management or consumption capabilities, and such features can be controlled, accessed, or regulated by one or more modules of the server 126 and / or by one or more processor cores 243 to intelligently consume, distribute, supply, or otherwise manage power.

[0079] One or more storage devices 250 can represent computer-readable storage media that include volatile and / or non-volatile, removable and / or non-removable media implemented in any method or technology for storing information such as processor-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), EEPROM, flash memory, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic tape cassette, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by the processor core 243.

[0080] One or more input / output devices 248 can represent any input or output device of server 126. In such an example, the input / output devices 248 can generate, receive, and / or process input from any type of device capable of detecting input from a person or machine. For example, one or more input / output devices 248 can generate, receive, and / or process input in the form of physical, audio, image, and / or visual input (such as a keyboard, microphone, camera). One or more input / output devices 248 can generate, present, and / or process output through any type of device capable of generating output. For example, one or more input / output devices 248 can generate, present, and / or process output in the form of tactile, audio, visual, and / or video output (such as a tactile response, sound, flash, and / or picture). Some devices can act as input devices, some devices can act as output devices, and certain devices can act as both input and output devices.

[0081] Memory 249 includes one or more computer-readable storage media, which can include random access memory (RAM), such as various forms of dynamic RAM (DRAM), such as DDR2 / DDR3 SDRAM or static RAM (SRAM), flash memory, or any other form of fixed or removable storage media that can be used to carry or store the required program code and program data in the form of instructions or data structures and can be accessed by a computer. Memory 249 provides a physical address space composed of addressable memory locations. In some examples, memory 249 can present a non-uniform memory access (NUMA) architecture to processor core 243. That is, processor core 243 may not have equal memory access times to the various storage media that make up memory 249. Processor core 243 can be configured in some instances to use the portion of memory 249 that provides lower memory latency for the core to reduce overall memory latency.

[0082] The kernel 209 can be an operating system kernel that executes in kernel space and can include, for example, a Linux, Berkeley Software Distribution (BSD), or another Unix variant kernel, or a Windows Server operating system kernel, which is available from Microsoft Corporation. Generally, the processor core 243, storage devices (such as the cache 245, memory 249, and / or storage device 250), and the kernel 209 can store instructions and / or data and can provide an operating environment for executing such instructions and / or modules of the server 126. These modules can be implemented as software, but in some examples can include any combination of hardware, firmware, and software. The combination of the processor core 243, storage devices (such as the cache 245, memory 249, and / or storage device 250) within the server 126, and the kernel 209 can retrieve, store, and / or execute the instructions and / or data of one or more applications, modules, or software. The processor core 243 and / or such storage devices can also be operatively coupled to one or more other software and / or hardware components, including but not limited to one or more components of the server 126 and / or one or more devices or systems connected to the server 126.

[0083] The hypervisor 210 is an operating system-level component that executes on the hardware platform 244 to create and run one or more virtual machines 148. In Figure 2 examples, the hypervisor 210 can integrate the functionality of the kernel 209 (such as a "Type 1 hypervisor"). In other examples, the hypervisor 210 can execute on top of the kernel 209 (such as a "Type 2 hypervisor"). In some cases, the hypervisor 210 can be referred to as a Virtual Machine Manager (VMM). Examples of hypervisors include the Kernel-based Virtual Machine (KVM) of the Linux kernel, Xen, ESXi available from VMware, Windows Hyper-V available from Microsoft, and other open-source and proprietary hypervisors.

[0084] In Figure 2 examples, the server 126 includes a virtual router 142 that executes within the hypervisor 210 and can operate in a manner consistent with the description provided in connection with Figure 1 . In Figure 2 examples, the virtual router 142 can manage one or more virtual networks, and each virtual network can provide a network environment to execute virtual machines 148 on the virtualization platform provided by the hypervisor 210. Each virtual machine 148 can be associated with one of the virtual networks.

[0085] The policy agent 205 may execute as part of the hypervisor 210, or may execute within kernel space or as part of the kernel 209. The policy agent 205 may monitor some or all of the performance metrics associated with the server 126. In accordance with the techniques described herein, among other metrics of the server 126, the policy agent 205 is configured to monitor metrics related to or descriptive of the use of resources shared within the processor 240 by each process 151 executing on a processor core 243 within the multi-core processor 240 of the server 126. In some examples, such internal processor metrics relate to the use of the cache 245 (e.g., L3 cache) or the use of bandwidth on the memory bus 242. The policy agent 205 may also be able to generate and maintain a mapping, such as by association with a process identifier (PID) or other information maintained by the kernel 209, that associates the processor metrics of the process 151 to one or more virtual machines 148. In other examples, the policy agent 205 may be able to assist the policy controller 201 in generating and maintaining such a mapping. The policy agent 205 may, under the direction of the policy controller 201, implement one or more policies 202 at the server 126 in response to usage metrics obtained for resources shared within the physical processor 240 and / or further based on other usage metrics of resources external to the processor 240.

[0086] In Figure 2 an example of, a virtual router agent 136 is included in the server 126. Referring to Figure 1 , the virtual router agent 136 may be included in each server 126 (although not shown in Figure 1 )). In Figure 2In the example, the virtual router agent 136 communicates with the SDN controller 132 and, in response thereto, instructs the virtual router 142 to control the overlay of the virtual network and coordinate the routing of data packets within the server 126. Generally, the virtual router agent 136 communicates with the SDN controller 132, which generates commands to control packet routing through the data center 110. The virtual router agent 136 can be executed in user space and act as a proxy for control plane messages between the virtual machine 148 and the SDN controller 132. For example, the virtual machine 148A can request to send a message using its virtual address via the virtual router agent 136, and the virtual router agent 136A can in turn send the message and request to receive a response to the message for the virtual address of the virtual machine 148A, which issued the first message. In some cases, the virtual machine 148A can call a process or function call presented by the application programming interface of the virtual router agent 136, and the virtual router agent 136 also handles the encapsulation of messages, including addressing.

[0087] In some example implementations, the server 126 can include an orchestration agent that communicates directly with the orchestration engine 130 ( Figure 2 (not shown in the figure). For example, in response to an instruction from the orchestration engine 130, the orchestration agent transmits the attributes of a particular virtual machine 148 that is executed on each of the corresponding servers 126 and can create or terminate individual virtual machines.

[0088] The virtual machines 148A, 148B to 148N (collectively referred to as "virtual machines 148") can represent example instances of the virtual machine 148. The server 126 can divide the virtual and / or physical address space provided by the memory 249 and / or the storage device 250 into a user space for running user processes. The server 126 can also divide the virtual and / or physical address space provided by the memory 249 and / or the storage device 250 into a kernel space, which is protected and inaccessible to user processes.

[0089] Generally, each virtual machine 148 can be any type of software application, and each virtual machine can be assigned a virtual address for use within the corresponding virtual network, where each virtual network can be a different virtual subnet provided by the virtual router 142. Each virtual machine 148 can be assigned its own virtual layer 3 (L3) IP address, for example, for sending and receiving communications, but does not know the IP address of the physical server on which the virtual machine is executed. In this way, a "virtual address" is an address for an application that is different from the logical address for the underlying physical computer system, such as Figure 1 the server 126A in the example of

[0090] Each virtual machine 148 may represent a tenant virtual machine that runs customer applications such as web servers, database servers, enterprise applications, or hosts virtualized services for creating service chains. In some cases, any one or more of the servers 126 (see Figure 1 ) or another computing device directly host customer applications, i.e., not as virtual machines. The virtual machines (e.g., virtual machine 148), servers 126, or individual computing devices that host customer applications as referred to herein may alternatively be referred to as "hosts". Additionally, although one or more aspects of the present disclosure are described in terms of virtual machines or virtual hosts, the techniques according to one or more aspects of the present disclosure described herein for such virtual machines or virtual hosts may also apply to containers, applications, processes, or other execution units (virtual or non-virtual) executing on the server 126.

[0091] Processes 151A, 151B to 151N (collectively "processes 151") may be executed in one or more virtual machines 148, respectively. For example, one or more of the processes 151A may correspond to the virtual machine 148A, or may correspond to an application or a thread of an application executing within the virtual machine 148A. Similarly, different sets of processes 151B may correspond to the virtual machine 148B, or to an application or a thread of an application executing within the virtual machine 148B. In some examples, each process 151 may be a thread of an execution unit or other execution unit controlled and / or created by an application associated with one of the virtual machines 148. Each process 151 may be associated with a process identifier that is used by the processor core 243 to identify each process 151 when reporting one or more metrics such as internal processor metrics collected by the policy agent 205.

[0092] In operation, the hypervisor 210 of the server 126 may create multiple processes that share the resources of the server 126. For example, the hypervisor 210 may instantiate or start one or more virtual machines 148 on the server 126 (e.g., under the guidance of the orchestration engine 130). Each virtual machine 148 may execute one or more processes 151, and each of these software processes may execute on one or more of the processor cores 243 of the hardware processor 240 of the server 126. For example, the virtual machine 148A may execute the process 151A, the virtual machine 148B may execute the process 151B, and the virtual machine 148N may execute the process 151N. In Figure 2In the example, processes 151A, 151B, and 151N (collectively referred to as "process 151") are all executed on the same physical host (such as server 126), and can share certain resources when executed on server 126. For example, processes executed on processor core 243 can share memory bus 242, memory 249, input / output device 248, storage device 250, cache 245, memory controller 246, input / output controller 247, and / or other resources.

[0093] Kernel 209 (or hypervisor 210 that implements kernel 209) can schedule processes to be executed on processor core 243. For example, kernel 209 can schedule processes 151 belonging to one or more virtual machines 148 to be executed on processor core 243. One or more processes 151 can be executed on one or more processor cores 243, and kernel 209 can preempt one or more processes 151 periodically to schedule another process 151. Therefore, kernel 209 can perform context switching periodically to start or resume the execution of a different one of processes 151. Kernel 209 can maintain a queue for identifying the next process to be scheduled for execution, and kernel 209 can put the previous process back into the queue for later execution. In some examples, kernel 209 can schedule processes based on polling or other means. When the next process in the queue starts to execute, the next process can access the shared resources used by the previous process, including for example cache 245, memory bus 242, and / or memory 249.

[0094] Figure 3 is a block diagram showing an example analysis cluster 300 according to one or more aspects of the present disclosure, such as Figure 1 data center 110.

[0095] In Figure 3 the example, multiple network devices 302A-N (hereinafter referred to as "network devices 302") are interconnected via network 304. Server 306 is a dedicated server, which is generally used as an intermediate device between network devices 302 and analysis cluster 300, so that any communication (such as a message) is first processed by an interface generated by components of server 306. The message flow initiated by one or more network devices 302 can be directed to this interface, and then server 306 allocates each message flow according to one or more performance management services.

[0096] In server 306, load balancer 310 refers to a collection of software programs that, when executed on various hardware (such as processing circuitry), run one or more performance management services. Load balancer 310 can coordinate various hardware / software components to enable Layer 2 load balancing in analytics cluster 300 (such as Layer 2 performance management microservices); as an example, load balancer 310 and the software layer operating the interface can be configured to distribute message flows according to one or more policies. Load balancer 310 can maintain routing data to identify which of servers 308 receives messages in a given message flow based on the protocol type of the message flow.

[0097] The multiple computing systems of analytics cluster 300 can run various applications to support the monitoring and management of network infrastructure resources. This disclosure refers to these computing systems as multiple servers 308A-C (hereinafter referred to as "servers 308"), where each server 308 can be Figure 1 an example embodiment of server 126. Via server 306 as an entry point, network device 302 can provide relevant data for resource monitoring and network infrastructure management. Various applications running on server 308 consume the relevant data, for example, by applying rules and generating information containing insights into an example data center. These insights provide meaningful information to improve network infrastructure resource monitoring and network management. For example, an application can use the relevant data to evaluate the health or risk of a resource and determine whether the health of the resource is good or bad, or at risk or low risk. Transmitting the relevant data stream to analytics cluster 300 enables network administrators to measure trends in link and node utilization and troubleshoot problems such as network congestion in real time. The relevant data streamed can include telemetry data such as physical interface statistics, firewall filter counter statistics, label switched path (LSP) statistics.

[0098] Although the above-described computing systems and other computing systems described herein are operable to perform resource monitoring and network infrastructure management of an example data center, this disclosure applies to any computing system that must follow policies to determine how to manage incoming messages to ensure that appropriate collector applications receive and process the messages for stored data. Even though the messages involved in this disclosure can be formatted according to any structure, message components can generally be formatted according to networking protocols; for example, some messages according to telemetry protocols store telemetry data for communication with collector applications in destination servers running in server 308.

[0099] In each server 308, multiple collector applications process and store relevant data for subsequent processing. Each collector application can be configured to receive messages according to the protocol types described herein, where each example message can be arranged in a specific message format. Each compatible collector application is operable to convert the message data into a workable data set for monitoring and managing network infrastructure resources. Each compatible collector application can arrange the data to conform to one or more data models compatible with the protocol type (e.g., telemetry protocol type). For example, based on the load balancer 310, server 308A can be configured with a compatible collector application to receive relevant data according to the SNMP, Junos Telemetry Interface (JTI), and SYSLOG protocol types. For each protocol, server 308A can run a single compatible collector application or multiple compatible collector applications. Server 308A can run more refined collector applications; for example, server 308A can run one or more collector applications that are configured to collect data for specific versions of the protocol type, such as separate collector applications for SNMP versions 2 and 3.

[0100] For a given protocol type, the present disclosure describes a set of requirements by which the relevant data stored in a message or sequence of messages is compliant. As described herein, different network devices such as routers, switches, etc. provide different mechanisms (e.g., "telemetry protocols") for collecting relevant data (e.g., performance telemetry data). By way of illustration, SNMP, gRPC, JTI, and SYSLOG are examples of different telemetry protocols for exporting performance telemetry data. Additionally, each telemetry protocol can have one or more unique requirements, such as one or more requirements for controlling the (telemetry) data processing load at one or more computing systems of an analytics cluster. Example requirements can include: maintaining the message load at or below a threshold (i.e., maximum value); adjusting the size and / or rate (e.g., periodically or in bursts) of a complete telemetry data set such that it is transmitted to and / or processed by one or more computing systems of an analytics cluster; regulating the size of a message (e.g., the size of the message payload), etc. There are many other requirements that can affect the load balancing microservices of an analytics cluster. Another requirement may indicate that the protocol is sensitive to message ordering / reordering or packet loss when messages are delivered to a collector application. In a cloud-scale environment, these requirements make it challenging to load balance high-bandwidth telemetry data across multiple collector applications while meeting the requirements of each telemetry protocol. The present disclosure presents techniques for collecting telemetry data using UDP-based telemetry protocols (e.g., SNMP, JTI, and SYSLOG) for further processing and analysis while meeting the specific requirements of different telemetry protocols.

[0101] Examples of such performance telemetry data can be generated by sensors, agents, devices, etc. in data center 110 and then communicated to server 306 in the form of messages. In server 308, various applications analyze the telemetry data stored in each message and obtain statistics or some other meaningful information. Some network infrastructure elements (e.g., routers) in data center 110 can be configured to transmit telemetry data in the form of messages to one or more destination servers.

[0102] Junos Telemetry Interface (JTI) sensors can generate telemetry data (e.g., Link State Path (LSP) traffic data, logical and physical interface traffic data) at the forwarding component (e.g., packet forwarding engine) of network device 302N and transmit probes through the data plane of network device 302N. In addition to connecting the routing engine of network device 302N to server 306, the data port can also be connected to server 306 or a suitable collector application running on one of servers 308. Other network devices 302 can use server 306 to reach a suitable collector application.

[0103] The techniques described herein for message management (e.g., distribution) not only enforce arbitrary policies but also ensure that an appropriate set of requirements is applied to incoming messages. One example set of requirements ensures that the data of an incoming message conforms to the data model described by the message protocol type. In this way, despite having different requirements, efficient resource monitoring and network infrastructure management in an example data center can be achieved by different protocol types (e.g., including different versions) that operate more or less simultaneously.

[0104] Figure 4A is a block diagram showing an example configuration of an example analysis cluster in an example data center 110 according to one or more aspects of the present disclosure. Figure 1

[0105] Similar to Figure 3 analysis cluster 300, analysis cluster 400 provides microservices to support network resource monitoring and network management, such as Figure 1Data center 110. In the exemplary data center 110, multiple network devices 402A-N (hereinafter referred to as "network devices 402") are interconnected via a network 404. The network 404 generally represents various network infrastructure elements and physical carrier media to facilitate communication in the exemplary data center. At least one network device, network device 402N, is communicatively coupled to an analysis cluster 400 via a server 406. The analysis cluster is formed by multiple computing systems on which various applications support network resource monitoring and network management. The present disclosure refers to these computing systems as multiple destination servers 408A-C (hereinafter referred to as "servers 408") for which the server 406 provides load balancing services. Similar to the server 306, the server 406 serves as an entry point through which the network devices 402 can provide a message that conveys relevant data for resource monitoring and network infrastructure management. Many applications running on the servers 408 consume the relevant data, for example, by applying rules and generating information containing insights into the exemplary data center 110.

[0106] Each server 408 includes a corresponding ingress engine 410A-C of the ingress engines (hereinafter referred to as "ingress engines 410") and a corresponding one of the configuration data 412A-C. Multiple collector applications 414A-C (hereinafter referred to as "collector applications 414") running on the servers 408 consume the relevant data, for example, by applying rules and generating information containing insights into the exemplary data center 110. Each ingress engine 410 generally references execution elements that include various computing hardware and / or software, which, in some cases, execute an exemplary configuration stored in the corresponding configuration data 412.

[0107] The exemplary configuration data 412 for any given ingress engine 410 enables the ingress engine 410 to deliver messages of a specific protocol type to one or more appropriate collector applications corresponding to that specific protocol type, where the specific protocol type meets one or more requirements of the data stored in such messages. For message data according to a specific protocol type, exemplary requirements for the viability of the data may include session affinity, proxies for modifying message data, scalability, etc. Given these exemplary requirements, exemplary settings for a specific protocol can identify appropriate collector applications for each message sequence (e.g., flow), attributes to push into the message header, failover nodes in the analysis cluster 400 for high availability, thresholds for load balancing and auto-scaling, etc. The exemplary configuration data 412 can further include initial or default settings for the number of collector applications of each protocol type to be run in the corresponding server of each ingress engine. The exemplary configuration data 412 can identify, for example, which collector application 414 is to process the data provided by the network device based on the protocol type of the data.

[0108] To illustrate by way of example, in server 408A, configuration data 412A stores the configuration settings of that server 408A in view of the corresponding set of requirements in policy 416. Typically, these settings enable ingress engine 410A to correctly manage incoming messages, for example, by maintaining the order of incoming messages, providing scalability (i.e., load balancing, auto-scaling, and high availability), and adding / maintaining source address information 414 when sending messages to one or more appropriate collector applications. Some example settings in configuration data 412A enable ingress engine 410A to identify which collector application 414 in server 408A is suitable for receiving messages according to a specific protocol type. Other settings indicate changes to specific message attributes, for example, in order to maintain the correct order of the message sequence. As described herein, by having multiple messages in the correct order, collector applications 414 and other applications that support network resource monitoring and network management can assemble a complete data set of the relevant data transmitted in these messages.

[0109] As described herein, policy 416 includes policy information for each protocol type or message type, and this policy information can be used to provide configuration data 412 for ingress engine 410 to server 408. With respect to server 408A, configuration data 412A can identify collector application 414A for messages according to SNMP, collector application 414B for messages according to JTI, and collector application 414C for messages according to SYSLOG. It should be noted that configuration data 412A can specify additional collector applications 414, which include additional collectors for SNMP, SYSLOG, and JTI messages and one or more collectors for other protocol types.

[0110] Server 406 runs load balancer 418 as an execution element that includes various hardware and / or software. In one example, load balancer 418 executes multiple mechanisms to create and then update routing data 420 with learned routes to direct messages to their destination servers in server 408. Load balancer 418 can store in routing data 420 the routes optimized for performance by server 408. Routing data 420 can indicate which of the servers 408 will receive a particular sequence (e.g., stream) of messages, such as messages generated by the same network device or associated with the same source network address; these indications do not represent requirements for the relevant data, and if followed, could prevent such data from becoming useless. Instead, each ingress engine 410 is configured to enforce these requirements on the relevant message data based on the message protocol type.

[0111] For illustration, given routing data 420, server 406 may deliver messages of a message flow to a particular server 408, and the corresponding ingress engine 410 may in turn deliver those messages to an appropriate collector application corresponding to the protocol type of the message, as required by the protocol type. The appropriate collector application may represent the same collector application for processing other messages of the same message flow, which may or may not be running on the particular server 408. Thus, even if a compatible collector application is running in the particular server 408, the corresponding ingress engine 410 may deliver messages to an appropriate collector application in a different server 408.

[0112] Figure 4B is a block diagram showing an example routing instance in an Figure 4A analytics cluster 400 according to one or more aspects of the present disclosure. The example routing instance is derived from the path traversed when the ingress engine 410 executes the corresponding configuration data 412 to deliver a message to an appropriate collector application 414. Thus, each example routing instance refers to a set of nodes forming a path that delivers each message belonging to a particular message protocol type to one or more destination servers in server 408.

[0113] In Figure 4B shown by a solid line, the first routing instance represents the path of messages arranged according to the first protocol type. A policy 416 that stores one or more requirements for message data according to the first protocol type may include an attribute indicating a high tolerance for reordering. This attribute may require multiple collector applications to process the message data rather than a single dedicated collector application. Due to the high tolerance attribute, there is no need for multiple collector applications to maintain the message order (e.g., by coordinating message processing), and they can freely process messages in any order without rendering the associated data useless.

[0114] Thus, when a series of messages carrying related telemetry data arrive at server 406, these messages are delivered to server 408A, where the ingress engine 410A identifies multiple collector applications 414A to receive / consume the related data (e.g., in any order). The messages may be distributed among multiple instances of the collector application 414A, where each instance receives a different portion of the related data for processing (e.g., in parallel, sequentially, or in any order). An example distribution of messages by the ingress engine 410A may be to deliver every ith message in a subsequence of n messages to the ith collector application 414A running in server 408. In Figure 4AIn this case, the ingress engine 410 delivers the first, second, and third messages in every three-message subsequence to the corresponding collector application 414A in servers 408A, 408B, and 408C. In other examples, copies of each message are distributed across multiple instances so that even if one copy is lost or some packets are dropped, the collector application 414A can still consume the relevant data and extract meaningful information to support network resource monitoring and management.

[0115] Any collector application 414 running in server 408 may fail, and messages (failover) need to be migrated to a different collector application 414 and / or a different server 408. To promote high availability in analytics cluster 400, the ingress engine 410A may create a new instance of the collector application 414A (i.e., the failover collector application 414A) in the same server 408A or in a different server 408B or 408C (i.e., the failover node). In response to a failure at server 408A, server 406 delivers the message to server 408B or 408C, where the corresponding ingress engine 410B or 410C further delivers the message to the failover collector application 414A. In response to a failure of the collector application 414A running at server 408A, the ingress engine 410A may deliver the message to the failover collector application 414A and resume message processing. Thus, high availability can be achieved by having another server running the same collector application, so that if the first server fails, the message traffic can failover to another server running the same type of collector application.

[0116] Maintaining failover nodes and applications provides many benefits. As an example, in the event of a failure, at least one instance of a compatible collector application is running on another server 408 in analytics cluster 400. Another benefit is that if a message is lost or message data is corrupted, message processing will be interrupted very little or not at all. Thus, the failure does not interrupt analytics cluster 400. Messages arranged according to the first protocol type may arrive out of order, be corrupted, be lost, and / or be completely discarded during transmission, and in most examples, the ingress engine of server 408 should be able to obtain meaningful information even from an incomplete set of data.

[0117] The second routing instance (in Figure 4BThe one shown with a dashed line (in the middle) results from the implementation of Policy 416 for messages arranged according to the second protocol type. Messages of the second protocol type are processed sequentially by the same collector application. According to Policy 416 for the second protocol type, one or more requirements must be followed because these messages have a low tolerance for reordering, otherwise meaningful information cannot be obtained or learned. An example requirement might be that only one collector application can process messages in message order. The second protocol type may conform to the UDP transport protocol, which does not use sequence numbers to maintain message order. Examples of the second protocol type include UDP-based telemetry protocols (such as JTI) where session infinity is required. The ingress engine 410 is configured to direct these messages to a specific instance of the collector application 414B, such as an instance running on server 408A. Because the analytics server 400 is dedicated to a specific instance of the collector application 414B to process messages in a given flow (such as the same source address, etc.) of the second protocol type. If any ingress engine other than the ingress engine 410A receives this message, that ingress engine (such as ingress engine 410B or 410C) will immediately redirect the message to the appropriate collector application, as Figure 4B shown.

[0118] The configuration data 412B can provide an identifier, a memory address, or another piece of information that describes how to access a specific instance of the collector application 414B running on server 408A for the second protocol type. Due to the configuration data 412B, the ingress engine 410B of server 408B can be fully programmed to have instructions to appropriately redirect messages according to the second protocol type. For example, the ingress engine 410B can redirect any message determined to be arranged according to the second protocol type to server 408A. Figure 4B This redirection is shown using a dashed line. Although in server 408B, another instance of the collector application 414B may be running (as Figure 4B shown), Policy 416 can still identify server 408A and / or the compatible collector application 414B running in server 408A to receive messages from server 406.

[0119] The ingress engine 410 generally applies the appropriate policies for incoming messages to determine which collector application or applications are appropriate. If no local appropriate collector application is available, the ingress engine 410 determines which servers 408 run the appropriate collector application. Additionally, the corresponding policy can specify one or more scalability requirements. For example, if a portion of the message flow according to a second protocol type is sent to the collector application 414B running in server 408A, the policy for the second protocol type may require the same instance of the collector application 414B to receive one or more subsequent messages of the same message flow. Once an ordered set of messages is assembled at server 408A, at some point, a sufficient amount of relevant data can be collected from the data stored in the ordered set of messages. If that amount exceeds a threshold, a new instance of the collector application 414B can be created to handle such overflow. In the event of a failure of the collector application 414B in server 408A, the ingress engine 410A migrates the message processing to a failover instance of the collector application 414B. Before migration, the ingress engine 410A can create a failover instance of the collector application 414B in server 408A or another server 408. If a message is lost from the sorted set for some reason, information cannot be extracted from the data stored in the message.

[0120] The load balancer 418 can be configured to maintain session affinity by sending messages of the same message flow to the same server 408 in which an instance of the appropriate collector application is running. The load balancer 418 stores the mapping between the same server 408 and the message flow in the routing data 420 for the server 406 to follow. Even if the appropriate collector application has migrated (e.g., due to a failure) to another server 408, the load balancer 418 can continue to send messages of the same flow to the same server 408. This can occur after the collector application 414B has been migrated to server 408B and when the load balancer 418 selects the same destination server (e.g., server 408B) for the same message flow arranged according to the second protocol type. To prevent message loss, the corresponding ingress engine 410B of server 408B can redirect the message to the collector application 414B in server 408A.

[0121] In Figure 4BThe third routing instance is shown by a dashed line, where the dashed line represents the path of messages according to the third protocol type. For this routing instance, policy 416 can establish proxies at each ingress engine to insert metadata into each message. As a requirement for any relevant data in these messages, policy 416 specifies one or more collector applications 414C, but leaves it up to the ingress engine 410C to choose how many collector applications 414C to keep open for message processing. An example policy 416 requirement can be to identify the collector application 414C as an appropriate collector, but at the same time, grant the ingress engine 410C the permission to select or create a duplicate collector application 414C to process the messages. The ingress engine 410C can distribute these messages (e.g., in an ordered sequence) among the instances of the collector application 414C running in the server 408, including an instance in 408C and another instance in 408C.

[0122] Alternatively, the ingress engine 410C can distribute the messages to multiple replicas of the collector application 414C running in the server 408C. Even though Figure 4B the server 408C is shown as having one replica of the collector application 414C, other examples of the server 408C may have multiple replicas, and other servers 408 may have other replicas. Any number of collector applications 414C can process the relevant data in these messages; by redistributing these messages among themselves, a complete data set of the relevant data in these messages can be assembled and processed.

[0123] Figure 4C is a block diagram showing an example failover in the analysis cluster 400 according to one or more aspects of the present disclosure. Generally speaking, an example failover refers to a state (e.g., an operating state) in which the analysis cluster 400 identifies a non-functional destination server and utilizes resources on one or more other destination servers to process messages originally destined for the non-functional destination server. Figure 4A

[0124] Figure 4C ​​As shown, the server 408A experiences a computing failure and becomes inoperable; thus, if a message is sent to the server 408A, the message will not reach and may be completely lost. The load balancer 418 is configured to ensure high availability of the analytics cluster 400 by redirecting these messages to another destination server in the server 408, thereby ensuring that these messages are processed for the relevant data. To replace the collector applications previously running in the server 408A, the ingress engines 410B and 410C can instantiate initial copies of these applications (if not already running) or instantiate other copies of these applications in other servers 408. For example, messages of the first protocol type (e.g., SYSLOG protocol type) can be redirected from the server 408A to the servers 408B and 408C, and each server has installed and is running a copy of the collector application 414A.

[0125] Messages according to the second protocol type (e.g., a telemetry protocol such as JTI) can be redirected after a failure, but doing so may risk disrupting telemetry data analysis and losing some insights into network infrastructure monitoring and management. The policy 416 for the second protocol type requires these messages to be processed in sequence, or at least certain relevant telemetry data cannot be extracted for operational statistics and other meaningful information. Thus, when the failure interrupts the server 408A, the server 408B can operate as a failover node, and a new instance (e.g., a copy / a single copy) of the collector application 414B can receive these messages and continue telemetry data analysis. However, in some cases, telemetry data analysis may be interrupted because the protocol type of the message requires the same instance of the collector application 414B to receive the message. For some second protocol types (e.g., JTI), the new instance of the collector application 414B can expect minimal message loss during failover until a steady state is reached. There may be other cases where the new instance of the collector application 414B currently running on the server 408B cannot respond to the failover on the server 408A and process incoming messages.

[0126] It should be noted that the present disclosure does not exclude providing high availability for messages of the second protocol type. In some examples, even if the second protocol type requires sequential message processing, the collector application 414B can be migrated from the server 408A to the server 408B. Even though the policy 416 indicates that an ordered sequence of messages needs to be maintained, the ingress engine 410B instantiates a second copy of the collector application 414B in the server 408B to process these messages. Although Figure 4Cis not shown, but multiple copies (e.g., replicas) of the collector application 414B can be instantiated in the server 408B. This can occur even if a first copy of the collector application 414B is already running in the server 408B. In one example, the first copy of the collector application 414B does not have any previously transmitted messages and, at least for this reason, cannot be used for incoming messages. In this way, the first copy of the collector application 414B is not burdened with excessive processing tasks, and data collection of messages for this first copy is not interrupted. On the other hand, if the policy 416 requires messages to arrive in sequence for processing at the same collector application 414B, a second copy of the collector application 414B can be specified for this purpose after resetting or restarting the message sequence. Some protocol types can meet this requirement while allowing, for example, new instances to assume responsibility during failover.

[0127] Figure 4D illustrates an example auto-scaling operation in the analysis cluster 400 in accordance with one or more aspects of the present disclosure Figure 4A block diagram.

[0128] In some examples, the policy 416 indicates to any ingress engine 410 in the server 408 how or when to perform load balancing. For messages of the SYSLOG protocol type, the ingress engine 410 can load balance these messages among all copies of the SYSLOG-compatible collector application running in the server 408. For messages of the SNMP TRAP (e.g., version 2 or 3) protocol type, the ingress engine 410 can load balance these messages among all copies of the SNMP TRAP-compatible collector application. For messages of SNMP TRAP version 2, the ingress engine 410 can preserve the source IP address, or otherwise the source IP address can be overwritten by the ingress engine 410 to achieve scalability across replicas. The ingress engine 410 with an SNMP agent can ingest the source IP address to preserve this source information until the replica receives the UDP message. For JTI messages, neither the load balancer 418 nor the ingress engine 410 provides scalability because the policy 416 for JTI poses a requirement that all incoming JTI-conforming messages from the same source IP address be routed to the same copy of the appropriate collector application. To maintain session affinity for JTI messages, the ingress engine 410 is excluded from routing these messages to other replicas, even if that replica is also a compatible instance of the same collector application.

[0129] For example, SYSLOG messages storing telemetry data are transmitted to at least server 408A for processing by collector application 414A. In response to receiving the SYSLOG message, ingress engine 410A of server 408A may create a new instance of collector application 414A (e.g., replica collector application 414A') for receiving telemetry data stored in messages of the SYSLOG protocol type. In addition to or instead of transmitting the message to a previous instance of collector application 414A, ingress engine 410A may transmit the message to new instance replica 414', and / or transmit the message to a new instance of collector application 414A running on a different computing system.

[0130] In some examples, in response to determining an overload at one or more of servers 408, the corresponding ingress engine 410 may analyze network traffic and, based on that analysis, identify one or more new collector applications to instantiate. The corresponding ingress engine 410 may determine which protocol types are being used to format most of the messages in the network traffic and then instantiate copies of the appropriate collector applications to process those messages. If analysis cluster 400 mostly receives JTI messages, the corresponding ingress engine 410 creates multiple new instances of appropriate collector application 414B in multiple destination servers 408. In other examples, policy 416 indicates when the corresponding ingress engine 410 performs load balancing among multiple destination servers 408.

[0131] Figure 5 is a flowchart showing example operations of computing systems in an example analysis cluster in accordance with one or more aspects of the present disclosure. The following describes Figures 4A - 4D in the context of server 406 of Figure 4B which server itself is an example of server 126 of Figure 5 and is a computing system that can be deployed in analysis cluster 400 of Figure 1 -D. In other examples, Figure 4A the operations described in Figure 5 may be performed by one or more other components, modules, systems, or devices. Additionally, in other examples, the operations described in conjunction with Figure 5 may be combined, performed in a different order, or omitted.

[0132] In Figure 5In the example of, and in accordance with one or more aspects of the present disclosure, the processing circuitry of server 406 accesses policy information (e.g., policy 416) and configures each ingress engine 410(500) running in server 408 of analytics cluster 400 based on the policy information. Based on the policy information for analytics cluster 400, collector applications are distributed among the servers 408 of analytics cluster 400 and are configured to receive messages from network devices of example data center 110 and then process the data stored in those messages. Messages transmitted from network devices arrive at server 406, and the server may or may not route those messages to an appropriate destination server running a compatible collector application to process the data stored in those messages. The policy information includes one or more requirements for data stored in messages to be considered valid and / or sufficient for analytics purposes.

[0133] The policy information in any of the example policies described herein generally includes one or more requirements for data stored in messages corresponding to a particular networking protocol type. One example requirement in the policy information determines one or more servers 408 to receive messages according to the protocol type based on the protocol type of the message data for storage. Demonstrating compliance with each requirement ensures that the stored message data is a viable data set for network resource monitoring and management. Some requirements ensure compliance with a particular network protocol type of the message, which may be of the type of UDP-based telemetry protocols (e.g., JTI, SYSLOG, SNMP TRAP versions 2 and 3, etc.), or of another category of networking protocol types. A particular network protocol may belong to a group of protocols, such as a group including any protocol type without a reordering mechanism. Since messages can correspond to multiple network protocol types, some example policies define two or more sets of requirements for different networking protocols, such as a set of requirements for ensuring compliance with telemetry protocols and a second set of requirements for ensuring compliance with transport protocols. As another example, based on the protocol type, the policy information identifies which collector application is suitable for receiving certain messages in the event of a migration event (such as a failover or load balancing). The protocol type can also establish an auto-scaling threshold for determining when to create a new instance of an appropriate collector application or destroy an old instance.

[0134] As described herein, the analytics cluster 400 can provide microservices such as for a data center 110, and some of the microservices involve implementing example policies to direct reordering of messages for network resource monitoring and network management. A policy controller 201 can define one or more policies, such as for the data center 110, that include the policies regarding message reordering described above. For example, a user interface device 129 can detect an input and output an indication of the input to the policy controller 201. A policy control engine 211 of the policy controller 201 can determine that the input corresponds to information sufficient to define one or more policies. The policy control engine 211 can define one or more policies and store the one or more policies in a policy data store 202. The policy controller 201 can deploy the one or more policies to one or more policy agents 205 executing on one or more servers 126. In one example, the policy controller 201 deploys one or more example policies to a policy agent 205 executing on a server 126 (such as server 406).

[0135] As described herein, a server 406 can receive a message and, based on the message's header and (possibly) other attributes, determine a destination server corresponding to the flow of the message. Processing circuitry of the server 406 identifies one of the servers 408 as the destination server based on routing data 420 and then transmits the message to the identified destination server 408. In turn, a corresponding ingress engine 410 runs on the processing circuitry of the identified destination server 408. The received message is received and the protocol type (502) is determined.

[0136] The corresponding ingress engine 410 of the identified destination server 408 identifies one or more appropriate collector applications (504) for processing the message. The identified destination server 408 can identify the application based on policy information for the protocol type. If the policy information requires the same collector application to receive an ordered set of messages, the corresponding ingress engine 410 of the identified destination server 408 can pass the message to an instance of a collector application running on the identified destination server 408 or a different second destination server 408.

[0137] If the policy information indicates a high tolerance for reordering and does not require messages to arrive at the same collector application, the processing circuitry of the server 406 can load balance messages of the same protocol type and from the same flow (i.e., having the same source address) across multiple collector applications running across the servers 408.

[0138] Alternatively, the ingress engine 410 of server 408 may modify the message data (506). Such modifications may or may not conform to the policy information of the protocol type. SNMP TRAP version 2, an example protocol type, may accept reordering of packets; however, the source IP address (e.g., packet header) in each packet needs to be preserved or it may be overwritten by the ingress engine 410. When the ingress engine 410 receives SNMP TRAP version 2 packetized data to preserve the source IP address, the ingress engine 410 inserts that address into each packet (e.g., UDP packet) and delivers those packets to the appropriate collector application.

[0139] The ingress engine 410 of server 408 delivers the message to one or more appropriate collector applications (508). Each appropriate collector application corresponding to the protocol type of the message conforms to at least one requirement of the policy information. As described herein, in response to receiving a message from a network device in data center 110, the ingress engine 410 may determine whether the message conforms to at least one requirement of the data stored in the message. If the policy information indicates a high tolerance for reordering of the protocol type of the message, the ingress engine 410 delivers the message to a particular one of the multiple appropriate collector applications. Based on the policy information, the ingress engine 410 distributes the message among the multiple appropriate collector applications such that each application processes different messages of the same message stream. In some examples, the ingress engine 410 delivers the message to the same instance of the collector application as other messages of the same protocol type and message stream.

[0140] For the processes, apparatuses, and other examples or illustrations described herein, including in any flowcharts or flow diagrams, certain operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be entirely added, combined, or omitted (e.g., not all described acts or events are necessary for implementing the technique). Additionally, in some examples, operations, actions, steps, or events may be performed in parallel rather than sequentially, such as by multi-threading, interrupt processing, or multiple processors. Certain other operations, actions, steps, or events may be automated even if not explicitly identified as such. Additionally, certain operations, actions, steps, or events described as automated may alternatively not be automated but, in some examples, such operations, actions, steps, or events may be performed in response to an input or another event.

[0141] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on and / or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, the computer-readable medium may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures to implement the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0142] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the medium of the definition includes coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. Disk and optical disks used herein include optical disks (CD), laser disks, optical disks, digital versatile disks (DVD), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, while optical disks reproduce data optically by laser. Combinations of the above should also be included within the scope of computer-readable media.

[0143] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the terms “processor” or “processing circuitry” as used herein may refer to any of the foregoing structures or any other structure suitable for implementation of the described techniques. Additionally, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Similarly, the technique may be fully implemented in one or more circuits or logic elements.

[0144] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or IC collections (such as chip sets). In the present disclosure, various components, modules or units are described to emphasize the functional aspects of the devices configured to execute the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units can be combined in a hardware unit or provided by a collection of interoperating hardware units, including one or more processors as described above in conjunction with suitable software and / or firmware.

Claims

1. A method for processing messages, comprising: receiving, by an ingress engine running on a processing circuit of a computing system, a plurality of messages from a network device in a data center, the data center including a plurality of network devices and the computing system, wherein the plurality of messages includes data configured according to a specific protocol type; a policy defined by a server communicatively coupled to the computing system, the policy including: a first attribute indicating a high tolerance for reordering of messages containing data configured according to a first protocol type, and a second attribute indicating a low tolerance for reordering of messages containing data configured according to a second protocol type; and in response to receiving the plurality of messages from the network device of the data center, transmitting, by the ingress engine, the plurality of messages to: a plurality of appropriate collector applications when the specific protocol type is the first protocol type; or a specific appropriate collector application when the specific protocol type is the second protocol type.

2. The method according to claim 1, wherein transmitting the plurality of messages by the ingress engine of the computing system further comprises: Based on at least one requirement, transmitting, by the ingress engine, the plurality of messages to a collector application running in a second computing system.

3. The method according to claim 1, wherein transmitting the plurality of messages by the ingress engine of the computing system further comprises: Transmitting the plurality of messages to a collector application running in the computing system and identical to other messages having the same protocol type or to a collector application running in a second computing system and identical to other messages having the same protocol type.

4. The method according to claim 1, wherein transmitting the plurality of messages by the ingress engine further comprises: Distributing, by the ingress engine, a message sequence to a plurality of collector applications running in a plurality of computing systems of an analysis cluster, the plurality of collector applications being compatible with the protocol type of the messages.

5. The method according to claim 1, wherein transmitting the plurality of messages by the ingress engine further comprises: Transmitting, by the ingress engine, the plurality of messages to a collector application running in a failover node of a cluster identical to the computing system, wherein the collector application corresponds to the protocol type of the messages and is a copy of the appropriate collector application.

6. The method according to claim 1, wherein transmitting the plurality of messages by the ingress engine of the computing system further comprises: Identifying a collector application to process stored data according to the telemetry protocol type of the plurality of messages, the stored data provided by the network device.

7. The method according to any one of claims 1 to 6, further comprising: creating, by the ingress engine, a new instance of the appropriate collector application for receiving telemetry data stored in messages of the specific protocol type; and at least one of: transmitting, by the ingress engine of the computing system, the plurality of messages to the new instance running on the computing system, or transmitting, by the ingress engine of the computing system, the plurality of messages to the new instance running on a different computing system.

8. The method according to any one of claims 1 to 6, wherein transmitting the plurality of messages by the ingress engine of the computing system further comprises: Transmitting, by a load balancer of a server communicatively coupled to the computing system, the plurality of messages to a second ingress engine running on a second computing system, wherein the second ingress engine redirects the plurality of messages to the ingress engine running in the computing system.

9. The method according to any one of claims 1-6, wherein transmitting the plurality of messages by the ingress engine of the computing system further comprises: Modifying a message header before transmitting the plurality of messages to the appropriate collector application.

10. The method according to any one of claims 1 to 6, wherein transmitting the plurality of messages by the ingress engine further comprises transmitting the plurality of messages by the ingress engine to an appropriate collector application corresponding to the protocol type of the messages based on a determination of how to manage the messages based on at least one requirement.

11. The method according to any one of claims 1 to 6, further comprising: Configuration data is generated by processing circuitry of a server communicatively coupled to the computing system, based on the policy for the protocol type, wherein the policy indicates at least one requirement for a set of viable data for telemetry data stored in the messages, and wherein the configuration data identifies one or more appropriate collector applications for receiving the messages.

12. The method according to any one of claims 1-6, wherein transmitting the plurality of messages by the ingress engine further comprises: The ingress engine determines whether the plurality of messages comply with at least one requirement for data stored in the messages by accessing configuration data, the configuration data including the number of collector applications for the ingress engine running in the computing system, and the at least one requirement for data stored in messages transmitted from one or more network devices to the computing system, wherein the at least one requirement identifies the appropriate collector application corresponding to the protocol type of the messages.

13. A computing system, comprising: a memory; and processing circuitry communicative with the memory, the processing circuitry being configured to execute logic including an ingress engine, the ingress engine being configured to: receive a plurality of messages from network devices in a data center, the data center including a plurality of network devices and the computing system, wherein the plurality of messages include data configured according to a particular protocol type; identify a policy defined by a server communicatively coupled to the computing system, the policy including: a first attribute indicating a high tolerance for reordering of messages containing data configured according to a first protocol type, and a second attribute indicating a low tolerance for reordering of messages containing data configured according to a second protocol type; and in response to receiving the plurality of messages from the network devices in the data center, transmit the plurality of messages to: a plurality of appropriate collector applications when the particular protocol type is the first protocol type; a particular appropriate collector application when the particular protocol type is the second protocol type.

14. The computing system according to claim 13, wherein the ingress engine distributes the messages of the plurality of messages to a plurality of collector applications corresponding to the protocol type of the messages, the plurality of collector applications running in a plurality of computing systems in a cluster.

15. The computing system according to claim 13, wherein the ingress engine uses the configuration data to identify an instance of the appropriate collector application based on the protocol type of the messages, the instance running in at least one of the computing system or a different computing system.

16. The computing system according to any one of claims 13 - 15, wherein in response to a failure, the ingress engine redirects the plurality of messages to a second appropriate collector application running in a failover node communicatively coupled to the computing system.

17. The computing system according to any one of claims 13 - 15, wherein the ingress engine modifies the plurality of messages to comply with one or more requirements.

18. The computing system according to any one of claims 13 - 15, wherein the computing system operates as a failover node for at least one second computing system in the same cluster, and at least one of a server communicatively coupled to the computing system or a second ingress engine running in the second computing system transmits the plurality of messages to the computing system in response to a failure at the appropriate collector application.

19. A computer - readable storage medium encoded with instructions for causing a programmable processor to perform the method according to any one of claims 1 - 12.

Citation Information

Patent Citations

  • Packet segmentation offload for virtual networks

    US9641435B1

  • Packet distributing system and method for distributing access packets to a plurality of server apparatuses

    US20030195919A1

  • Adaptable network event monitoring configuration in datacenters

    US20180063195A1