Network device, method for flow monitoring, and computer-readable storage medium
Through network equipment that adaptively adjusts the sampling rate, the problem of waste and incomplete monitoring of flow monitoring resources in virtualized data centers is solved, and efficient and economical network traffic monitoring and management is achieved.
Patent Information
- Application Number
- CN202210845876.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-24
- Filing Date
- 2022-07-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-19
AI Technical Summary
In virtualized data centers, it is difficult for the prior art to adaptively adjust the sampling rate of stream monitoring, resulting in waste of resources or incomplete monitoring.
The sampling rate is adaptively adjusted by network equipment, and dynamically adjusted the sampling rate based on changes in flow parameters to optimize flow monitoring, avoid resource overload and ensure monitoring integrity.
It realizes efficient and economical monitoring of network traffic in virtualized data centers, reduces resource waste, and improves the degree of automation and troubleshooting of network management.
Smart Images

Figure CN115914012B_ABST
Abstract
Description
[0001] This application claims the benefit of U.S. Patent Application No. 17 / 410,887, filed on August 24, 2021, the entire content of which is incorporated herein by reference. Technical Field
[0002] This disclosure relates to the analysis of computer networks. Background Art
[0003] Virtualized data centers are becoming a core foundation of modern information technology (IT) infrastructure. Modern data centers widely utilize virtualized environments in which virtual hosts such as virtual machines or containers are deployed and run on underlying computing platforms of physical computing devices.
[0004] Virtualization within large-scale data centers can provide several advantages, including efficient use of computing resources and simplified network configuration. Thus, in addition to the efficiency and increased return on investment (ROI) provided by virtualization, enterprise IT personnel typically prefer virtualized computing clusters in data centers due to their management advantages. However, virtualization can pose some challenges when analyzing, evaluating, and / or troubleshooting network operations. Summary of the Invention
[0005] This disclosure describes techniques for adaptive flow monitoring. Flow monitoring includes the process of monitoring traffic flows within a network. Flow monitoring can enable network administrators to gain a better understanding of the networks they manage, automate specific network management tasks, and / or perform other activities.
[0006] Ideally, one skilled in the art could sample each packet, but sampling each packet can be expensive to implement, can burden processing resources, and can additionally increase the footprint of network equipment. Thus, techniques for sampling flows such as sampled flow (sFlow) have been developed to sample flows at a given sampling rate. Currently, a network administrator can set this sampling rate in a flow collector for a given interface of another network device, such as a top-of-rack (ToR) switch or other network device. However, flows change over time and a manually set sampling rate may become ineffective or may also burden processing resources. For example, a manually set sampling rate may not sample all flows of an interface quickly enough or may be too fast for the processing circuitry to handle. Thus, it may be desirable to sample flows adaptively.
[0007] According to the technology of the present disclosure, a network device may change the sampling rate of a flow from an interface based on a change in the sampling parameter of the flow or a lack of change (or lack of significant change) in the sampling parameter of the flow. For example, if changing to a low sampling rate results in monitoring approximately the same number of flows as a high sampling rate, the low sampling rate may be better than the high sampling rate because the high sampling rate may be considered a waste of processing resources. On the other hand, if changing to a low sampling rate results in monitoring a significantly lower number of flows than a high sampling rate, the high sampling rate may be better than the low sampling rate because the low sampling rate does not allow all flows to be monitored.
[0008] In one example, the present disclosure describes a method including: receiving, from an interface of a network device, a first sample of a flow sampled at a first sampling rate; determining a first parameter based on the first sample; receiving, from the interface, a second sample of the flow sampled at a second sampling rate, where the second sampling rate is different from the first sampling rate; determining a second parameter based on the second sample; determining a third sampling rate based on the first parameter and the second parameter; sending a signal indicating the third sampling rate to the network device; and receiving, from the interface, a third sample of the flow sampled at the third sampling rate.
[0009] In another example, the present disclosure describes a network device including: a memory configured to store a plurality of sampling rates; a communication unit configured to send signals and receive samples of a data stream; and a processing circuit communicatively coupled to the memory and the communication unit, the processing circuit being configured to: receive, from an interface of another network device, a first sample of a flow sampled at a first sampling rate; determine a first parameter based on the first sample; receive, from the interface, a second sample of the flow sampled at a second sampling rate, where the second sampling rate is different from the first sampling rate; determine a second parameter based on the second sample; determine a third sampling rate based on the first parameter and the second parameter; control the communication unit to send a signal indicating the third sampling rate to another network device; and receive, from the interface, a third sample of the flow sampled at the third sampling rate.
[0010] In another example, the present disclosure describes a computer-readable medium including instructions for causing a programmable processor to: receive, from an interface, a first sample of a flow sampled at a first sampling rate; determine a first parameter based on the first sample; receive, from the interface, a second sample of the flow at a second sampling rate, where the second sampling rate is different from the first sampling rate; determine a second parameter based on the second sample; determine a third sampling rate based on the first parameter and the second parameter; control the communication unit to send a signal indicating the third sampling rate to the network device; and receive, from the interface, a third sample of the flow sampled at the third sampling rate.
[0011] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, the drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1A is a conceptual diagram of an exemplary network including a system for analyzing traffic flows between networks and / or within a data center in accordance with one or more aspects of the present disclosure.
[0013] Figure 1B is a conceptual diagram of an exemplary component of a system for analyzing traffic flows between networks and / or within a data center in accordance with one or more aspects of the present disclosure.
[0014] Figure 2 is a block diagram of an exemplary network analysis system in accordance with one or more aspects of the present disclosure.
[0015] Figure 3 is a block diagram of an exemplary network device in accordance with one or more aspects of the present disclosure.
[0016] Figure 4 is a flowchart of an exemplary method in accordance with one or more techniques of the present disclosure. DETAILED DESCRIPTION
[0017] Data centers using virtual environments provide efficiency, cost, and organizational advantages, where virtual hosts such as virtual machines or containers are deployed and run on the underlying computing platforms of physical computing devices. However, in managing any data center fabric, it is undoubtedly important to obtain meaningful insights into application workloads. Collecting flow datagrams (which may include traffic samples) from network devices can help provide such insights.
[0018] Sampling each packet can be prohibitively expensive and may prevent the network device from performing the network device's primary functions, such as routing, switching, or processing packets. Sampling of packets on an interface can be performed at a sampling rate that provides a balance between the expense of sampling each packet and allowing the network device to focus its processing capabilities on its primary purpose. However, a statically provided sampling rate may become obsolete as the flow changes. Thus, it may be desirable to provide adaptive flow monitoring.
[0019] According to one or more aspects of the present disclosure, a network device may adapt the sampling rate of an interface based on changes in flow parameters, such as changes in the number of flows determined at different sampling rates. In the various examples described herein, the network device may recursively receive samples of flows sampled at different sampling rates from an interface of another network device and determine the parameters of the flow until the last determined parameter is different or substantially different from the parameter determined immediately before. In some examples, the determined parameter is the number of flows. In some examples, the different sampling rates decrease step by step, each sampling rate being lower than the previous sampling rate. In some examples, once the network device determines that the last determined parameter is different or substantially different from the parameter determined immediately before, the network device may instruct another network device to sample at a sampling rate higher than the last sampling rate. Samples of the flows may provide insights into the network and provide tools for network discovery, investigation, and troubleshooting for users, administrators, and / or other personnel.
[0020] Figure 1A is a conceptual diagram of an exemplary network including a system for analyzing traffic flows between networks and / or within a data center according to one or more aspects of the present disclosure. Figure 1A An exemplary implementation of network system 100 and data center 101 is shown, where data center 101 hosts one or more computing networks, computing domains or projects, and / or cloud-based computing networks, commonly referred to herein as cloud computing clusters. The cloud-based computing clusters may be co-located in a common overall computing environment (such as a single data center) or distributed across environments (such as across different data centers). For example, the cloud-based computing clusters may be different cloud environments, such as various combinations of OpenStack cloud environments, Kubernetes cloud environments, or other computing clusters, domains, networks, etc. In other instances, other implementations of network system 100 and data center 101 may be suitable. This implementation may include Figure 1A a subset of the components included in the example of Figure 1A and / or may include
[0021] In Figure 1A the example of, data center 101 provides an operating environment for applications and services to customer 104 coupled to data center 101 via service provider network 106. Although the functions and operations described in connection with Figure 1A network system 100 may be shown as being distributed across Figure 1A multiple devices in Figure 1AThe features and technologies of one or more of the devices. Similarly, one or more of such devices may include specific components and may perform various technologies that are otherwise attributed to one or more other devices in this disclosure. Further, this disclosure may be combined with Figure 1A Describe specific operations, technologies, features, and / or functions, or describe specific operations, technologies, features, and / or functions otherwise performed by specific components, devices, and / or modules. In other examples, other components, devices, or modules may perform the operation, technology, feature, and / or function. Accordingly, even if not specifically described in this manner herein, however, some operations, technologies, features, and / or functions attributed to one or more components, devices, or modules may be attributed to other components, devices, and / or modules.
[0022] Data center 101 houses infrastructure equipment such as network and storage systems, redundant power supplies, and environmental controls. Service provider network 106 may be coupled to one or more networks managed by other providers and may thus form part of a large-scale public network infrastructure, e.g., the Internet.
[0023] In some examples, data center 101 may represent one of a plurality of geographically distributed network data centers. As Figure 1A shown in the example of, data center 101 is a facility that provides network services to customer 104. Customer 104 may be a general entity such as an enterprise and government or an individual. For example, a network data center may house network services for several enterprises and end users. Other exemplary services may include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. In some examples, data center 101 is an individual network server, network peer, or other.
[0024] In Figure 1AIn an example, data center 101 includes a collection of storage systems, application servers, compute nodes, or other devices, including network devices 110A through 110N (collectively referred to as "network devices 110", representing any number of network devices). The devices 110 can be interconnected via a high-speed switching fabric 121 provided by one or more layers of physical network switches and routers. In some examples, the devices 110 can be included within the fabric 121, however, for ease of illustration, the fabric 121 is shown separately. The network devices 110 can be any number of different types of network devices (core switches, top-of-rack (TOR) switches, spine network devices, leaf network devices, edge network devices, or other network devices), however, in some examples, one or more of the devices 110 can be used as physical compute nodes of the data center. For example, one or more of the devices 110 can provide an operating environment for the running of one or more customer-specified virtual machines or other virtual instances (such as containers). In this example, one or more of the devices 110 can alternatively be referred to as host compute devices or more simply as hosts. Thus, the network devices 110 can run one or more virtual instances, such as virtual machines, containers, or other virtual runtimes for running one or more services, such as virtual network functions (VNFs).
[0025] Generally, each of the network devices 110 can be any type of device that operates on a network and can generate data (e.g., flow datagrams, sFlow datagrams, NetFlow datagrams, etc.) that can be accessed via telemetry or other means, which can include any type of computing device, sensor, camera, node, monitoring device, or other device. Further, some or all of the network devices 110 can represent components of another device, where the component can generate data that is collected via telemetry or other means. For example, some or all of the network devices 110 can represent physical or virtual network devices, such as switches, routers, hubs, gateways, security devices such as firewalls, intrusion detection, and / or intrusion prevention devices,
[0026] The data center 101 can include one or more non-edge switches, routers, hubs, gateways, security devices such as firewalls, intrusion detection, and / or intrusion prevention devices, servers, computer terminals, laptops, printers, databases, wireless mobile devices such as cellular phones or personal digital assistants, wireless access points, bridges, cable modems, application accelerators, or other network devices. The switching fabric 121 can perform Layer 3 routing via a service provider network 106 to route network traffic between the data center 101 and the customer 104. The gateway 108 is used to forward and receive packets between the switching fabric 121 and the service provider network 106.
[0027] A software-defined network (“SDN”) controller 132 provides a logically (and in some cases, physically) centralized controller that facilitates the operation of one or more virtual networks within a data center 101 according to one or more examples of the present disclosure. In some examples, the SDN controller 132 operates in response to configuration inputs received via a northbound API 131 from an orchestration engine 130, which in turn may operate in response to configuration inputs received from an administrator 128 interacting with and / or operating a user interface device 129.
[0028] The user interface device 129 may be implemented as any suitable device for presenting output and / or accepting user input. For example, the user interface device 129 may include a display. The user interface device 129 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or by the administrator 128. For example, the user interface device 129 may represent a workstation, a laptop or notebook computer, a desktop computer, a tablet computer, or any other computing device that operates and / or presents a user interface according to one or more aspects of the present disclosure. In some examples, the user interface device 129 may be physically separated from and / or located at a different location than the controller 201. In this example, the user interface device 129 may communicate with the controller 201 via a network or other communication means. In other examples, the user interface device 129 may be a local peripheral of the controller 201 or may be integrated into the controller 201.
[0029] In some examples, the orchestration engine 130 manages the functions of a data center, such as computing, storage, networking, and application resources. For example, the orchestration engine 130 may create virtual networks for tenants within the data center 101 or across data centers. The orchestration engine 130 may attach virtual machines (VMs) to a tenant's virtual network. The orchestration engine 130 may connect a tenant's virtual network to an external network, e.g., the Internet or a VPN. The orchestration engine 130 may enforce security policies across groups of VMs or at the boundary of a tenant network. The orchestration engine 130 may deploy network services (e.g., load balancers) within a tenant's virtual network.
[0030] In some examples, the SDN controller 132 manages network and network services such as load balancing, security, and can allocate resources from the device 110 used as a host device to various applications via the southbound API 133. That is, the southbound API 133 represents a set of communication protocols used by the SDN controller 132 to make the actual state of the network equal to the desired state specified by the orchestration engine 130. For example, the SDN controller 132 can implement high-level requests from the orchestration engine 130 by configuring physical switches (e.g., TOR switches, chassis switches, and switching fabric 121), physical routers, physical service nodes such as firewalls and load balancers, and virtual services such as virtual firewalls in VMs. The SDN controller 132 maintains routing, physical, and configuration information in the state database.
[0031] The network analysis system 140 interacts with one or more devices 110 (and / or other devices) to collect flow datagrams from the data center 101 and / or the network system 100. A flow datagram refers to a datagram that includes data representing a flow of network traffic. For example, agents operating within the data center 101 and / or the network system 100 can sample the flow of packets within the data center 101 and / or the network system 100 and encapsulate the sampled packets into flow datagrams. The agents can forward the flow datagrams to the network analysis system 140, thereby enabling the network analysis system 140 to collect flow datagrams.
[0032] According to one or more aspects of the present disclosure, Figure 1A the network analysis system 140 in can configure each device 110 to sample packets at respective adaptive sampling rates and generate flow datagrams. For example, in the example described in reference Figure 1A , the network analysis system 140 outputs a signal to each device 110. Each device 110 receives the signal and interprets the signal as a command to sample at a specified sampling rate and generate a flow datagram (including sampled packets). Thereafter, when each device 110 processes data packets, each device 110 communicates the flow datagram including the flow data to the network analysis system 140. In the Figure 1A example, other network devices, including network devices within the switching fabric 121 (and not specifically shown), can also be configured to generate flow datagrams. The network analysis system 140 receives the flow datagrams.
[0033] The network analysis system 140 can store rule data regarding one or more applications. In the present disclosure, an "application" is a label for a specific type of traffic data. In some examples, an application can be a general service, an internal service, or an external application. General services can be identified based on calculations of ports and protocols. Examples of general services can include Transmission Control Protocol (TCP), port 80 and Hypertext Transfer Protocol (HTTP), port 443 and Hypertext Transfer Protocol Secure (HTTPS), port 22 and Secure Shell (SSH), etc. An internal service can be a custom service deployed on a virtual machine (VM) or a set of VMs. An internal service can be identified by a combination of Internet Protocol (IP) address, port, protocol, and virtual network (VN). An external application can be a global service name associated with traffic. An external application can be identified by a combination of port, IP address, Domain Name Service (DNS) domain, etc.
[0034] As described above, the network analysis system 140 can receive a sequence of flow datagrams. The network analysis system 140 can use the rule data regarding applications to identify the flow datagrams associated with an application within the sequence of flow datagrams.
[0035] When sampling a flow, the processing resources of a network device such as the network device 110A need to add a flow header (such as an sFlow header) and, after receiving a sample from, for example, a Packet Forwarding Engine (PFE), transmit the sample to a network analysis system (such as the flow collector of the network analysis system 140). It may be desirable to ensure that the processing resource usage of the flow daemon does not affect other functions of the network device 110A. A relatively low sampling rate can be used to avoid overburdening the processing resources of the network device. However, due to the relatively low sampling rate, the network analysis system 140 may not capture some flows.
[0036] Accordingly, it may be desirable to find the sampling rate of the interfaces of network devices such as network device 110A, i.e., without overburdening the processing capabilities of network device 110A, but still allowing network analysis system 140 to receive samples from each flow processed by the interface to obtain accurate statistics. For example, network analysis system 140 may send a command to network device 110A to sample the flows at an initial first sampling rate. In some examples, this initial first sampling rate may be relatively high to capture samples of each flow processed by the interface of network device 110A. Network device 110A may send flow datagrams to network analysis system 140, including samples of the flows sampled at the initial sampling rate. Network analysis system 140 may determine a first parameter associated with the flow datagram. Network analysis system 140 may send a command to network device 110A to sample the flows at a lower second sampling rate. Network device 110A may send flow datagrams to network analysis system 140, including samples of the flows sampled at the second sampling rate. Network analysis system 140 may determine a second parameter associated with the flow datagram based on the samples of the flows sampled at the second sampling rate. Network analysis system 140 may compare the first parameter with the second parameter and determine whether the first parameter is substantially the same as the second parameter. If the first parameter is the same or substantially the same as the second parameter (e.g., within a predetermined threshold of each other), network analysis system 140 may send a command to network device 110A to sample the flows of the interface at an even lower third sampling rate. Network analysis system 140 and network device 110A may continue this process recursively until the parameters are different or substantially different (e.g., outside a predetermined threshold of each other), at which point network analysis system 140 may send a command to network device 110A to sample at a higher sampling rate.
[0037] For example, the network analysis system 140 may send a command to the network device 110A to sample the flow at an initial first sampling rate. The network device 110A may send flow datagrams to the network analysis system 140, including samples of the flow sampled at the initial sampling rate. The network analysis system 140 may determine a first flow quantity based on the samples sampled at the first sampling rate. The network analysis system 140 may send a command to the network device 110A to sample the flow at a lower second sampling rate. The network device 110A may send flow datagrams to the network analysis system 140, including samples of the flow sampled at the second sampling rate. The network analysis system 140 may determine a second flow quantity based on the samples sampled at the second sampling rate. The network analysis system 140 may compare the first flow quantity with the second flow quantity. If the first flow quantity is equal to or relatively equal to (e.g., within a predetermined threshold of each other) the second flow quantity, the network analysis system 140 may send a command to the network device 110A to sample the flow of the interface at an even lower third sampling rate. The network analysis system 140 and the network device 110A may recursively continue this process until the determined flow quantities are less or substantially less (e.g., outside a predetermined threshold of each other), at which time the network analysis system 140 may send a command to the network device 110A to sample at a higher sampling rate. In some examples, the network analysis system 140 may periodically repeat this process via the network device 110. In this way, the network analysis system 140 may determine a suitable sampling rate for any given interface of the network device 110, thereby avoiding overloading the processing resources of the network device and keeping the device operating at a relatively optimal level even when the traffic pattern of the interface changes.
[0038] In some examples, the network analysis system 140 may determine the maximum number of flows that the network analysis system 140 can handle and the sampling rate may be further based on the maximum number of flows to avoid overloading the network analysis system 140.
[0039] In some examples, a particular interface of network device 110 may handle a relatively large number of flows, which may require a higher sampling rate than other interfaces to sample the individual flows handled by the particular interface. In such cases, a less aggressive scheme may be used to reduce the sampling rate of these interfaces. For example, network analysis system 140 may receive samples of flows sampled at an initial first sampling rate from a second interface of network device 110A. Network analysis system 140 may determine the number of flows based on the samples. Network analysis system 140 may determine that the number of flows is greater than or equal to a predetermined threshold. Network analysis system 140 may determine a new fourth sampling rate based on the number of flows being greater than or equal to the predetermined threshold and send a signal indicating the new fourth sampling rate to network device 110A. The new fourth sampling rate may be higher than the second sampling rate mentioned above. Thus, network analysis system less aggressively reduces the sampling rate of the second interface that is handling a relatively large number of flows.
[0040] In some examples, a particular interface may be considered more important because it can handle traffic, i.e., flows that are more meaningful, for example, for the quality of experience (QoE) of an application. In such cases, a less aggressive scheme may be used to reduce the sampling rate of these particular interfaces. For example, network analysis system 140 may receive samples of flows sampled at an initial first sampling rate from a second interface of network device 110A. Network analysis system 140 may determine that the second interface is handling flows that are more meaningful for QoS than the first interface. For example, network analysis system 140 may perform deep packet inspection to determine that the second interface is handling flows that are more meaningful for QoS than the first interface. Network analysis system 140 may determine a new fourth sampling rate based on the second interface handling flows that are more meaningful for quality of experience than the first interface and send a signal indicating the new fourth sampling rate to network device 110A. The new fourth sampling rate may be higher than the second sampling rate mentioned above. Thus, network analysis system less aggressively reduces the sampling rate of the second interface that is handling flows that are more meaningful for QoS than the first interface.
[0041] For example, network analysis system 140 may independently and dynamically adjust the sampling rate of each interface of network device 110 to identify an appropriate sampling rate for each interface. In some examples, network analysis system 140 may send a command to a network device (e.g., network device 110A) to maintain the sampling rate at the appropriate sampling rate for a given interface.
[0042] For example, an initial first sampling rate can be one packet out of every 100 packets. The network analysis system 140 can send a command to the network device 110A to sample the traffic flow at a rate of one packet out of every 100 packets. The network analysis system 140 can receive from the network device 110 a flow datagram including samples of the traffic flow sampled at a rate of one packet out of every 100 packets. The network analysis system 140 can determine the traffic flow quantity based on these samples. Then, the network analysis system 140 can send a command to sample the traffic flow at a lower sampling rate, such as one packet out of every 200 packets. The network analysis system 140 can receive from the network device 110 a flow datagram including samples of the traffic flow sampled at a rate of one packet out of every 200 packets. The network analysis system 140 can determine the traffic flow quantity based on these samples. The network analysis system 140 can compare the determined traffic flow quantities. If the determined traffic flow quantities are equal or relatively equal, the network analysis system 140 can continue to lower the sampling rate, such as to one packet out of every 400 packets. This process can be repeated until the determined traffic flow quantities are not equal or are substantially different, and then the network analysis system 140 can send a command to increase the sampling rate. In some examples, when determining the appropriate sampling rate for a given interface, the network analysis system can stop applying adaptive sampling to that interface or apply adaptive sampling to that interface less frequently. Although the sampling rates discussed herein are based on the number of packets, in some examples, the sampling rate can be based on time, for example, one sample per tenth of a second.
[0043] Figure 1B is a conceptual diagram showing exemplary components of a system for analyzing traffic flows between networks and / or within a data center in accordance with one or more aspects of the present disclosure. Figure 1B including in combination Figure 1A the multiple identical elements described. Figure 1B The elements shown in Figure 1A can correspond to the elements shown in Figure 1A and identified by like-numbered reference numerals in Figure 1A . Generally, in some examples, although the element may relate to alternative implementations having more, fewer, and / or different capabilities and attributes, the like-numbered element can be implemented in a manner consistent with the description of the corresponding element provided in combination with
[0044] However, different from Figure 1A , Figure 1BShows the components of the network analysis system 140. The network analysis system 140 is shown as including a load balancer 141, a flow collector 142, a queue & event store 143, a topology & metric source 144, a data store 145, and a flow API 146. Generally, the network analysis system 140 and the components of the network analysis system 140 are designed and / or configured to ensure high availability and the ability to handle a large amount of flow data. In some examples, multiple instances of the components of the network analysis system 140 can be orchestrated (e.g., by the orchestration engine 130) to run on different physical servers to ensure that there is no single point of failure for any component of the network analysis system 140. In some examples, the network analysis system 140 or its components can be independently and horizontally scaled to enable efficient and / or effective processing of the required traffic (e.g., flow data).
[0045] As Figure 1A , Figure 1B The network analysis system 140 in
[0046] In Figure 1B , the load balancer 141 of the network analysis system 140 receives flow data packets from the devices 110. The load balancer 141 can distribute the flow data packets across multiple flow collectors to ensure an active / active failover strategy for the flow collectors. In some examples, multiple load balancers 141 may be required to ensure high availability and scalability.
[0047] The flow collector 142 collects flow datagrams from the load balancer 141. For example, the flow collector 142 of the network analysis system 140 receives flow datagrams from the device 110 and processes the flow datagrams (after being processed by the load balancer 141). The flow collector 142 forwards the upstream flow datagrams to the queue & event store 143. In some examples, the flow collector 142 can address, process, and / or adapt unified datagrams in sFlow, NetFlow v9, IPFIX, jFlow, Contrail Flow, and other formats. The flow collector 142 is capable of parsing internal headers in sFlow datagrams and other flow datagrams (i.e., headers of at least partially encapsulated packets). The flow collector 142 is capable of processing message overflows, rich flow records with topological information (e.g., AppFormix topological information), and other types of messages and datagrams. Before writing or forwarding the data to the queue & event store 143, the flow collector 142 is also capable of converting the data into binary format. The underlying flow data of the "sFlow" type (referring to "sampled flow") is a packet export standard for Layer 2 of the OSI model. It provides a way to export truncated packets together with interface counters for network monitoring purposes.
[0048] According to the techniques of the present disclosure, the flow collector 142 can receive a first sample of a flow sampled at a first sampling rate from an interface of the network device 110A. The flow collector 142 can determine a first parameter based on the first sample. The flow collector 142 can receive a second sample of the flow sampled at a second sampling rate from the interface, where the second sampling rate is different from the first sampling rate. The flow collector 142 can determine a second parameter based on the second sample. The flow collector 142 can determine a third sampling rate based on the first parameter and the second parameter. The flow collector 142 can send a signal indicating the third sampling rate to the network device. The flow collector 142 can receive a third sample of the flow sampled at the third sampling rate from the interface.
[0049] The queue & event store 143 can receive data from one or more flow collectors 142, store the data, and make the data available for ingestion in the data store 145. In some examples, this can separate the task of receiving and storing large amounts of data from the task of indexing the data and preparing it for analytical queries. In some examples, the queue & event store 143 can also enable independent users to directly consume the stream records. In some examples, the queue & event store 143 can be used to detect anomalies and generate real-time alerts. In some examples, the flow data can be parsed by reading the encapsulated packets, including VXLAN, MPLS over UDP, and MPLS over GRE.
[0050] The topology & metrics source 144 can enrich or augment the datagram with topology information and / or metrics information. For example, the topology & metrics source 144 can provide network topology metadata, which can include identified nodes or network devices, configuration information, configurations, established links, and other information about the nodes and / or network devices. In some examples, the topology & metrics source 144 can use AppFormix topology data or can be a running AppFormix module. The information received from the topology & metrics source 144 can be used to enrich the flow datagrams collected by the flow collector 142 and support the flow API 146 when processing queries to the data store 145.
[0051] The data store 145 can be configured to store data (such as datagrams) in an indexed format received from the queue & event store 143 and the topology & metrics source 144, enabling fast aggregation queries and fast random access data retrieval. In some examples, the data store 145 can achieve fault tolerance and high availability by sharing and replicating the data.
[0052] The flow API 146 can process query requests transmitted by one or more user interface devices 129. For example, in some examples, the flow API 146 can receive a query request from a user interface device 129 via an HTTP POST request. In this example, the flow API 146 converts the information included in the request into a query to the data store 145. To create the query, the flow API 146 can use the topology information from the topology & metrics source 144. The flow API 146 can use one or more of these queries to perform analysis on behalf of the user interface device 129. The analysis can include traffic deduplication, overlay - underlay correlation, traffic path identification, and / or heatmap traffic calculation. Specifically, the analysis may involve: correlating the underlying flow data with the overlay flow data, thereby enabling identification of which underlying network devices are related to the transit traffic flowing through the virtual network and / or flowing between two virtual machines. Through techniques according to one or more aspects of the present disclosure, the network analysis system 140 can associate data flows with applications in a data center, such as a multi - tenant data center.
[0053] Figure 2 is a block diagram illustrating an exemplary network analysis system in accordance with one or more aspects of the present disclosure.
[0054] The network analysis system 140 can be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, devices, cloud computing systems, and / or other computing systems capable of performing the operations and / or functions described in one or more aspects of the present disclosure. In some examples, the network analysis system 140 represents a cloud computing system, server farm, and / or server cluster (or a portion thereof) that provides services to client devices and other devices or systems. In other examples, the network analysis system 140 can represent one or more virtual computing instances (e.g., virtual machines, containers) of a data center, cloud computing system, server farm, and / or server cluster or can be implemented via one or more virtual computing instances (e.g., virtual machines, containers) of a data center, cloud computing system, server farm, and / or server cluster.
[0055] In Figure 2 examples, the network analysis system 140 can include a power supply 241, processing circuitry 243, one or more communication units 245, one or more input devices 246, one or more output devices 247, and one or more storage devices 250. The one or more storage devices 250 can include one or more collector modules 252, a command interface module 254, an API server 156, and a streaming database 259. In some examples, the network analysis system 140 includes additional components, fewer components, or different components.
[0056] One or more devices, modules, storage areas, or other components in the network analysis system 140 can be interconnected to enable communication (physically, communicatively, and / or operationally) between components. In some examples, this connectivity can be provided via a communication channel (e.g., communication channel 242), a system bus, a network connection, an interprocess communication data structure, or any other method for communicating data.
[0057] The power supply 241 can provide power to one or more components in the network analysis system 140. The power supply 241 can receive power from a primary alternating current (AC) power supply at a data center, building, residence, or other location. In other examples, the power supply 241 can be a battery or a device that supplies direct current (DC). In yet further examples, the network analysis system 140 and / or the power supply 241 can receive power from another power source. One or more of the devices or components shown within the network analysis system 140 can be connected to the power supply 241 and / or can receive power from the power supply 241. The power supply 241 can have intelligent power management or consumption capabilities and can be controlled, accessed, or adjusted via one or more modules in the network analysis system 140 and / or via the processing circuitry 243 to intelligently consume, distribute, supply, or otherwise manage power.
[0058] The processing circuitry 243 of the network analysis system 140 can implement and / or execute functions and / or instructions associated with the network analysis system 140 or with one or more modules shown and / or described herein. The processing circuitry 243 can be, can include, and / or can form part of a processing circuitry that performs operations in accordance with one or more aspects of the present disclosure. Examples of the processing circuitry 243 include a microprocessor, an application processor, a display controller, an auxiliary processor, one or more sensor hubs, and any other hardware configured to function as a processor, processing unit, or processing device. The network analysis system 140 can utilize the processing circuitry 243 to perform operations in accordance with one or more aspects of the present disclosure using software, hardware, firmware, or a combination of hardware, software, and firmware that resides in and / or runs on the network analysis system 140.
[0059] For example, the processing circuitry 243 can receive a first sample of a stream sampled at a first sampling rate from an interface of another network device (e.g., network device 110A). The processing circuitry 243 can determine a first parameter based on the first sample. For example, the processing circuitry 243 can count a first stream quantity based on the first sample. In some examples, the processing circuitry 243 can control one or more communication units 245 to send a signal indicating a second sampling rate to the network device 110A. The processing circuitry 243 can receive a second sample of the stream sampled at the second sampling rate from the interface, where the second sampling rate is different from the first sampling rate. The processing circuitry 243 can determine a second parameter based on the second sample. For example, the processing circuitry 243 can count a second stream quantity based on the second sample. The processing circuitry 243 can determine a third sampling rate based on the first parameter and the second parameter. For example, if the first parameter is approximately the same as the second parameter, the processing circuitry can determine a third sampling rate lower than the first sampling rate and the second sampling rate. If the first parameter is substantially different from the second parameter, the processing circuitry 243 can determine a third sampling rate higher than the second sampling rate. The processing circuitry 243 can control one or more communication units 245 to send a signal indicating the third sampling rate to the network device 110A. The processing circuitry 243 can receive a third sample of the stream sampled at the third sampling rate from the interface.
[0060] One or more communication units 245 of the network analysis system 140 may communicate with devices located outside the network analysis system 140 by sending and / or receiving data, and in some aspects, may operate as input and output devices. For example, one or more communication units 245 may send signals indicating respective sampling rates to network devices, such as network device 210. One or more communication units 245 may also receive datagrams including data of packets sampled at respective sampling rates.
[0061] In some examples, one or more communication units 245 may communicate with other devices over a network. In other examples, one or more communication units 245 may transmit and / or receive radio signals over a radio network such as a cellular radio network. Examples of one or more communication units 245 include network interface cards (e.g., such as Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device capable of transmitting and / or receiving information. Other examples of communication unit 245 may include devices capable of communicating via mobile devices as well as those found in universal serial bus (USB) controllers, etc. GPS, NFC, ZigBee, with cellular networks (e.g., 3G, 4G, 5G), and radios. This communication may follow, implement, or comply with appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, Bluetooth, NFC, or other technologies or protocols.
[0062] One or more input devices 246 may represent any input device of the network analysis system 140 not otherwise separately described herein. One or more input devices 246 may generate, receive, and / or process inputs from any type of device capable of detecting human or machine input. For example, one or more input devices 246 may generate, receive, and / or process inputs in the form of electrical, physical, audio, image, and / or visual inputs (e.g., peripherals, keyboards, microphones, cameras).
[0063] One or more output devices 247 may represent any output device of the network analysis system 140 not otherwise separately described herein. One or more output devices 247 may generate, receive, and / or process inputs from any type of device capable of detecting human or machine input. For example, one or more output devices 247 may generate, receive, and / or process outputs in the form of electrical and / or physical outputs (e.g., peripherals, actuators).
[0064] One or more storage devices 250 within the network analysis system 140 may store information for processing during operation of the network analysis system 140. The one or more storage devices 250 may store program instructions and / or data associated with one or more modules described in accordance with one or more aspects of the present disclosure. The processing circuitry 243 and the one or more storage devices 250 may provide an operating environment or platform for the module, which may be implemented as software, however, in some examples, may include any combination of hardware, firmware, and software. The processing circuitry 243 may execute instructions and the one or more storage devices 250 may store the instructions and / or data for one or more modules. The combination of the processing circuitry 243 and the one or more storage devices 250 may retrieve, store, and / or execute the instructions and / or data for one or more applications, modules, or software. The processing circuitry 243 and / or the one or more storage devices 250 may also be operatively coupled to one or more other software and / or hardware components, including but not limited to one or more components of the network analysis system 140 and / or one or more devices or systems shown as being connected to the network analysis system 140.
[0065] In some examples, the one or more storage devices 250 are implemented via a transient memory, which means that the primary purpose of the one or more storage devices is not long-term storage. One or more storage devices 250 of the network analysis system 140 may be configured as volatile memory for short-term storage of information and thus, if deactivated, do not retain the stored content. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art. In some examples, the one or more storage devices 250 further include one or more computer-readable storage media. The one or more storage devices 250 may be configured to store a larger amount of information than volatile memory. The one or more storage devices 250 may be further configured for long-term storage of information as non-volatile memory space and retain the information after power-on / off cycles. Examples of non-volatile memory include magnetic hard disks, optical disks, flash memory, or forms of electrically erasable memory (EPROM) or electrically erasable and programmable (EEPROM) memory.
[0066] The collector module 252 may perform functions related to: receiving streaming datagrams; determining parameters associated with the sampled stream, such as the number of streams sampled at a given sampling rate; and performing load balancing as needed to ensure high availability, throughput, and scalability for collecting the streaming data when executed by the processing circuitry 243. The collector module 252 may process the data and prepare the data for storage in the stream database 259. In some examples, the collector module 252 may store the streaming data in the stream database 259.
[0067] The command interface module 254 may perform functions related to: generating a user interface for presenting the results of the analytical queries performed by the API server 256 when executed by the processing circuitry 243. In some examples, the command interface module 254 may generate information sufficient to generate a set of user interfaces.
[0068] The API server 256 may perform analytical queries involving data stored in the stream database 259 derived from a collection of streaming datagrams. In some examples, the API server 256 may receive requests in the form of information derived from HTTP POST requests and, in response, may convert the requests into queries to be performed on the stream database 259. Further, in some examples, the API server 256 may obtain topology information related to the network device 110A and perform analyses including data deduplication, overlay-underlay correlation, traffic path identification, and heatmap traffic calculation.
[0069] The stream database 259 may represent any suitable data structure or storage medium for storing information related to the data stream information, including the storage of streaming datagrams. The stream database 259 may store data in an indexed format, which enables fast data retrieval and execution of queries. The information stored in the stream database 259 is searchable and / or classifiable such that one or more modules within the network analysis system 140 may provide an input requesting information from the stream database 259 and, in response to the input, receive the information stored in the stream database 259. The stream database 259 is maintained primarily by the collector module 252. The stream database 259 may be implemented by multiple hardware devices and may achieve fault tolerance and high availability through sharing and replicating data. In some examples, the stream database 259 may be implemented using the open-source ClickHouse column-oriented database management system.
[0070] For example, the command interface module 254 of the network analysis system 140 may receive a query from the user interface device 129 ( Figures 1A to 1B)). The communication unit 245 of the network analysis system 140 detects signals and provides information to the command interface module 254, which in turn provides a query to the API server 256 based on the provided information. The query can be a request for information about the network system 100 within a given time window. The API server 256 processes the query regarding the data in the flow database 259. For example, a user of the user interface device 129 (e.g., the administrator 128) may wish to determine which network devices 110 are involved in the flows associated with a specific application. The API 146( Figure 1B ) can operate in the same manner as the API server 256.
[0071] The network analysis system 140 can cause a user interface to be presented at the user interface device 129 based on the query results. For example, the API server 256 can output information about the query results to the command interface module 254. In this example, the command interface module 254 uses the information from the API server 256 to generate data sufficient to create a user interface. Further, in this example, the command interface module 254 causes one or more communication units 245 to output signals. In this example, the user interface device 129 detects the signals and generates a user interface based on the signals. The user interface device 129 presents the user interface at a display associated with the user interface device 129.
[0072] Figure 2 The modules (collector module 252, command interface module 254, API server 256) shown and / or described in this disclosure or elsewhere can perform the operations described using software, hardware, firmware, or a combination of hardware, software, and firmware residing in and / or executed on one or more computing devices. For example, a computing device can run one or more of these modules using multiple processors or multiple devices. A computing device can run one or more of these modules as a virtual machine running on underlying hardware. One or more of these modules can run as one or more services in an operating system or computing platform. One or more of these modules can run as one or more runnable programs in the application layer of a computing platform. In other examples, the functions provided by the modules can be implemented by dedicated hardware devices.
[0073] Although specific modules, data stores, components, programs, executables, data items, functional units, and / or other items included within one or more storage devices may be shown individually, one or more of such items may be combined and operate as a single module, component, program, executable, data item, or functional unit. For example, one or more modules or data stores may be combined or partially combined such that they operate as a single module or provide a function. Further, one or more modules may interact with each other and / or operate in conjunction with each other such that, for example, one module serves as a service or an extension of another module. Additionally, each module, data store, component, program, executable, data item, functional unit, or other item shown within the storage device may include a number of components, sub-components, modules, sub-modules, data stores, and / or other components or modules or data stores not shown.
[0074] Further, each module, data store, component, program, executable, data item, functional unit, or other item shown within the storage device may be implemented in a variety of ways. For example, each module, data store, component, program, executable, data item, functional unit, or other item shown within the storage device may be implemented as a downloadable or pre-installed application or “app”. In other examples, each module, data store, component, program, executable, data item, functional unit, or other item shown within the storage device may be implemented as part of an operating system running on a computing device.
[0075] Figure 3 is a block diagram illustrating an exemplary network device in accordance with one or more aspects of the present disclosure. In some examples, network device 110A may be an example of a TOR switch or a chassis switch.
[0076] In Figure 3In an illustrative example, network device 110A includes a control unit 32 that, in some examples, provides control plane functionality for network device 110A. Control unit 32 can include processing circuitry 40, a routing engine 42 that includes routing information 44 and a resource module 46, a software plug-in 48, a hypervisor 50, and VMs 52A–52N (collectively “VM 52”). In some examples, processing circuitry 40 can include one or more processors configured to implement and / or process functions and / or instructions for execution within control unit 32. For example, processing circuitry 40 is capable of processing instructions stored in a storage device. For example, processing circuitry 40 can include a microprocessor, a DSP, an ASIC, an FPGA, or equivalent discrete or integrated logic circuitry, or a combination of any of the foregoing devices or circuits. Accordingly, processing circuitry 40 can include any suitable architecture, whether in hardware, software, firmware, or any combination thereof, to perform the functions ascribed herein to processing circuitry 40.
[0077] In some cases, processing circuitry 40 can include a set of compute nodes. A compute node can be a component of a central processing unit (CPU) that receives information, outputs information, performs calculations, performs actions, manipulates data, or any combination thereof. Additionally, in some examples, a compute node can include transient storage, networking, memory, and processing resources for running one or more VMs, containers, a container pool, or other types of workloads. As such, a compute node can represent compute resources for running a workload. A compute node used for running a workload can be referred to herein as an “in-use compute node”. Additionally, a compute node not used for running a workload can be referred to herein as an “unused compute node”. Thus, if processing circuitry 40 includes a relatively large number of unused compute nodes, processing circuitry 40 can have a relatively large amount of compute resources available for running workloads, and if processing circuitry 40 includes a relatively small number of unused compute nodes, processing circuitry 40 can have a relatively small amount of compute resources available for running workloads.
[0078] Processing circuitry 40 can sample flows processed by network device 110A and can create a flow datagram that includes sampled packets of the flow. This sampling and creating of the flow datagram uses the processing resources of processing circuitry 40. By adaptively changing the sampling rate, the techniques of the present disclosure can balance the use of the processing resources of processing circuitry 40 with the monitoring of flows by flow collector 142.
[0079] In some examples, the amount of available compute resources in control unit 32 depends on network device 110A in Figure 1A and Figure 1BRoles within data center 101. For example, if network device 110A represents the master switch of a logical switch, it may consume a large amount of computing resources within processing circuitry 40 to provide control plane functionality to the respective switches within the logical switch (e.g., respective network devices 110B - 110N). Thus, if network device 110A represents a line card, it may consume a smaller amount of computing resources within processing circuitry 40 compared to the case where network device 10A represents the master switch, since the line card can receive control plane functionality from the master switch.
[0080] In some examples where control unit 32 provides control plane functionality to network device 110A and / or other network devices, control unit 32 includes a routing engine 42 configured to communicate with Figure 2 forwarding unit 60 of a network device not shown in and other forwarding units. In some cases, where the network device is part of a logical switch, routing engine 42 may represent control plane management for packet forwarding throughout the network device. For example, network device 110A includes interface cards 70A - 70N (collectively "IFC 70") that receive packets via an inbound link and transmit packets via an outbound link. IFC 70 typically has one or more physical network interface ports. In some examples, each network interface port (also referred to herein as an interface) may be sampled at a sampling rate independently determined by flow collector 142 based on a comparison of parameters associated with sampling at different sampling rates, the number of flows the interface is handling, and / or whether the flows the interface is handling are more significant for QoS than another interface. In some examples, after receiving a packet via IFC 70, network device 110A uses forwarding unit 60 to forward the packet to the next destination based on operations performed by routing engine 42.
[0081] Routing engine 42 may provide an operating environment for various protocols ( Figure 2 not shown in) executed at different layers of the network stack. Routing engine 42 may be responsible for the maintenance of routing information 44 to reflect the current topology of the network and other network entities connected to network device 110A. Specifically, the routing protocol periodically updates routing information 44 based on routing protocol messages received by network device 110A to accurately reflect the topology of the network and other entities. The protocol may be a software process running on processing circuitry 40. As such, routing engine 42 may occupy a set of computing nodes within processing circuitry 40 such that the set of computing nodes is unavailable for running VMs. For example, routing engine 42 may include a bridging port extension protocol such as IEEE 802.1BR. Routing engine 42 may also include network protocols operating at the network layer of the network stack. In Figure 2In an example, the network protocol may include one or more control and routing protocols, such as, Border Gateway Protocol (BGP), Interior Gateway Protocol (IGP), Label Distribution Protocol (LDP), and / or Resource Reservation Protocol (RSVP). In some examples, the IGP may include the Open Shortest Path First (OSPF) protocol or the Intermediate System to Intermediate System (IS-IS) protocol. The routing engine 42 may also include one or more daemons, including user-level processes that run network management software, execute routing protocols to communicate with peer routers or switches, maintain and update one or more routing tables, and create one or more forwarding tables for installation into the forwarding unit 60, and other functions.
[0082] For example, the routing information 44 may include routing data that describes the various routes within the network, and corresponding next-hop data for each route that indicates the appropriate neighboring devices within the network. The network device 110A updates the routing information 44 based on the received announcements to accurately reflect the topology of the network. Based on the routing information 44, the routing engine 42 may generate forwarding information ( Figure 2 (not shown in the figure) and install the forwarding data structure into the forwarding information within the forwarding unit 60. The forwarding information associates network destinations with the specified next-hop and the corresponding interface ports within the forwarding plane. In addition, the routing engine 42 may include one or more resource modules 46 for configuring the resources of the expansion ports and uplink ports on the satellite devices interconnected with the network device 110A. The resource module 46 may include a scheduler module for configuring Quality of Service (QoS) policies, a firewall module for configuring firewall policies, or other modules for configuring the resources of the network device.
[0083] In some examples, the processing circuit 40 runs the software plug-in 48 and the hypervisor 50. In some examples, the software plug-in 48 enables the network device 110A to communicate with the orchestration engine 130. The software plug-in 48 may be configured to interface with a lightweight message broker and the hypervisor 50 such that the software plug-in 48 serves as a mediator between the orchestration engine 130 and the hypervisor 50. The orchestration engine 130 outputs instructions to manage the VMs 52, and the hypervisor 50 is configured to implement the instructions received from the orchestration engine 130.
[0084] The software plug-in 48 may interface with a lightweight message broker such as RabbitMQ to exchange messages with the orchestration engine 130. Thus, the software plug-in may implement any combination of AMQP, STOMP, MQTT, and HTTP. The software plug-in 48 may be configured to generate and process the messages sent via the lightweight message broker.
[0085] In some cases, the hypervisor 50 can be configured to communicate with the software plug-in 48 to manage the life cycle of the VM 52. In some examples, the hypervisor 50 represents a Kernel-based Virtual Machine (KVM) hypervisor, a special operating mode of Quick Emulator (QEMU) that allows the Linux Kernel to be used as a hypervisor. The hypervisor 50 can perform hardware-assisted virtualization to create a virtual machine that emulates the functions of computer hardware when running the corresponding virtual machine on a computing node using the processing circuitry 40. The hypervisor 50 can interface with the software plug-in 48.
[0086] The forwarding unit 60 represents the hardware and logical functions that provide high-speed forwarding of network traffic. In some examples, the forwarding unit 60 can be implemented as a programmable forwarding plane. The forwarding unit 60 can include one or more forwarding chipsets programmed with forwarding information that maps network destinations to a specified next hop and a corresponding output interface port. In some examples, the forwarding unit 60 includes the forwarding information. Based on the routing information 44, the forwarding unit 60 maintains the forwarding information that associates network destinations with a specified next hop and a corresponding interface port. For example, the routing engine 42 can analyze the routing information 44 and generate the forwarding information based on the routing information 44. The forwarding information can be maintained in the form of one or more tables, linked lists, radix trees, databases, flat files, or any other data structure. Although Figure 2 not shown therein, however, the forwarding unit 60 can include a CPU, a memory, and one or more ASICs.
[0087] shown for illustrative purposes only Figure 2 the architecture of the network device 110A shown in. The present disclosure is not limited to this architecture. In other examples, the network device 110A can be configured in various ways. In one example, some of the functions of the routing engine 42 and the forwarding unit 60 can be distributed within the IFC 70.
[0088] Elements in control unit 32 may be implemented solely in software, or in hardware, or as a combination of software, hardware, or firmware. For example, control unit 32 may include one or more processors that execute software instructions, one or more microprocessors, DSPs, ASICs, FPGAs, or any other equivalent integrated or discrete logic circuitry, or any combination thereof. In such cases, the various software modules of control unit 32 may include executable instructions stored, embedded, or encoded in a computer-readable medium, such as a computer-readable storage medium that contains instructions. For example, when executed, the instructions embedded or encoded in a computer-readable medium may cause a programmable processor or other processor to perform a method. The computer-readable storage medium may include random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), non-volatile random access memory (NVRAM), flash memory, a hard disk, a CD-ROM, a floppy disk, magnetic tape, a solid state drive, magnetic media, optical media, or other computer-readable media. The computer-readable medium may be encoded with instructions corresponding to various aspects of network device 110A, such as protocols. In some examples, control unit 32 retrieves and executes instructions from memory for these aspects.
[0089] Figure 4 is a flowchart showing an exemplary method in accordance with one or more techniques of the present disclosure. Processing circuitry 243 of network analysis system 140 ( Figure 2 ) may receive a first sample (400) of a stream sampled at a first sampling rate from an interface of another network device. For example, processing circuitry 40 ( Figure 3 ) of network device 110A may sample a stream being processed in IFC 70A ( Figure 3 ), create a first datagram including the sample, and forward the first datagram to network analysis system 140 that receives the first sample in the first datagram. Processing circuitry 243 may determine a first parameter based on the first sample (402). For example, processing circuitry 243 may count a first number of flows in the first sample to determine the first parameter.
[0090] The processing circuit 243 may receive second samples (404) of a stream sampled at a second sampling rate from the interface, where the second sampling rate is different from the first sampling rate. For example, the processing circuit 40 may sample the stream being processed by the IFC 70A at the second sampling rate, create a second datagram including the second samples, and forward the second datagram to the network analysis system 140 that receives the second samples in the second datagram. The processing circuit 243 may determine a second parameter based on the second samples (406). For example, the processing circuit 243 may count the number of second streams in the second samples to determine the second parameter.
[0091] The processing circuit 243 may determine a third sampling rate based on the first parameter and the second parameter (408). The third sampling rate may be different from the second sampling rate (e.g., higher or lower). For example, the processing circuit 243 may compare the first parameter (e.g., the number of first streams) with the second parameter (e.g., the number of second streams) to determine whether the first parameter is equal to the second parameter or whether the first parameter is within a predetermined threshold of the second parameter. For example, if the first parameter is equal to the second parameter or within the predetermined threshold of the second parameter, the processing circuit 243 may determine that the third sampling rate is a sampling rate lower than the second sampling rate. If the first parameter is not equal to the second parameter or not within the predetermined threshold of the second parameter, the processing circuit 243 may determine that the third sampling rate is a sampling rate higher than the second sampling rate. For example, the third sampling rate may be equal to the first sampling rate or may be a sampling rate between the second sampling rate and the first sampling rate.
[0092] The processing circuit 243 may control one or more communication units 245 to send a signal indicating the third sampling rate to another network device (410). For example, the network management system 140 may instruct the network device 110A to sample the IFC 70A at the third sampling rate. The processing circuit 243 may receive third samples (412) of a stream sampled at the third sampling rate from the interface. For example, the processing circuit 40 may sample the stream being processed by the IFC 70A at the third sampling rate, create a third datagram including the third samples, and forward the third datagram to the network management system 140 that receives the third samples.
[0093] In some examples, the first parameter refers to the first flow quantity and the second parameter refers to the second flow quantity. In some examples, IFC 70A is the first interface, and the network management system 140 can receive a fourth sample of the flow sampled at the first sampling rate from a second interface (e.g., IFC 70B). The network management system 140 can determine a third flow quantity based on the fourth sample. The network management system 140 can determine that the third flow quantity is greater than or equal to a predetermined threshold. The network management system 140 can determine a fourth sampling rate based on the third flow quantity that is greater than or equal to the predetermined threshold. The network management system 140 can send a signal indicating the fourth sampling rate to the network device 110A, where the fourth sampling rate is higher than the second sampling rate. The network management system 140 can receive a fifth sample of the flow sampled at the fourth sampling rate from the second interface.
[0094] In some examples, IFC 70A is the first interface, and the network management system 140 can receive a fourth sample of the flow sampled at the first sampling rate from a second interface (e.g., IFC 70B). The network management system 140 can determine that the second interface is processing flows that are more significant in terms of quality of experience than the first interface. The network management system 140 can determine a fourth sampling rate based on the second interface processing flows that are more significant in terms of quality of experience than the first interface. The network management system 140 can send a signal indicating the fourth sampling rate to the network device 110A, where the fourth sampling rate is higher than the second sampling rate. The network management system 140 can receive a fifth sample of the flow sampled at the fourth sampling rate from the second interface.
[0095] In some examples, the first flow quantity is equal to the second flow quantity or within a predetermined threshold of the second flow quantity, and the third sampling rate is lower than the second sampling rate. In some examples, the predetermined threshold is static or based on the flow quantity determined at one of multiple sampling rates. For example, the predetermined threshold can be the flow quantity or a percentage of the flow quantity determined at a certain sampling rate, such as the second sampling rate. The predetermined threshold can be a user-configured value or determined based on device capabilities or device roles. When determined based on device capabilities, a more powerful (capable) device can have a relatively higher threshold than a less powerful device. When determined based on device roles, various device roles such as server blades, border leads, gateways, spines, etc. can have predetermined corresponding thresholds.
[0096] In some examples, the first sampling rate is higher than the second sampling rate, the first flow quantity is greater than the second flow quantity, and the third sampling rate is higher than the second sampling rate. In some examples, the third sampling rate is equal to the first sampling rate.
[0097] In some examples, the first sample, the second sample, and the third sample are sFlow samples. In some examples, the network management system 140 can periodically repeat Figure 4 the techniques in
[0098] For the processes, apparatuses, and other examples or illustrations included in any of the flowcharts or flow diagrams described herein, the specific operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be added, combined, or entirely omitted (e.g., not all of the described actions or events are necessary for the implementation of the present technique). Moreover, in certain examples, for instance, operations, actions, steps, or events may be performed by multithreading, interrupt handling, or simultaneously by multiple processors rather than sequentially. Further specific operations, actions, steps, or events may be performed automatically even if not specifically identified as such. Additionally, alternatively, specific operations, actions, steps, or events described as being performed automatically may not be performed automatically, but rather, in some examples, may be performed in response to an input or another event.
[0099] The figures included herein each illustrate at least one exemplary implementation of an aspect of the present disclosure. However, the scope of the present disclosure is not limited to this implementation. Accordingly, in other instances, other examples or alternative implementations of the systems, methods, or techniques described herein may be appropriate in addition to those shown in the figures. This implementation may include a subset of the devices and / or components included in the figures and / or may include additional devices and / or components not shown in the figures.
[0100] The detailed description set forth above is intended as a description of various configurations and is not intended to represent the only configuration in which the concepts described herein may be implemented. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, the concepts may be implemented without these specific details. In some instances, well-known structures and components are shown in block diagram form in the reference figures to avoid obscuring the concepts.
[0101] Accordingly, although one or more implementations of various systems, apparatuses, and / or components may be described with reference to specific figures, the systems, apparatuses, and / or components may be implemented in many different ways. For example, one or more apparatuses shown as separate apparatuses in the figures herein may alternatively be implemented as a single apparatus; one or more components shown as separate components may alternatively be implemented as a single component. Additionally, in some examples, one or more apparatuses shown as a single apparatus in the figures herein may alternatively be implemented as multiple apparatuses; one or more components shown as a single component may alternatively be implemented as multiple components. Each of the multiple apparatuses and / or components may be directly coupled via wired or wireless communication and / or remotely coupled via one or more networks. Additionally, one or more apparatuses or components shown in the respective figures herein may alternatively be implemented as part of another apparatus or component not shown in the figure. In this and other ways, two or more apparatuses or components may perform some of the functions described herein via distributed processing.
[0102] Further, specific operations, techniques, features, and / or functions may be described herein as being performed by specific components, apparatuses, and / or modules. In other examples, the operation, technique, feature, and / or function may be performed by different components, apparatuses, or modules. Accordingly, in other examples, even if not specifically described in this way herein, some operations, techniques, features, and / or functions described herein as being attributable to one or more components, apparatuses, or modules may be attributable to other components, apparatuses, and / or modules.
[0103] Although specific advantages have been identified in connection with the description of some examples, each of the other examples may include some, none, or all of the recited advantages. Other advantages, whether technical or otherwise, may become apparent to those of ordinary skill in the art in this disclosure. Further, although specific examples have been disclosed herein, any number of techniques may be used to implement aspects of the disclosure, whether currently known or not, and accordingly, the disclosure is not limited to the examples specifically described and / or shown in this disclosure.
[0104] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on a computer-readable medium as one or more instructions or code and / or the functionality may be transmitted via the computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium (corresponding to a non-transitory medium such as a data storage medium) or a communication medium including any medium that facilitates transfer of a computer program from one place to another (e.g., according to a communication protocol). Thus, the computer-readable medium generally may correspond to: (1) a tangible computer-readable storage medium, i.e., non-transitory; or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that is accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0105] By way of example, and not limitation, the computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage medium and data storage medium do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transitory, tangible storage media. Disk and optical disks used herein include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0106] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor" or "processing circuit" as used herein may refer to any of the foregoing structures or any other structure suitable for implementing the described techniques. In addition, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Further, the techniques may be implemented entirely in one or more circuits or logic elements.
[0107] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or IC sets (e.g., chip sets). In the present disclosure, each component, module, or unit is described as enhancing a functional aspect of a device configured to perform the disclosed techniques, however, it is not necessarily required that each be implemented by distinct hardware units. Rather, as described above, each unit may be combined in hardware units or provided by a collection of interoperating hardware units in conjunction with suitable software and / or firmware, including the one or more processors described above.
Claims
1. A method for flow monitoring, the method comprising: Receiving a first sample of a flow sampled at a first sampling rate from an interface of a network device; Determining a first parameter based on the first sample; Receiving a second sample of a flow sampled at a second sampling rate from the interface, wherein the second sampling rate is different from the first sampling rate; Determining a second parameter based on the second sample; Determining a third sampling rate based on the first parameter and the second parameter; Sending a signal indicating the third sampling rate to the network device; and Receiving a third sample of a flow sampled at the third sampling rate from the interface.
2. The method according to claim 1, wherein The first parameter refers to a first flow quantity, and wherein the second parameter refers to a second flow quantity.
3. The method according to claim 2, wherein The interface is a first interface, and the method further comprises: Receiving a fourth sample of a flow sampled at the first sampling rate from a second interface of the network device; Determining a third flow quantity based on the fourth sample; Determining that the third flow quantity is greater than or equal to a predetermined threshold; Determining a fourth sampling rate based on the third flow quantity that is greater than or equal to the predetermined threshold; Sending a signal indicating the fourth sampling rate to the network device, wherein the fourth sampling rate is higher than the second sampling rate; and Receiving a fifth sample of a flow sampled at the fourth sampling rate from the second interface.
4. The method according to claim 2, wherein, The interface is a first interface, and the method further comprises: Receiving a fourth sample of a flow sampled at the first sampling rate from a second interface of the network device; Determining that the second interface is processing flows that are more significant in terms of quality of experience than the first interface; Determining a fourth sampling rate based on the second interface processing flows that are more significant in terms of quality of experience than the first interface; Sending a signal indicating the fourth sampling rate to the network device, wherein the fourth sampling rate is higher than the second sampling rate; and Receiving a fifth sample of a flow sampled at the fourth sampling rate from the second interface.
5. The method according to claim 2, wherein, The first flow quantity is equal to the second flow quantity or within a predetermined threshold of the second flow quantity, and wherein the third sampling rate is lower than the second sampling rate.
6. The method according to claim 5, wherein The predetermined threshold is static or based on a flow quantity determined at one of a plurality of sampling rates.
7. The method according to claim 2, wherein The first sampling rate is higher than the second sampling rate, wherein the first flow quantity is greater than the second flow quantity, and wherein the third sampling rate is higher than the second sampling rate.
8. The method according to claim 7, wherein The third sampling rate is equal to the first sampling rate.
9. The method according to any one of claims 1 to 8, wherein The first sample, the second sample, and the third sample are sFlow samples.
10. The method according to any one of claims 1 to 8, wherein, The method repeats periodically.
11. A network device, comprising: A memory configured to store a plurality of sampling rates; A communication unit configured to send signals and receive samples of a data stream; And A processing circuit communicatively coupled to the memory and the communication unit, the processing circuit being configured to: Receive a first sample of a flow sampled at a first sampling rate from an interface of another network device; Determine a first parameter based on the first sample; Receive a second sample of a stream sampled at a second sampling rate from the interface, wherein the second sampling rate is different from the first sampling rate; Determine a second parameter based on the second sample; Determine a third sampling rate based on the first parameter and the second parameter; Control the communication unit to send a signal indicating the third sampling rate to the other network device; and Receive a third sample of a stream sampled at the third sampling rate from the interface.
12. The network device according to claim 11, wherein, The first parameter refers to a first flow quantity, and wherein the second parameter refers to a second flow quantity.
13. The network device according to claim 12, wherein, The interface is a first interface, and wherein the processing circuit is further configured to: Receive a fourth sample of a stream sampled at the first sampling rate from a second interface of the other network device; Determine a third flow quantity based on the fourth sample; Determine that the third flow quantity is equal to or greater than a predetermined threshold; Determine a fourth sampling rate based on the third flow quantity that is greater than or equal to the predetermined threshold; Control the communication unit to send a signal indicating the fourth sampling rate to the other network device, wherein the fourth sampling rate is higher than the second sampling rate; and Receive a fifth sample of a stream sampled at the fourth sampling rate from the second interface.
14. The network device according to claim 12, wherein, The interface is a first interface, and wherein the processing circuit is further configured to: Receive a fourth sample of a stream sampled at the first sampling rate from a second interface of the other network device; Determine that the second interface is processing a stream that is more significant in terms of quality of experience than the first interface; Determine a fourth sampling rate based on the second interface processing a stream that is more significant in terms of quality of experience than the first interface; Control the communication unit to send a signal indicating the fourth sampling rate to the other network device, wherein the fourth sampling rate is higher than the second sampling rate; and Receive a fifth sample of a stream sampled at the fourth sampling rate from the second interface.
15. The network device according to claim 12, wherein, The first flow quantity is equal to the second flow quantity or within a predetermined threshold of the second flow quantity, and wherein the third sampling rate is lower than the second sampling rate.
16. The network device according to claim 15, wherein, The predetermined threshold is static or based on a flow quantity determined at one of the plurality of sampling rates.
17. The network device according to claim 12, wherein, The first sampling rate is higher than the second sampling rate, wherein the first flow quantity is greater than the second flow quantity, and wherein the third sampling rate is higher than the second sampling rate.
18. The network device according to claim 12, wherein, The third sampling rate is equal to the first sampling rate.
19. The network device according to any one of claims 11 to 18, wherein, The first sample, the second sample, and the third sample are sFlow samples.
20. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Adaptive monitoring of telecommunications networks
CN103379002A