Underlayer-overlay correlation

By collecting and correlating the underlying streaming data and overlay streaming data in a virtualized data center, the challenges of network failure analysis in a virtualized environment are solved, enabling efficient troubleshooting and analysis.

CN120179718APending Publication Date: 2025-06-20HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249111.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2019-11-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In virtualized data centers, prior art faces challenges when analyzing and troubleshooting network operation failures, especially when identifying the physical path taken by data streams through the network.

Method used

By collecting and correlating underlying stream data and overlay stream data, you can gain insight into network operations and performance. The underlying stream data and samples covering stream data support high availability and large capacity stream data collection and respond when analyzing queries.

Benefits of technology

Efficient troubleshooting and analysis of virtualized networks is achieved, enabling faster identification of connectivity issues and reducing the number of underlying network devices associated with the issue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179718A_ABST
    Figure CN120179718A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to underlayer-coverage correlation. This disclosure describes techniques that include collecting underlying streaming data along with overlay streaming data within a network, and correlating the data to enable insight into network operation and performance. In one example, the disclosure describes a method comprising the steps of: collecting streaming data for a network, the network having a plurality of network devices and a plurality of virtual networks established within the network; storing the streaming data in a data repository; receiving an information request regarding the data stream, where the information request specifies a source virtual network for the data stream and also specifies a destination virtual network for the data stream; and querying the data repository with the specified source virtual network and the specified destination virtual network to identify one or more network devices that have processed at least one packet in the data stream based on the stored stream data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application date of November 6, 2019, application number 201911076051.1, and invention title "Underlay - Overlay Correlation". Technical Field

[0002] This disclosure relates to the analysis of computer networks, including analyzing the paths taken by data as it traverses the network. Background Art

[0003] Virtualized data centers are becoming a core foundation of modern information technology (IT) infrastructure. In particular, modern data centers have widely utilized virtualized environments in which virtual hosts such as virtual machines or containers are deployed and executed on the underlying computing platforms of physical computing devices.

[0004] Virtualization within large data centers can provide multiple advantages, including efficient use of computing resources and simplified network configuration. Thus, enterprise IT staff generally prefer these clusters not only because of the efficiency and improved return on investment (ROI) provided by virtualization, but also because of the management advantages of virtualized computing clusters within the data center. However, virtualization can pose some challenges when analyzing, evaluating, and / or troubleshooting network operation failures. Summary of the Invention

[0005] This disclosure describes techniques that include collecting information about physical network infrastructure (e.g., underlay flow data) and information about network virtualization (e.g., overlay flow data), and correlating the data to enable insights into network operation and performance. In some examples, samples of both underlay flow data and overlay flow data are collected and stored in a manner that not only supports high availability and high-volume flow data collection but also supports analyzing such data in response to analytical queries. Before or in response to such queries, the underlay flow data can be enriched, augmented, and / or supplemented with overlay flow data to enable insights into, identification of, and / or analysis of the underlying network infrastructure that may correspond to the overlay data flows. Graphs and other information illustrating which components of the underlying network infrastructure correspond to various overlay data flows can be presented in a user interface.

[0006] The techniques described herein can provide one or more technical advantages. For example, by providing information about how the underlying network infrastructure relates to various overlay data flows, it is possible to create useful tools for discovery and investigation. In some examples, such tools can be used for efficient and pipelined troubleshooting and analysis of virtualized networks. As an example, the techniques described herein can allow for more efficient troubleshooting of connectivity, at least because these techniques enable the identification of a significantly reduced number of underlying network devices that may be related to the connectivity issue.

[0007] In some examples, the present disclosure describes operations performed by a network analysis system or other network system in accordance with one or more aspects of the present disclosure. In a particular example, the present disclosure describes a method that includes the steps of: collecting flow data for a network that has a plurality of network devices and a plurality of virtual networks established within the network, where the flow data includes underlying flow data and overlay flow data, the underlying flow data includes a plurality of underlying data streams, the overlay flow data includes a plurality of overlay data streams, where the underlying flow data identifies, for each underlying data stream included within the underlying flow data, the network device that has processed network packets associated with the underlying data stream, and where the overlay flow data identifies, for each overlay data stream included within the overlay flow data, one or more of the virtual networks within the virtual network associated with the overlay data stream; storing the flow data in a data repository; receiving an information request regarding a data stream, where the information request specifies a source virtual network for the data stream and also specifies a destination virtual network for the data stream; and querying the data repository with the specified source virtual network and the specified destination virtual network to identify, based on the stored flow data, one or more network devices that have processed at least one packet within the data stream.

[0008] In another example, the present disclosure describes a system that includes processing circuitry configured to perform the operations described herein. In another example, the present disclosure describes a non-transitory computer-readable storage medium that includes instructions that, when executed, configure the processing circuitry of a computing system to perform the operations described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1A is a conceptual diagram illustrating an example network in accordance with one or more aspects of the present disclosure, the example network including a system for analyzing traffic flows across a network and / or within a data center.

[0010] Figure 1B is a conceptual diagram illustrating example components of a system for analyzing traffic flows across a network and / or within a data center in accordance with one or more aspects of the present disclosure.

[0011] Figure 2 is a block diagram illustrating an example network for analyzing traffic flows across a network and / or within a data center in accordance with one or more aspects of the present disclosure.

[0012] Figure 3 is a conceptual diagram illustrating an example query performed on stored underlying flow data and overlay flow data in accordance with one or more aspects of the present disclosure.

[0013] Figure 4is a conceptual diagram illustrating an example user interface presented by a user interface device in accordance with one or more aspects of the present disclosure.

[0014] Figure 5 is a flowchart illustrating operations performed by an example network analysis system in accordance with one or more aspects of the present disclosure. DETAILED DESCRIPTION

[0015] Data centers using virtualized environments offer efficiency, cost, and organizational advantages where virtual hosts such as virtual machines or virtual containers are deployed and executed on the underlying computing platforms of physical computing devices. However, gaining meaningful insights into application workloads is crucial for managing any data center fabric. Collecting traffic samples from networked devices can help provide such insights. In the various examples described herein, traffic samples are collected and then processed by analysis algorithms such that information about the overlay traffic can be correlated with the underlying infrastructure. In some examples, a user interface can be generated to support visualizing the collected data and how the underlying infrastructure relates to various overlay networks. The presentation of such data in the user interface can provide insights into the network and provide tools for network discovery, investigation, and troubleshooting to users, administrators, and / or other personnel.

[0016] Figure 1A is a conceptual diagram of an example network in accordance with one or more aspects of the present disclosure, the example network including a system for analyzing traffic flows across a network and / or within a data center. Figure 1A Illustrates an example implementation of network system 100 and data center 101, the data center 101 hosting one or more computing networks, computing domains or projects, and / or cloud-based computing networks, which are generally referred to herein as cloud computing clusters. Cloud-based computing clusters can co-reside in a common overall computing environment (such as a single data center) or be distributed across environments (such as across different data centers). Cloud-based computing clusters can be, for example, different cloud environments such as OpenStack cloud environments, Kubernetes cloud environments, or various combinations of other computing clusters, domains, networks, etc. In other instances, other implementations of network system 100 and data center 101 can be appropriate. Such implementations can include a subset of the components included in the Figure 1A example and / or can include additional components not shown in the Figure 1A example.

[0017] In the Figure 1A example, data center 101 provides an operating environment for applications and services to customer 104 coupled to data center 101 via service provider network 106. Although in conjunction with Figure 1AThe functions and operations described for network system 100 can be illustrated as being distributed across multiple devices in Figure 1A but in other examples, features and techniques attributable to one or more devices in Figure 1A can be performed internally by local components of one or more such devices. Similarly, one or more such devices can include specific components and perform various techniques that can otherwise be attributed to one or more other devices in the description herein. Further, specific operations, techniques, features, and / or functions can be described in conjunction with Figure 1A or otherwise performed by specific components, devices, and / or modules. In other examples, such operations, techniques, features, and / or functions can be performed by other components, devices, or modules. Thus, some operations, techniques, features, and / or functions attributed to one or more components, devices, or modules can be attributed to other components, devices, and / or modules even if not specifically described in this manner herein.

[0018] Data center 101 hosts infrastructure devices such as networking and storage systems, redundant power, and environmental controls. Service provider network 106 can be coupled to one or more networks managed by other providers and can thus form part of a large-scale public network infrastructure such as the Internet.

[0019] In some examples, data center 101 can represent one of many geographically distributed network data centers. As shown in the example of Figure 1A data center 101 is a facility that provides network services to customers 104. Customers 104 can be collective entities such as enterprises and governments or individuals. For example, a network data center can host network services for multiple enterprises and end users. Other exemplary services can include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. In some examples, data center 101 is a standalone network server, network peer, or other.

[0020] In Figure 1AIn the example, data center 101 includes a set of storage systems, application servers, computing nodes, or other devices, including network devices 110A through network device 110N (collectively referred to as "network devices 110", representing any number of network devices). Devices 110 can be interconnected via a high-speed switching fabric 121 provided by one or more layers of physical network switches and routers. In some examples, devices 110 can be included within fabric 121 but are shown separately for ease of illustration. Network devices 110 can be any one of many different types of network devices (core switches, backbone network devices, leaf network devices, edge network devices, or other network devices), but in some examples, one or more of devices 110 can act as physical computing nodes of the data center. For example, one or more of devices 110 can provide an operating environment for executing one or more customer-specific virtual machines or other virtualized instances (such as containers). In such examples, one or more of devices 110 can alternatively be referred to as host computing devices or more simply as hosts. Network devices 110 can thus execute one or more virtualized instances, such as virtual machines, containers, or other virtual execution environments for running one or more services such as virtualized network functions (VNFs).

[0021] Generally, each network device 110 can be any type of device that can operate on a network and can generate data (such as flow data or sFlow data) accessible via telemetry or other means, and this device can include any type of computing device, sensor, camera, node, monitoring device, or other device. Further, some or all of network devices 110 can represent components of another device, where such components can generate data that can be collected via telemetry or other means. For example, some or all of network devices 110 can represent physical or virtual network devices, such as switches, routers, hubs, gateways, security devices (such as firewalls), intrusion detection and / or intrusion prevention devices.

[0022] Although not specifically shown, switching fabric 121 can include top-of-rack (TOR) switches coupled to the distribution layer of a chassis switch, and data center 101 can include one or more non-edge switches, routers, hubs, gateways, security devices (such as firewalls), intrusion detection and / or intrusion prevention devices, servers, computer terminals, laptops, printers, databases, wireless mobile devices (such as cellular phones or personal digital assistants), wireless access points, bridges, cable modems, application accelerators, or other network devices. Switching fabric 121 can perform layer 3 routing to route network traffic between data center 101 and customer 104 via service provider network 106. Gateway 108 is used to forward and receive packets between switching fabric 121 and service provider network 106.

[0023] According to one or more examples of the present disclosure, a software-defined network (“SDN”) controller 132 provides a logically and in some cases physically centralized controller for facilitating the operation of one or more virtual networks within a data center 101. In some examples, the SDN controller 132 operates in response to configuration input operations received from an orchestration engine 130 via a northbound API 131, which in turn may respond to configuration input operations received from an administrator 128 interacting with and / or operating a user interface device 129.

[0024] The user interface device 129 may be implemented as any suitable device for presenting output and / or accepting user input. For example, the user interface device 129 may include a display. The user interface device 129 may be a computing system, such as a mobile or non-mobile computing device operated by a user and / or administrator 128. According to one or more aspects of the present disclosure, the user interface device 129 may, for example, represent a workstation, a laptop or notebook computer, a desktop computer, a tablet computer, or any other computing device that can be operated by a user and / or present a user interface. In some examples, the user interface device 129 may be physically separated from and / or located in a different location than the controller 201. In such examples, the user interface device 129 may communicate with the controller 201 via a network or other communication means. In other examples, the user interface device 129 may be a local peripheral of the controller 201 or may be integrated into the controller 201.

[0025] In some examples, the orchestration engine 130 manages the functions of the data center 101, such as compute, storage, networking, and application resources. For example, the orchestration engine 130 may create virtual networks for tenants within or across the data center 101. The orchestration engine 130 may attach virtual machines (VMs) to a tenant's virtual network. The orchestration engine 130 may connect a tenant's virtual network to an external network, such as the Internet or a VPN. The orchestration engine 130 may implement security policies across a group of VMs or at the boundary of a tenant network. The orchestration engine 130 may deploy network services (e.g., load balancers) within a tenant's virtual network.

[0026] In some examples, the SDN controller 132 manages network and networking services such as load balancing, security, and can allocate resources from the devices 110 acting as host devices to various applications via the southbound API 133. That is, the southbound API 133 represents a set of communication protocols that the SDN controller 132 uses to make the actual state of the network equal to the desired state specified by the orchestration engine 130. For example, the SDN controller 132 can implement high-level requests from the orchestration engine 130 by configuring: physical switches such as TOR switches, chassis switches, and fabric 121; physical routers; physical service nodes such as firewalls and load balancers; and virtual services such as virtual firewalls in VMs. The SDN controller 132 maintains routing, network, and configuration information in a state database.

[0027] The network analysis system 140 interacts with one or more of the devices 110 (and / or other devices) to collect flow data across the data center 101 and / or the network system 100. Such flow data can include underlying flow data and overlay flow data. In some examples, the underlying flow data can be collected by means of samples of flow data collected at layer 2 of the OSI model. The overlay flow data can be data (e.g., data samples) derived from overlay traffic across one or more virtual networks established within the network system 100. The overlay flow data can include, for example, information identifying the source virtual network and the destination virtual network.

[0028] According to one or more aspects of the present disclosure, Figure 1A the network analysis system 140 can configure each device 110 to collect flow data. For example, in the example that can be described with reference to Figure 1A , the network analysis system 140 outputs a signal to each device 110. Each device 110 receives the signal and interprets the signal as a command to collect flow data, which includes underlying flow data and / or overlay flow data. Thereafter, as data packets are processed by each device 110, each device 110 transmits the underlying flow data and / or overlay flow data to the network analysis system 140. The network analysis system 140 receives the flow data, prepares it for use in response to an analysis query, and stores the flow data. In Figure 1A the example, other network devices including network devices (not specifically shown) within the fabric 121 can also be configured to collect underlying and / or overlay flow data.

[0029] The network analysis system 140 can process queries. For example, in the described example, the user interface device 129 detects an input and outputs information about the input to the network analysis system 140. The network analysis system 140 determines that the information corresponds to a request for information about the network system 100 from a user of the user interface device 129. The network analysis system 140 processes the request by querying the stored flow data. The network analysis system 140 generates a response to the query based on the stored flow data and outputs information about the response to the user interface device 129.

[0030] In some examples, the request received from the user interface device 129 can include source and / or destination virtual networks. In such an example, the network analysis system 140 can, in response to such a request, identify one or more possible data paths on the underlying network devices that packets traveling from the source virtual network to the destination virtual network might have taken. To identify the possible data paths, the network analysis system 140 can correlate the collected overlay flow data with the collected underlying flow data such that the underlying network devices used by the overlay data flows can be identified.

[0031] Figure 1B is a conceptual diagram illustrating example components of a system for analyzing traffic flows across a network and / or within a data center in accordance with one or more aspects of the present disclosure. Figure 1B Includes many of the same elements described in conjunction with Figure 1A described. Figure 1B The elements shown may correspond to elements Figure 1A identified by the same reference numerals in Figure 1A shown. In general, such elements with the same reference numerals may be implemented in a manner consistent with the description of the corresponding elements provided in conjunction with Figure 1A but in some examples, such elements may relate to alternative implementations having more, fewer, and / or different capabilities and attributes.

[0032] However, different from Figure 1A Figure 1B ​Illustrates the components of network analysis system 140. Network analysis system 140 is shown as including a load balancer 141, a flow collector 142, a queue and event repository 143, a topology and metrics source 144, a data repository 145, and a flow API 146. Generally, network analysis system 140 and the components of network analysis system 140 are designed and / or configured to ensure high availability and the ability to handle large volumes of flow data. In some examples, multiple instances of the components of network analysis system 140 can be orchestrated (e.g., by an orchestration engine 130) to execute on different physical servers to ensure that no single component of network analysis system 140 has a single point of failure. In some examples, network analysis system 140 or its components can be scaled independently and horizontally to enable efficient and / or effective handling of traffic (e.g., flow data) of a desired capacity.

[0033] Same as Figure 1A in, Figure 1B the network analysis system 140 can configure each device 110 to collect flow data. For example, network analysis system 140 can output signals to each device 110 to configure each device 110 to collect flow data, which includes underlying flow data and overlay flow data. Subsequently, one or more of the devices 110 can collect the underlying flow data and overlay flow data and report such flow data to network analysis system 140.

[0034] In Figure 1B it is the load balancer 141 of network analysis system 140 that receives the flow data from each device 110. For example, in Figure 1B it, the load balancer 141 can receive the flow data from each device 110. The load balancer 141 can distribute traffic across multiple flow collectors to ensure an active / active failover strategy for the flow collectors. In some examples, multiple load balancers 141 may be required to ensure high availability and scalability.

[0035] The flow collector 142 collects data from the load balancer 141. For example, the flow collector 142 of the network analysis system 140 receives and processes flow packets (after being processed by the load balancer 141) from each device 110. The flow collector 142 sends the flow packets upstream to the queue and event repository 143. In some examples, the flow collector 142 can address, process, and / or accommodate unified data from sFlow, NetFlow v9, IPFIX, jFlow, Contrail Flow, and other formats. The flow collector 142 can be capable of parsing the internal headers from sFlow packets and other data flow packets. The flow collector 142 can be capable of handling message overflows, rich flow records with topology information (such as AppFormix topology information), etc. The flow collector 142 can also be capable of converting the data into binary format before writing or sending the data to the queue and event repository 143. The underlying flow data of the "sFlow" type (referring to "sampled flow") is a standard for packet output at layer 2 of the OSI model. It provides a means for outputting truncated packets and interface counters for network monitoring purposes.

[0036] The queue and event repository 143 processes the collected data. For example, the queue and event repository 143 can receive data from one or more flow collectors 142, store the data, and make the data available for ingestion in the data repository 145. In some examples, this enables the separation of the task of receiving and storing large volumes of data from the task of indexing the data and preparing it for analytical queries. In some examples, the queue and event repository 143 can also enable independent users to directly consume the stream of flow records. In some examples, the queue and event repository 143 can be used to detect anomalies and generate alerts in real time. In some examples, the flow data can be parsed by reading the encapsulated packets, which include VXLAN, MPLS over UDP, and MPLS over GRE. The queue and event repository 143 resolves the source IP, destination IP, source port, destination port, and protocol from the internal (underlying) packets. Some types of flow data (including sFlow data) only include fragments of the sampled network traffic (such as the first 128 bytes), so in some cases, the flow data may not include all internal fields. In such examples, this data can be marked as missing.

[0037] The topology and metrics source 144 can enrich or augment data with topology information and / or metrics information. For example, the topology and metrics source 144 can provide network topology metadata, which can include identified nodes or network devices, configuration information, configurations, established links, and other information about such nodes and / or network devices. In some examples, the topology and metrics source 144 can use AppFormix topology data, or can be an AppFormix module in execution. Information received from the topology and metrics source 144 can be used to enrich the flow data collected by the flow collector 142 and support queries to the data repository 145 by the flow API 146.

[0038] The data repository 145 can be configured to store data received from the queue and event repository 143 and the topology and metrics source 144 in an indexed format to support fast aggregation queries and fast random access data retrieval. In some examples, the data repository 145 can achieve fault tolerance and high availability by sharding and replicating data.

[0039] The flow API 146 can process query requests sent by one or more user interface devices 129. For example, in some examples, the flow API 146 can receive a query request from the user interface device 129 by means of an HTTP POST request. In such an example, the flow API 146 converts the information included in the request into a query to the data repository 145. To create the query, the flow API 146 can use the topology information from the topology and metrics source 144. The flow API 146 can use one or more of such queries to perform analysis on behalf of the user interface device 129. Such analysis can include traffic deduplication, overlay-underlay correlation, traffic path identification, and / or heatmap traffic calculation. In particular, such analysis can involve correlating the underlying flow data with the overlay flow data to support identifying which underlying network devices are associated with the traffic flowing on the virtual network and / or between two virtual machines.

[0040] By means of techniques according to one or more aspects of the present disclosure, such as by correlating the underlying flow data with the overlay flow data, the network analysis system 140 can be able to determine for a given data flow which tenant in the multi-tenant data center the data flow belongs to. Further, the network analysis system 140 can also be able to determine which virtual computing instances (e.g., virtual machines or containers) are the source and / or destination virtual computing instances of such a flow. Even further, correlating the underlying flow data with the overlay flow data, such as by enriching the underlying flow data with the overlay flow data, can facilitate troubleshooting of performance or other problems that may occur in the network system 100.

[0041] For example, in some cases, connectivity issues may arise during a specific time frame in which limited information is available but information about the source and destination virtual networks is known. Resolving such issues can be challenging because, given a source virtual network and a destination virtual network, it may be difficult to ascertain the physical path taken by the data flow through the network. Since it may additionally not be readily known what the actual physical path is through the underlying infrastructure, there may be many network devices or physical links that could be the underlying cause of the connectivity issue. However, by collecting underlying flow data and overlay flow data and enriching the underlying flow data with the overlay flow data collected during the same time period, it can be determined which underlying network devices processed the data flow and the physical links traversed by the data flow, enabling determination of the data path—or most likely or a set of likely data paths—taken by the data flow through the network, or at least determination of a smaller number of likely data paths for the data flow. Thus, resolving such connectivity issues can be significantly more efficient, at least because the number of underlying network devices associated with the connectivity issue can be substantially reduced.

[0042] Figure 2 is a block diagram illustrating an example network for analyzing traffic flows across a network and / or within a data center in accordance with one or more aspects of the present disclosure. Figure 2 The network system 200 can be described as Figure 1A or Figure 1B an example or alternative implementation of the network system 100. One or more aspects herein may be described in the context of FIG. 1. Figure 2 One or more aspects.

[0043] Although data centers, such as Figure 1A , Figure 1B and Figure 2 the data center shown, can be operated by any entity, some data centers are operated by service providers, where the business model of such service providers is to provide computing capabilities to their clients. For this reason, data centers typically contain a large number of computing nodes or host devices. For efficient operation, these hosts must be connected to each other and to the outside world, and this capability is provided by physical network devices that can be interconnected in a leaf-spine topology. The collection of these physical devices (such as network devices and hosts) forms the underlying network.

[0044] Each host device in such a data center typically runs multiple virtual machines thereon, which are referred to as workloads. Clients of the data center typically have access to these workloads and can use such workloads to install applications and perform other operations. Workloads running on different host devices but accessible by a particular client are organized into a virtual network. Each client typically has at least one virtual network. These virtual networks are also referred to as overlay networks. In some cases, a client of the data center may experience connectivity issues between two applications running on different workloads. Due to the deployment of workloads in a large multi-tenant data center, resolving such issues is often complex.

[0045] In Figure 2 the example of, network 205 connects network analysis system 240, host device 210A, host device 210B, and host device 210N. Network analysis system 240 may correspond to Figure 1A and Figure 1B the example or alternative implementation of network analysis system 140 shown. Host devices 210A, 210B to 210N may be collectively referred to as "host device 210", representing any number of host devices 210.

[0046] Each host device 210 may be Figure 1A and Figure 1B an example of device 110 of, but in Figure 2 the example of, each host device 210 is implemented as a server or host device operating as a computing node of a virtualized data center, as opposed to a network device. Thus, in Figure 2 the example of, each host device 210 executes multiple virtual computing instances, such as virtual machines 228.

[0047] Similar to Figure 1A and Figure 1B , a user interface device 129 that can be operated by an administrator 128 is also connected to network 205. In some examples, user interface device 129 may present one or more user interfaces at a display device associated with user interface device 129, some of which may have a form similar to user interface 400.

[0048] Figure 2 Underlying flow data 204 and overlay flow data 206 flowing within network system 200 are also illustrated. In particular, underlying flow data 204 leaving backbone device 202A and flowing towards network analysis system 240 is shown. Similarly, overlay flow data 206 leaving host device 210A and flowing across 205 is shown. In some examples, overlay flow data 206 is transmitted to network analysis system 240 via network 205, as described herein. For simplicity,Figure 2 Illustrated are a single instance of underlying flow data 204 and a single instance of overlay flow data 206. However, it should be understood that each of the backbone devices 202 and leaf devices 203 can generate the underlying flow data 204 and transmit it to the network analysis system 240, and in some examples, each host device 210 (and / or other devices) can generate the underlying flow data 204 and transmit such data across the network 205 to the network analysis system 240. Further, it should be understood that each host device 210 (and / or other devices) can generate the overlay flow data 206 and transmit such data across the network 205 to the network analysis system 240.

[0049] The network 205 can correspond to Figure 1A and Figure 1B either the switching fabric 121 and / or any one of the service provider networks 106, or alternatively, can correspond to a combination of the switching fabric 121, the service provider network 106, and / or another network. The network 205 can also include Figure 1A and Figure 1B some components of

[0050] Illustrated within the network 205 are backbone devices 202A and 202B (collectively referred to as "backbone devices 202" and representing any number of backbone devices 202), and leaf devices 203A, 203B, and leaf device 203C (collectively referred to as "leaf devices 203" and also representing any number of leaf devices 203). Although the network 205 is illustrated as having backbone devices 202 and leaf devices 203, other types of network devices can be included in the network 205, including core switches, edge network devices, top-of-rack devices, and other network devices.

[0051] Generally, the network 205 can be the Internet, or can include or represent any public or private communication network or other network. For example, the network 205 can be a cellular network, ZigBee, Bluetooth, near field communication (NFC), satellite, enterprise, service provider, and / or other types of networks that support the transfer of traffic data between computing systems, servers, and computing devices. One or more of the client devices, server devices, or other devices can use any suitable communication technology to transmit and receive data, commands, control signals, and / or other information across the network 205. The network 205 can include one or more network hubs, network switches, network routers, satellite dishes, or any other network devices. Such devices or components can be operatively coupled to each other to provide information exchange between computers, devices, or other components (e.g., between one or more client devices or systems and one or more server devices or systems).Figure 2 Each device or system shown can be operably coupled to network 205 using one or more network links. The link coupling such a device or system to network 205 can be Ethernet, Asynchronous Transfer Mode (ATM), or other types of network connections, and such connections can be wireless and / or wired connections. Figure 2 One or more devices or systems shown or otherwise on network 205 can be in a remote location relative to one or more other devices or systems shown.

[0052] The network analysis system 240 can be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, appliances, cloud computing systems, and / or other computing systems capable of performing the operations and / or functions described in accordance with one or more aspects of the present disclosure. In some examples, the network analysis system 240 represents a cloud computing system, server farm, and / or server cluster (or a portion thereof) that provides services to client devices and other devices or systems. In other examples, the network analysis system 240 can represent one or more virtualized computing instances (e.g., virtual machines, containers) of a data center, cloud computing system, server farm, and / or server cluster or be implemented therewith.

[0053] In Figure 2 an example, the network analysis system 240 can include a power supply 241, one or more processors 243, one or more communication units 245, one or more input devices 246, and one or more output devices 247. The storage device 250 can include one or more collector modules 252, a user interface module 254, a flow API 256, and a data repository 259.

[0054] One or more of the devices, modules, storage areas, or other components of the network analysis system 240 can be interconnected to enable communication (physically, communicatively, and / or operatively) between components. In some examples, such connectivity can be provided by a communication channel (e.g., communication channel 242), a system bus, a network connection, an interprocess communication data structure, or any other method for transferring data.

[0055] Power supply 241 can supply power to one or more components of network analysis system 240. Power supply 241 can receive power from a main alternating current (AC) power supply in a data center, building, home, or other location. In other examples, power supply 241 can be a battery or a device that supplies direct current (DC). In still other examples, network analysis system 240 and / or power supply 241 can receive power from another source. One or more devices or components illustrated within network analysis system 240 can be connected to power supply 241 and / or can receive power from power supply 241. Power supply 241 can have intelligent power management or consumption capabilities, and such features can be controlled, accessed, or regulated by one or more modules of network analysis system 240 and / or by one or more processors 243 to intelligently consume, distribute, supply, or otherwise manage power.

[0056] One or more processors 243 of network analysis system 240 can implement functionality and / or execute instructions associated with network analysis system 240 or with one or more modules illustrated and / or described herein. One or more processors 243 can be a processing circuit that performs operations according to one or more aspects of the present disclosure, can be a part of, and / or can include the processing circuit. Examples of processors 243 include microprocessors, application processors, display controllers, auxiliary processors, one or more sensor hubs, and any other hardware configured to act as a processor, processing unit, or processing device. Central monitoring system 210 can use one or more processors 243 to perform operations according to one or more aspects of the present disclosure using software, hardware, firmware, or a combination of hardware, software, and firmware residing within and / or executed at network analysis system 240.

[0057] One or more communication units 245 of network analysis system 240 can communicate with devices external to network analysis system 240 by transmitting and / or receiving data, and in some aspects can operate as both an input device and an output device. In some examples, communication unit 245 can communicate with other devices over a network. In other examples, communication unit 245 can transmit and / or receive radio signals over a radio network such as a cellular radio network. Examples of communication units 245 include network interface cards (e.g., such as Ethernet cards), optical transceivers, radio frequency transceivers, GPS receivers, or any other type of device that can transmit and / or receive information. Other examples of communication unit 245 can include devices capable of communicating via GPS, NFC, ZigBee, and cellular networks (e.g., 3G, 4G, 5G), as well as those found in mobile devices and universal serial bus (USB) controllers, etc. Radio transceiver device. Such communication may embrace, implement, or comply with appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, Bluetooth, NFC, or other technologies or protocols.

[0058] One or more input devices 246 may represent any input device of the network analysis system 240 not otherwise separately described herein. One or more input devices 246 may generate, receive, and / or process input from any type of device capable of detecting input from a person or machine. For example, one or more input devices 246 may generate, receive, and / or process input in the form of electrical, physical, audio, image, and / or visual input (e.g., peripherals, keyboards, microphones, cameras).

[0059] One or more output devices 247 may represent any output device of the network analysis system 240 not otherwise separately described herein. One or more output devices 247 may generate, receive, and / or process input from any type of device capable of detecting input from a person or machine. For example, one or more output devices 247 may generate, receive, and / or process output in the form of electrical and / or physical output (e.g., peripherals, actuators).

[0060] One or more storage devices 250 within the network analysis system 240 may store information for processing during operation of the network analysis system 240. The storage device 250 may store program instructions and / or data associated with one or more modules described in accordance with one or more aspects of the present disclosure. One or more processors 243 and one or more storage devices 250 may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 243 may execute instructions, and one or more storage devices 250 may store the instructions and / or data of one or more modules. The combination of the processor 243 and the storage device 250 may retrieve, store, and / or execute the instructions and / or data of one or more applications, modules, or software. The processor 243 and / or the storage device 250 may also be operatively coupled to one or more other software and / or hardware components, including but not limited to one or more components of the network analysis system 240 and / or one or more devices or systems illustrated as being connected to the network analysis system 240.

[0061] In some examples, one or more storage devices 250 are implemented with a temporary memory, which can mean that the primary purpose of one or more storage devices is not long-term storage. The storage device 250 of the network analysis system 240 can be configured to store short-term information as volatile memory, and thus does not retain the stored content when deactivated. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art. In some examples, the storage device 250 further includes one or more computer-readable storage media. The storage device 250 can be configured to store a larger amount of information than volatile memory. The storage device 250 can also be configured for long-term information storage as non-volatile storage space and retains information after an activation / shutdown cycle. Examples of non-volatile memory include magnetic hard disks, optical disks, flash memory, or forms of electrically programmable memory (EPROM) or electrically erasable programmable (EEPROM) memory.

[0062] The collector module 252 can perform functions related to receiving both the underlying stream data 204 and the overlay stream data 206 and performing load balancing as needed to ensure high availability, throughput, and scalability for collecting such stream data. The collector module 252 can process the data and prepare it for storage in the data repository 259. In some examples, the collector module 252 can store the data in the data repository 259.

[0063] The user interface module 254 can perform functions related to generating a user interface for presenting the results of the analytical queries performed by the stream API 256. In some examples, the user interface module 254 can generate information sufficient to generate a set of user interfaces and cause the communication unit 215 to output such information over the network 205 for use by the user interface device 129 to present one or more user interfaces at a display device associated with the user interface device 129.

[0064] The stream API 256 can perform analytical queries related to the data stored in the data repository 259, which is derived from the collection of the underlying stream data 204 and the overlay stream data 206. In some examples, the stream API 256 can receive requests in the form of information originating from an HTTP POST request and, in response, convert the request into a query to be executed on the data repository 259. Further, in some examples, the stream API 256 can obtain topology information related to the device 110 and perform analyses including data deduplication, overlay-underlying correlation, traffic path identification, and heatmap traffic calculation.

[0065] The data repository 259 can represent any suitable data structure or storage medium for storing information related to data stream information (including the storage of data derived from the underlying stream data 204 and the overlay stream data 206). The data repository 259 can be responsible for storing data in an indexed format, which enables rapid retrieval of data and execution of queries. The information stored in the data repository 259 can be searchable and / or categorized such that one or more modules within the network analysis system 240 can provide an input requesting information from the data repository 259 and, in response to the input, receive the information stored within the data repository 259. The data repository 259 can be maintained primarily by the collector module 252. The data repository 259 can be implemented with multiple hardware devices and can achieve fault tolerance and high availability by sharding and replicating data. In some examples, the data repository 259 can be implemented using the open-source ClickHouse column-oriented database management system.

[0066] Each host device 210 represents a physical computing device or computing node that provides an execution environment for virtual hosts, virtual machines, containers, and / or other virtualized computing resources. In some examples, each host device 210 can be a component of a cloud computing system, a server farm, and / or a server cluster (or a portion thereof) that provides services to client devices and other devices or systems.

[0067] Specific aspects of the host device 210 are described herein with respect to the host device 210A. Other host devices 210 (e.g., host devices 210B through 210N) can be described similarly and can also include the same, similar, or corresponding components, devices, modules, functionality, and / or other features. Accordingly, the description herein with respect to the host device 210A can be correspondingly applied to one or more other host devices 210 (e.g., host devices 210B through host device 210N).

[0068] In Figure 2In the example of, host device 210A includes basic physical computing hardware, which includes a power supply 211, one or more processors 213, one or more communication units 215, one or more input devices 216, one or more output devices 217, and one or more storage devices 220. The storage device 220 may include a hypervisor 221, which includes a kernel module 222, a virtual router module 224, and a proxy module 226. Virtual machines 228A through 228N (collectively referred to as "virtual machines 228," representing any number of virtual machines 228) execute on or are controlled by the hypervisor 221. Similarly, the virtual router agent 229 may execute on or under the control of the hypervisor 221. One or more of the devices, modules, storage areas, or other components of the host device 210 may be interconnected to enable communication between components (physically, communicatively, and / or operationally). In some examples, such connectivity may be provided via a communication channel (e.g., communication channel 212), a system bus, a network connection, an interprocess communication data structure, or any other method for transferring data.

[0069] The power supply 211 may power one or more components of the host device 210. The processor 213 may implement functionality and / or execute instructions associated with the host device 210. The communication unit 215 may communicate on behalf of the host device 210 with other devices or systems. The one or more input devices 216 and output devices 217 may represent any other input and / or output devices associated with the host device 210. The storage device 220 may store information for processing during operation of the host device 210A. Each of such components may be implemented in a manner similar to the manner described herein in connection with the network analysis system 240 or otherwise.

[0070] The hypervisor 221 can act as a module or system that instantiates, creates, and / or executes one or more virtual machines 228 on a base host hardware device. In some contexts, the hypervisor 221 can be referred to as a virtual machine manager (VMM). The hypervisor 221 can execute within an execution environment provided by a storage device 220 and a processor 213 or on an operating system kernel (e.g., kernel module 222). In some examples, the hypervisor 221 is an operating system-level component that executes on a hardware platform (e.g., host 210) to provide a virtualized operating environment and orchestration controller for the virtual machines 228 and / or other types of virtual computing instances. In other examples, the hypervisor 221 can be a software and / or firmware layer that provides a lightweight kernel and operates to provide a virtualized operating environment and orchestration controller for the virtual machines 228 and / or other types of virtual computing instances. The hypervisor 221 can incorporate the functionality of the kernel module 222 (e.g., as a "type 1 hypervisor"), as Figure 2 shown. In other examples, the hypervisor 221 can execute on the kernel (e.g., as a "type 2 hypervisor").

[0071] The virtual router module 224 can execute multiple routing instances for corresponding virtual networks within the data center 101 and can route packets to the appropriate virtual machines executing within the operating environment provided by the device 110. The virtual router module 224 can also be responsible for collecting overlay flow data, such as Contrail Flow data when used in an infrastructure employing Contrail SDN. Thus, each host device 210 can include a virtual router. Packets received by the virtual router module 224 of the host device 210A, for example, from the underlying physical network fabric can include an outer header to allow the physical network fabric to tunnel the payload or "inner packet" to the physical network address of the network interface for the host device 210A. The outer header can include not only the physical network address of the server network interface but also a virtual network identifier, such as a VxLAN tag or a Multiprotocol Label Switching (MPLS) tag that identifies one of the virtual networks and the corresponding routing instance executed by the virtual router. The inner packet includes an inner header that has a destination network address that conforms to the virtual network addressing space for the virtual network identified by the virtual network identifier.

[0072] The proxy module 226 can execute as part of the hypervisor 221 or can execute within the kernel space or as part of the kernel module 222. The proxy module 226 can monitor some or all of the performance metrics associated with the host device 210A and can implement and / or execute policies that can be received from a policy controller ( Figure 2A received policy (not shown in the figure) is received. The proxy module 226 may configure the virtual router module 224 to transmit the overlay flow data to the network analysis system 240.

[0073] Virtual machines 228A through 228N (collectively referred to as "virtual machines 228," representing any number of virtual machines 228) may represent example instances of virtual machines 228. The host device 210A may divide the virtual and / or physical address space provided by the storage device 220 into a user space for running user processes. The host device 210A may also divide the virtual and / or physical address space provided by the storage device 220 into a kernel space, which is protected and inaccessible to user processes.

[0074] Generally, each virtual machine 228 may be any type of software application, and a virtual address may be assigned to each virtual machine for use within a corresponding virtual network, where each virtual network may be a different virtual subnet provided by the virtual router module 224. For example, its own virtual layer three (L3) IP address may be assigned to each virtual machine 228 for sending and receiving communications, but each virtual machine does not know the IP address of the physical server on which the virtual machine is executed. In this way, a "virtual address" is an address for an application that is different from the logical address for the underlying physical computer system (such as Figure 2 the host device 210A in the example).

[0075] Each virtual machine 228 may represent a tenant virtual machine running a customer application (such as a web server, a database server, an enterprise application) or hosting a virtualized service for creating a service chain. In some cases, any one or more of the host devices 210 or another computing device directly (i.e., not as a virtual machine) hosts a customer application. Although one or more aspects of the present disclosure are described in view of virtual machines or virtual hosts, the techniques according to one or more aspects of the present disclosure described herein for such virtual machines or virtual hosts may also be applied to containers, applications, processes, or other execution units (virtualized or non-virtualized) executing on the host device 210.

[0076] The virtual router agent 229 is in Figure 2is included within the host device 210A in the example and can communicate with the SDN controller 132 and the virtual router module 224 to control the overlay of the virtual network and coordinate the routing of data packets within the host device 210A. In general, the virtual router agent 229 communicates with the SDN controller 132, which generates commands that control the routing of control packets through the data center 101. The virtual router agent 229 can be executed in user space and operates as a proxy for control plane messages between the virtual machine 228 and the SDN controller 132. For example, the virtual machine 228A can request to send a message using its virtual address via the virtual router agent 229, and the virtual router agent 229 can in turn send the message and request to receive a response to the message for the virtual address of the virtual machine 228A, which initiates the first message. In some cases, the virtual machine 228A can call a procedure or function call presented by the application programming interface of the virtual router agent 229, and in such an example, the virtual router agent 229 also processes the encapsulation of the message, including addressing.

[0077] The network analysis system 240 can configure each of the spine devices 202 and the leaf devices 203 to collect the underlying flow data 204. For example, in the example that can be referred to Figure 2 described, the collector module 252 of the network analysis system 240 causes the communication unit 215 to output one or more signals over the network 205. Each of the spine devices 202 and the leaf devices 203 detects the signal and interprets the signal as a command that supports the collection of the underlying flow data 204. For example, when detecting a signal from the network analysis system 240, the spine device 202A configures itself to collect sFlow data and transmits the sFlow data (as the underlying flow data 204) over the network 205 to the network analysis system 240. As another example, when detecting a signal from the network analysis system 240, the leaf device 203A detects the signal and configures itself to collect sFlow data and transmits the sFlow data over the network 205 to the network analysis system 240. Further, in some examples, each host device 210 can detect a signal from the network analysis system 240 and interpret the signal as a command that supports the collection of sFlow data. Thus, in some examples, the sFlow data can be collected by the collector module executed on the host device 210.

[0078] Thus, in the described example, the spine device 202, the leaf device 203 (and possibly one or more of the host devices 210) collect sFlow data. However, in other examples, one or more such devices may collect other types of underlying flow data 204, such as IPFIX and / or NetFlow data. Collecting any such underlying flow data may involve collecting five-tuple data that includes a source IP address and a destination IP address, a source port number and a destination port number, and the network protocol used.

[0079] The network analysis system 240 may configure each host device 210 to collect overlay flow data 206. For example, continuing with the example Figure 2 described, the collector module 252 causes the communication unit 215 to output one or more signals over the network 205. Each host device 210 detects the signal, which is interpreted as a command to collect overlay flow data 206 and transmit the overlay flow data 206 to the network analysis system 240. For example, referring to host device 210A, the communication unit 215 of host device 210A detects a signal over the network 205 and outputs information about the signal to the hypervisor 221. The hypervisor 221 outputs the information to the agent module 226. The agent module 226 interprets the information from the hypervisor 221 as a command to collect overlay flow data 206. The agent module 226 configures the virtual router module 224 to collect overlay flow data 206 and transmit the overlay flow data 206 to the network analysis system 240.

[0080] In at least some examples, the overlay flow data 206 includes five-tuple information about the source and destination addresses, ports, and protocols. Additionally, the overlay flow data 206 may include information about the virtual network associated with the flow, the virtual network including a source virtual network and a destination virtual network. In some examples, particularly for networks configured with Contrail SDN available from Juniper Networks, Inc., of Sunnyvale, California, the overlay flow data 206 may correspond to Contrail Flow data.

[0081] In the described example, the proxy module 226 configures the virtual router module 224 to collect overlay flow data 206. However, in other examples, the hypervisor 221 can configure the virtual router module 224 to collect overlay flow data 206. Further, in other examples, the overlay flow data 206 can be collected by another module (alternatively or additionally), such as the proxy module 226, or even by the hypervisor 221 or the kernel module 222. Thus, in some examples, the host device 210 can collect both the underlying flow data (sFlow data) and the overlay flow data (e.g., Contrail Flow data).

[0082] The network analysis system 240 can receive both the underlying flow data 204 and the overlay flow data 206. For example, continuing this example and referring to Figure 2 , the spine device 202A samples, detects, senses, and / or collects the underlying flow data 204. The spine device 202A outputs a signal over the network 205. The communication unit 215 of the network analysis system 240 detects the signal from the spine device 202A and outputs information about the signal to the collector module 252. The collector module 252 determines that the signal includes information about the underlying flow data 204.

[0083] Similarly, the virtual router module 224 of the host device 210A samples, detects, senses, and / or collects the overlay flow data 206 at the host device 210A. The virtual router module 224 causes the communication unit 215 of the host device 210A to output a signal over the network 205. The communication unit 215 of the network analysis system 240 detects the signal from the host device 210A and outputs information about the signal to the collector module 252. The collector module 252 determines that the signal includes information about the overlay flow data 206.

[0084] The network analysis system 240 can process both the underlying flow data 204 and the overlay flow data 206 received from various devices within the network system 100. For example, still continuing the same example, the collector module 252 processes the signals received from the spine device 202A, the host device 210A, and other devices by distributing the signals across multiple collector modules 252. In some examples, each collector module 252 can execute on a different physical server and can scale independently and horizontally to handle the desired or peak capacity of traffic from the spine device 202, the leaf device 203, and the host device 210. Each collector module 252 stores each instance of the underlying flow data 204 and the overlay flow data 206 and makes the stored data available for ingestion in the data repository 259. The collector module 252 indexes the data and prepares the data for use in analysis queries.

[0085] The network analysis system 240 can store the underlying flow data 204 and the overlay flow data 206 in the data repository 259. For example, in Figure 2 , the collector module 252 outputs information to the data repository 259. The data repository 259 determines that the information corresponds to the underlying flow data 204 and the overlay flow data 206. The data repository 259 stores the data in an indexed format to support fast aggregation queries and fast random access data retrieval. In some examples, the data repository 145 can achieve fault tolerance and high availability by sharding and replicating the data across multiple storage devices, which can be located among multiple physical hosts.

[0086] The network analysis system 240 can receive a query. For example, continuing with the same example and referring to Figure 2 , the user interface device 129 detects the input and outputs a signal derived from the input over the network 205. The communication unit 215 of the network analysis system 240 detects the signal and outputs information about the signal to the flow API 256. The flow API 256 determines that the signal corresponds to a query by a user of the user interface device 129 for information about the network system 200 within a given time window. For example, a user of the user interface device 129 (e.g., the administrator 128) may have noticed that a particular virtual machine within a particular virtual network appears to be dropping packets at an abnormal rate and may seek to resolve the issue. One way to resolve the issue is to identify which network devices (e.g., which underlying router) are on the data path that appears to be dropping packets. Thus, the administrator 128 may seek to identify the possible paths taken between the source virtual machine and the destination virtual machine by querying the network analysis system 240.

[0087] The network analysis system 240 can process the query. For example, again continuing in Figure 2In the context of the example described, the flow API 256 determines that the signal received from the user interface device 129 includes information about the source virtual network and / or the destination virtual network. The flow API 256 queries the data repository 259 by enriching the underlying flow data stored within the data repository 259 from the identified time window in the query to include virtual network data from the overlay flow data. To perform the query, the flow API 256 narrows the data to the specified time window, and for each relevant underlying flow data 204 record, the flow API 256 adds any source virtual network and / or destination virtual network information from the overlay flow data 206 record that has a value matching the value of the corresponding underlying flow data 204 record. The flow API 256 identifies one or more network devices identified by the enriched underlying flow data. The flow API 256 determines one or more possible paths to be taken between the specified source virtual network and the destination virtual network based on the identified network devices. In some examples, a global join technique (e.g., available in the ClickHouse database management system) can be used for augmentation. In such an example, the flow API 256 collects the overlay flow data and broadcasts this data to all nodes. The data is then used as a lookup table independently for each node. To minimize the size of the table, the flow API 256 can perform filter criteria pushdown to the predicates of the subqueries.

[0088] The network analysis system 240 can cause a user interface depicting the possible paths between the source virtual network and the destination virtual network to be presented at the user interface device 129. The flow API 256 outputs information about the determined possible paths to the user interface module 254. The user interface module 254 uses the information from the flow API 256 to generate data sufficient to create a user interface that presents information about the possible paths between the source virtual network and the destination virtual network. The user interface module 254 causes the communication unit 215 to output a signal over the network 205. The user interface device 129 detects the signal over the network 205 and determines that the signal includes information sufficient to generate a user interface. The user interface device 129 generates a user interface (e.g., user interface 400) and presents it at a display associated with the user interface device 129. In some examples, the user interface 400 (also illustrated in Figure 4 presents information depicting one or more possible paths between virtual machines and can include information about how much data has been transferred or is being transferred between these virtual machines.

[0089] Figure 2The modules shown (e.g., virtual router module 224, proxy module 226, collector module 252, user interface module 254, flow API 256) and / or the modules illustrated or described elsewhere in this disclosure may perform operations described using software, hardware, firmware, or a hybrid of hardware, software, and firmware residing in and / or executed at one or more computing devices. For example, a computing device may execute one or more such modules using multiple processors or multiple devices. A computing device may execute one or more such modules as a virtual machine executing on underlying hardware. One or more such modules may execute as one or more services of an operating system or computing platform. One or more such modules may execute as one or more executable programs at the application layer of a computing platform. In other examples, the functionality provided by a module may be implemented by a dedicated hardware device.

[0090] Although specific modules, data repositories, components, programs, executable files, data items, functional units, and / or other items included within one or more storage devices may be illustrated separately, one or more of such items may be combined and operate as a single module, component, program, executable file, data item, or functional unit. For example, one or more modules or data repositories may be combined or partially combined such that they operate as a single module or provide functionality. Further, one or more modules may interact with and / or cooperate with each other such that, for example, one module acts as a service or extension of another module. Moreover, each module, data repository, component, program, executable file, data item, functional unit, or other item illustrated within a storage device may include multiple components, sub-components, modules, sub-modules, data repositories, and / or other components or modules or data repositories not illustrated.

[0091] Further, each module, data repository, component, program, executable file, data item, functional unit, or other item illustrated within a storage device may be implemented in various ways. For example, each module, data repository, component, program, executable file, data item, functional unit, or other item illustrated within a storage device may be implemented as a downloadable or pre-installed application or “app”. In other examples, each module, data repository, component, program, executable file, data item, functional unit, or other item illustrated within a storage device may be implemented as part of an operating system executing on a computing device.

[0092] Figure 3 is a conceptual diagram illustrating example queries performed on stored underlying flow data and overlay flow data in accordance with one or more aspects of the present disclosure. Figure 3 Illustrates data table 301, query 302, and output table 303. Data table 301 illustrates what may be stored inFigure 2 Records of both the underlying flow data and the overlay flow data within the data repository 259. The query 302 represents a query that can be generated by the flow API 256 in response to a request received by the network analysis system 240 from the user interface device 129. The output table 303 represents data generated from the data table 301 in response to the execution of the query 302.

[0093] In Figure 3 the example of Figure 2 and in accordance with one or more aspects of the present disclosure, Figure 2 and Figure 3 both, the network analysis system 240 collects both the underlying flow data 204 and the overlay flow data 206 from various devices within the network system 200. The collector module 252 of the network analysis system 240 stores the collected data within the data repository 259. In Figure 3 the example of

[0094] the network analysis system 240 can execute a query after the data corresponding to the data table 301 is stored within the data repository 259. For example, still referring to Figure 2 and Figure 3 the example of Figure 3 the communication unit 215 of the network analysis system 240 detects a signal determined by the flow API 256 to correspond to the query 302. In

[0095] the example of

[0096] the query 302 is the following SQL-like query:

[0095] SELECT networkDevice, bytes, srcVn WHERE timestamp in <9;9>.

[0096] The flow API 256 applies the query 302 to the data table 301, thereby selecting rows from the data table 301 that identify network devices and have timestamps greater than or equal to 7 and less than or equal to 9. The flow API 256 only identifies two network devices that meet this criterion: network device "a7" (from row 3 of the data table 301) and network device "a8" (row 5 of the data table 301).

[0097] The network analysis system 240 can correlate the overlay data with the underlying data to identify which source virtual networks used the identified network devices during the relevant time frame. For example, in Figure 3In the example, the flow API 256 of the network analysis system 240 determines whether any overlay flow data rows in the same time frame (timestamps 7 - 9) have the same five - tuple data as rows 3 and 5 (i.e., source and destination addresses and port numbers, and protocol). The flow API 256 determines that the overlay data from row 6 is within the specified time frame and has five - tuple data that matches the five - tuple data of row 3. Thus, the flow API 256 determines that the source virtual network for device “a7” is source virtual network “e” (see the “srcvn” column in row 1 of output table 303). Similarly, the flow API 256 determines that the overlay data from row 4 of data table 301 is within the specified time frame and has five - tuple data that matches the five - tuple data of row 5 of data table 301. Thus, the flow API 256 determines that the source virtual network for device “a8” is source virtual network “c” (see row 2 of output table 303).

[0098] If more than one instance (row) of overlay flow data is available, any or all such data can be used to identify the source virtual network. This is based on the assumption that virtual network configurations do not change frequently. The enrichment process described herein can be used for queries that request “the top N” network attributes. The enrichment process can also be used to identify paths, as Figure 4 shown.

[0099] Figure 4 is a conceptual diagram illustrating an example user interface presented by a user interface device in accordance with one or more aspects of the present disclosure. Figure 4 User interface 400 is illustrated. Although user interface 400 is shown as a graphical user interface, in other examples other types of interfaces can be presented, including text - based user interfaces, console - or command - based user interfaces, voice - prompted user interfaces, or any other suitable user interface. As Figure 4 shown, user interface 400 can correspond to a user interface generated by the user interface module 254 of the network analysis system 240 and presented at Figure 2 the user interface device 129. One or more aspects related to the generation and / or presentation of user interface 400 can be described herein in the context of Figure 2 .

[0100] In accordance with one or more aspects of the present disclosure, the network analysis system 240 can execute queries to identify paths. For example, in the example that can be referred to Figure 2 the user interface device 129 detects an input and outputs a signal via network 205. The communication unit 215 of the network analysis system 240 detects the signal corresponding to the query that the flow API 256 determines corresponds to network information. The flow API 256 executes the query (e.g., in conjunction with Figure 3in the described manner) and outputs information about the result to the user interface module 254. To find a path between two virtual machines, the flow API 256 can determine the most likely path (and the traffic traveling through the determined path). Additionally, the flow API 256 can perform additional queries to exclusively evaluate the overlay data flow to identify traffic registered on the virtual router module 224, enabling the identification and display of traffic between the relevant virtual machines and host devices. The flow API 256 can identify host-virtual machine and virtual machine-host paths in a similar manner.

[0101] The network analysis system 240 can generate a user interface for presentation at a display device, such as user interface 400. For example, still referring to Figure 2 and Figure 4 , the user interface module 254 generates the information underlying user interface 400 and causes the communication unit 215 to output a signal over network 205. The user interface device 129 detects the signal and determines that the signal includes information sufficient to present the user interface. The user interface device 129 presents the user interface 400 at a display device associated with the user interface device 129 in the Figure 4 manner shown.

[0102] In Figure 4 , the user interface 400 is presented within the display window 401. The user interface 400 includes a sidebar region 404, a main display region 406, and an options region 408. The sidebar region 404 provides an indication of which user interface mode is being presented within the user interface 400, which, in the Figure 4 example, corresponds to the "Structure" mode. Other modes may be available for other network analysis scenarios as appropriate. Along the top of the main display region 406 is a navigation interface component 427, which can also be used to select the type or mode of network analysis to perform. The status notification display element 428 can provide information about alerts or other status information related to one or more networks, users, elements, or resources.

[0103] The main display region 406 presents a network diagram and can provide the topology of the various network devices included within the network being analyzed. In the Figure 4 example shown, the network is illustrated as having network devices, edge network devices, hosts, and instances, as indicated by the "Legend" along the bottom of the main display region 406. The actual or potential data paths between network devices and other components are illustrated within the main display region 406. Although in Figure 4A limited number of different types of network devices and components are shown, but in other examples, other types of devices or components or elements may be presented and / or specifically illustrated, including core switch devices, backbone devices, leaf devices, physical or virtual routers, virtual machines, containers, and / or other devices, components, or elements. Further, some data paths or components (e.g., instances) of the network may be hidden or minimized within the user interface 400 to facilitate the illustration and / or presentation of the components or data paths most relevant to a given network analysis.

[0104] The option area 408 provides, along the right hand side of the user interface 400, a plurality of input fields related to the underlying network being analyzed (e.g., underlying quintuple input fields) and a plurality of input fields related to the overlay network being analyzed (e.g., source and destination virtual networks and IP address input fields). The user interface 400 accepts input via user interaction with one or more of the displayed input fields, and based on the data entered into the input fields, the user interface module 254 presents response information regarding the network being analyzed.

[0105] For example, in Figure 4 the example of, the user interface 400 accepts input in the option area 408 regarding a specific time frame (e.g., time range), source and destination virtual networks, and source and destination IP addresses. In the example shown, the underlying information in the user interface 400 has not been specified by user input. Using the input already provided in the option area 408, the network analysis system 240 determines information regarding one or more possible data paths (e.g., the most likely data path) through the underlying network devices. The network analysis system 240 determines such possible data paths during the time range specified in the option area 408, based on data collected by the network analysis system 240 (e.g., the collector module 252). The user interface module 254 of the network analysis system 240 generates data in support of the presentation of the user interface 400, wherein one possible data path is highlighted (by drawing each segment of the data path in bold), as Figure 4 shown. In some examples, more than one data path from the source virtual network to the destination virtual network may be highlighted. Further, in some examples, a heat map color scheme may be used to present one or more data paths in the main display area 406, meaning that the data paths are illustrated with a color (or shade of gray) corresponding to the amount of data being communicated through the path or corresponding to the degree of use of the corresponding path. While Figure 4The data paths are illustrated using a heat map color (or grayscale shading) scheme, but in other examples, data regarding utilization or traffic on or through a network device may be presented in other suitable ways (e.g., applying colors to other elements of the main display area 406, presenting a pop-up window, or presenting other user interface elements).

[0106] In some examples, the options area 408 (or other areas of the user interface 400) may include a graph or other indicator that provides information regarding utilization or traffic on one or more paths. In such examples, these graphs may be related to or generated in response to user input entered into input fields within the options area 408.

[0107] Figure 5 is a flowchart that illustrates operations performed by an example network analysis system in accordance with one or more aspects of the present disclosure. The present disclosure describes Figure 2 in the context of the network analysis system 240 of Figure 5 . In other examples, Figure 5 the operations described may be performed by one or more other components, modules, systems, or devices. Further, in other examples, the operations described in connection with Figure 5 may be combined, performed in a different order, omitted, or may include additional operations not specifically illustrated or described.

[0108] In Figure 5 the process shown, and in accordance with one or more aspects of the present disclosure, the network analysis system 240 may collect underlying flow data (501) and overlay flow data (502). For example, in Figure 2 , each backbone device 202 and each leaf device 203 output respective signals (e.g., sFlow data) over the network 205. The communication unit 215 of the network analysis system 240 detects the signals that the collector module 252 determines include the underlying flow data 204. Similarly, the virtual router module 224 within each host device 210 outputs a signal over the network 205. The communication unit 215 of the network analysis system 240 detects the additional signals that the collector module 252 determines include the overlay flow data 206. In some examples, the collector module 252 may load balance the reception of signals across multiple collector modules 252 to ensure that a large number of signals can be processed without delay and / or without loss of data.

[0109] The network analysis system 240 can store underlying flow data and overlay flow data (503). For example, the collector module 252 can output information about the collected flow data (e.g., underlying flow data 204 and overlay flow data 206) to the data repository 259. The data repository 259 stores the flow data in an indexed format and, in some examples, in a structure that supports fast aggregation queries and / or fast random access data retrieval.

[0110] The network analysis system 240 can receive an information request regarding a data flow (the "yes" path from 504). For example, the user interface device 129 detects an input. In one such example, the user interface device 129 outputs a signal over the network 205. The communication unit 215 of the network analysis system 240 detects the signal corresponding to the information request of the user from the user interface device 129 as determined by the flow API 256. Alternatively, the network analysis system 240 can continue to collect and store the underlying flow data 204 and overlay flow data 206 until an information request regarding the data flow is received (the "no" path from 504).

[0111] The network analysis system 240 can execute a query to identify information regarding a data flow (505). For example, when the network analysis system 240 receives an information request, the flow API 256 parses the request and identifies the information that can be used to execute the query. In some cases, the information can include the source and destination virtual networks and / or the associated time frame. In other examples, the information can include other information such as the underlying source or destination IP address or source or destination port number. The flow API 256 uses the information included in the request to query the data repository 259 for information about one or more relevant data flows. The data repository 259 processes the query and outputs the identity of one or more network devices used by the traffic between the source virtual network and the destination virtual network to the flow API 256. In some examples, the identity of the network device can enable the flow API 256 to determine one or more possible data paths traversed by the traffic between the source virtual network and the destination virtual network.

[0112] To determine the identity of the network device used by the traffic between the source virtual network and the destination virtual network, the flow API 256 can query the data repository 259 for the underlying flow data of network devices having the same five-tuple data (i.e., source and destination addresses and port numbers and protocol) as specified in the query. The network devices identified in the underlying flow data that match the five-tuple data are identified as the possible network devices used by the traffic between the source virtual network and the destination virtual network. The network analysis system 240 can output information regarding the data flow (506). For example, again referring to Figure 2, the streaming API 256 can output information about the data path determined by the streaming API 256 in response to a query to the user interface module 254. The user interface module 254 generates information sufficient to render a user interface including information about the data stream. The user interface module 254 causes the communication unit 215 to output a signal via the network 205, the signal including information sufficient to render the user interface. In some examples, the user interface device 129 receives the signal, parses the information, and renders a user interface that illustrates information about the data stream.

[0113] For the processes, apparatuses, and other examples or illustrations described herein, including in any flowcharts, the specific operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be added, combined, or entirely omitted (e.g., not all described actions or events are necessary for practicing these techniques). Additionally, in a particular example, operations, actions, steps, or events may be performed concurrently rather than sequentially, such as by means of multi-threading, interrupt handling, or multiple processors. Other specific operations, actions, steps, or events may be automated even if not specifically identified as such. Moreover, specific operations, actions, steps, or events described as automated may alternatively not be automated, but rather, in some examples, such operations, actions, steps, or events may be performed in response to an input or another event.

[0114] For ease of illustration, only a limited number of devices (e.g., user interface device 129, backbone device 202, leaf device 203, host device 210, network analysis system 240, and other devices) are shown in the figures included herein and / or in other illustrations referenced herein. However, more such systems, components, devices, modules, and / or other items may be used to perform the techniques according to one or more aspects of the present disclosure, and a collective reference to such systems, components, devices, modules, and / or other items may represent any number of such systems, components, devices, modules, and / or other items.

[0115] Each of the figures included herein illustrates at least one example implementation of an aspect of the present disclosure. However, the scope of the present disclosure is not limited to such implementations. Thus, other examples or alternative implementations of the systems, methods, or techniques described herein that are beyond those illustrated in the figures may be appropriate in other instances. Such implementations may include a subset of the devices and / or components included in the figures and / or may include additional devices and / or components not shown in the figures.

[0116] The detailed description set forth above is intended as a description of various configurations and is not intended to represent the only configuration in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, the concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in the figures in order to avoid obscuring such concepts.

[0117] Accordingly, while one or more implementations of various systems, devices, and / or components may be described with reference to specific figures, such systems, devices, and / or components may be implemented in a variety of different ways. For example, one or more devices illustrated as separate devices in the figures herein (e.g., FIGS. 1 and / or Figure 2 ) may alternatively be implemented as a single device; one or more components illustrated as separate components may alternatively be implemented as a single component. Moreover, in some examples, one or more devices illustrated as a single device in the figures herein may alternatively be implemented as multiple devices; one or more components illustrated as a single component may alternatively be implemented as multiple components. Each of such multiple devices and / or components may be directly coupled via wired or wireless communication and / or remotely coupled via one or more networks. Further, one or more devices or components illustrated in the various figures herein may alternatively be implemented as part of another device or component not shown in such figures. In this and other ways, some of the functions described herein may be performed via distributed processing of two or more devices or components.

[0118] Further, specific operations, techniques, features, and / or functions may be described herein as being performed by particular components, devices, and / or modules. In other examples, such operations, techniques, features, and / or functions may be performed by different components, devices, or modules. Thus, some operations, techniques, features, and / or functions described herein as being attributable to one or more components, devices, or modules may in other examples be attributable to other components, devices, and / or modules, even if not specifically described in this manner herein.

[0119] While specific advantages have been identified with respect to the description of some examples, various other examples may include some of the recited advantages, none of the recited advantages, or all of the recited advantages. Other advantages, whether technical or otherwise, may become apparent to those of ordinary skill in the art from this disclosure. Further, while specific examples have been disclosed herein, aspects of the present disclosure may be implemented using any number of techniques, whether currently known or not, and thus, the present disclosure is not limited to the examples specifically described and / or illustrated herein.

[0120] In one or more examples, the functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on and / or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another (e.g., according to a communication protocol). In this manner, a computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0121] By way of example, and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks generally reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0122] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, as used herein, the terms "processor" or "processing circuitry" may refer to any of the foregoing structures or any other structure suitable for implementation of the described techniques. Additionally, in some examples, the functionality may be provided within dedicated hardware and / or software modules. Moreover, the techniques may be fully implemented in one or more circuits or logic elements.

[0123] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs) or a group of ICs (e.g., a chip set). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but are not necessarily implemented by distinct hardware units. Rather, as described above, the various units may be combined in a hardware unit or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

Claims

1. A method, comprising: A network analysis system collects flow data including underlying flow data and overlay flow data on a network having multiple network devices; The network analysis system receives an information request regarding a data flow, where the information request specifies a source virtual address for the data flow and also specifies a destination virtual address for the data flow; The network analysis system identifies, based on the collected flow data, the network devices that have processed at least one packet in the data flow; The network analysis system determines, based on the identified network devices, an underlying data path from a source virtual network associated with the source virtual address to a destination virtual network associated with the destination virtual address; and The network analysis system outputs information regarding the underlying data path.

2. The method according to claim 1, wherein identifying the network device comprises: Identify the correlation between the underlying flow data and the overlay flow data; And Further identify the network devices based on the correlation.

3. The method according to any one of claims 1-2, wherein identifying the network device comprises: Evaluate the overlay flow data to identify traffic registered by a virtual router; And Based on the traffic registered by the virtual router, identify the traffic between one or more virtual machines and one or more network devices.

4. The method according to claim 1, wherein the information request regarding the data stream further comprises a time frame, and wherein identifying the network device comprises: Based on the time frame, determine which of the identified network devices have processed at least one packet in the data flow during the time frame.

5. The method according to claim 4, wherein the underlying stream data comprises a plurality of underlying stream records, and wherein identifying the network device comprises: Correlate the overlay flow data with the underlying flow records during the time frame; And Add overlay flow data related to each corresponding underlying flow record to at least some of the underlying flow records.

6. The method according to claim 5, wherein adding the overlay stream data comprises: Add source virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records, and Add destination virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records.

7. The method according to any one of claims 1, 2 or 4-6, wherein identifying the network device comprises: Identify network devices having five-tuple data that matches the source virtual address or the destination virtual address.

8. The method according to any one of claims 1, 2 or 4-6, wherein the plurality of network devices comprises a plurality of host devices, each host device executing a plurality of virtual computing instances, and wherein collecting the stream data comprises: Configure each of the multiple network devices to transmit underlying flow data to the network analysis system over the network; And Configure each of the host devices to transmit the overlay flow data to the network analysis system over the network.

9. The method according to claim 8, wherein configuring each of the host devices to transmit the overlay stream data comprises: Configure the virtual router executed on each of the host devices to transmit the overlay flow data.

10. The method according to claim 8, wherein the network analysis system includes a plurality of flow collector instances, and wherein collecting the flow data includes: Load balance the collection of the flow data by distributing the flow data across the multiple flow collector instances.

11. A system comprising a storage system and a processing circuit having access to the storage system, wherein the processing circuit is configured to: Collect flow data including underlying flow data and overlay flow data on a network having a plurality of network devices; Receive an information request regarding a data stream, wherein the information request specifies a source virtual address for the data stream and also specifies a destination virtual address for the data stream; Based on the collected flow data, identify the network devices that have processed at least one packet in the data flow; Based on the identified network devices, determine an underlying data path from a source virtual network associated with the source virtual address to a destination virtual network associated with the destination virtual address; and Output information regarding the underlying data path.

12. The system according to claim 11, wherein in order to identify the network device, the processing circuit is further configured to: Identify the correlation between the underlying flow data and the overlay flow data; and Further identify the network device based on the correlation.

13. The system according to any one of claims 11 - 12, wherein in order to identify the network device, the processing circuit is further configured to: Evaluate the overlay flow data to identify traffic registered by a virtual router; and Based on the traffic registered by the virtual router, identify traffic between one or more virtual machines and one or more network devices.

14. The system according to claim 11, wherein the information request regarding the data stream further includes a time frame, and wherein in order to identify the network device, the processing circuit is further configured to: Based on the time frame, determine which of the identified network devices have processed at least one packet in the data stream during the time frame.

15. The system according to claim 14, wherein the underlying flow data includes a plurality of underlying flow records, and wherein in order to identify the network device, the processing circuit is further configured to: Correlate the overlay flow data with the underlying flow records during the time frame; and Add overlay flow data related to each corresponding underlying flow record to at least some of the underlying flow records.

16. The system according to claim 15, wherein in order to add overlay flow data, the processing circuit is further configured to: Add source virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records in the underlying flow record, and Add destination virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records in the underlying flow record.

17. The system according to any one of claims 11, 12, or 14 - 16, wherein, to identify the network device, the processing circuit is further configured to: Identify a network device having five - tuple data that matches the source virtual address or the destination virtual address.

18. The system according to any one of claims 11, 12, or 14 - 16, wherein the plurality of network devices includes a plurality of host devices, each host device executing a plurality of virtual computing instances, and wherein, to collect flow data, the processing circuit is further configured to: Configure each of the plurality of network devices to transmit underlying flow data to the system over the network; and Configure each of the host devices to transmit the overlay flow data to the system over the network.

19. The system according to claim 18, wherein, to configure each of the host devices to transmit the overlay flow data, the processing circuit is further configured to: Configure a virtual router executed on each of the host devices to transmit the overlay flow data.

20. The system according to claim 18, wherein the system further includes a plurality of flow collector instances, and wherein, to collect the flow data, the processing circuit is further configured to: Load - balance the collection of the flow data by distributing the flow data across the plurality of flow collector instances.

21. A non - transient computer - readable storage medium, including instructions that, when executed, configure a processing circuit of a computing system to: Collect flow data including underlying flow data and overlay flow data on a network having a plurality of network devices; Receive an information request regarding a data flow, wherein the information request specifies a source virtual address for the data flow and also specifies a destination virtual address for the data flow; Based on the collected flow data, identify the network devices that have processed at least one packet in the data flow; Determine an underlying data path from a source virtual network associated with the source virtual address to a destination virtual network associated with the destination virtual address based on the identified network device; and Output information about the underlying data path.

22. The non-transitory computer-readable medium according to claim 21, wherein the instructions that cause the processing circuit to identify the network device further include instructions that, when executed, further cause the processing circuit to: Identify the correlation between the underlying flow data and the overlay flow data; and Further identify the network device based on the correlation.

23. The non-transitory computer-readable medium according to any one of claims 21-22, wherein the instructions that cause the processing circuit to identify the network device further include instructions that, when executed, further cause the processing circuit to: Evaluate the overlay flow data to identify traffic registered by the virtual router; and Based on the traffic registered by the virtual router, identify the traffic between one or more virtual machines and one or more network devices.

24. The non-transitory computer-readable medium according to claim 21, wherein the information request regarding the data stream further includes a time frame, and wherein the instructions that cause the processing circuit to identify the network device further include instructions that, when executed, further cause the processing circuit to: Based on the time frame, determine which of the identified network devices have processed at least one packet in the data stream during the time frame.

25. The non-transitory computer-readable medium according to claim 24, wherein the underlying flow data includes a plurality of underlying flow records, and wherein the instructions that cause the processing circuit to identify the network device further include instructions that, when executed, further cause the processing circuit to: Correlate the overlay flow data with the underlying flow records during the time frame; and Add overlay flow data related to each corresponding underlying flow record to at least some of the underlying flow records.

26. The non-transitory computer-readable medium according to claim 25, wherein the instructions that cause the processing circuit to add overlay flow data further include instructions that, when executed, further cause the processing circuit to: Add source virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records, and Add destination virtual network data from the overlay flow data collected during the time frame to at least some of the underlying flow records.

27. The non-transitory computer-readable medium according to any one of claims 21, 22, or 24-26, wherein the instructions that cause the processing circuit to identify the network device further include instructions that, when executed, further cause the processing circuit to: Identify a network device having five-tuple data that matches the source virtual address or the destination virtual address.

28. The non-transitory computer-readable medium according to any one of claims 21, 22, or 24-26, wherein the plurality of network devices includes a plurality of host devices, each host device executing a plurality of virtual computing instances, and wherein the instructions that cause the processing circuit to collect flow data further include instructions that, when executed, further cause the processing circuit to: Configure each of the plurality of network devices to transmit underlying flow data to the computing system over the network; and Configure each of the host devices to transmit the overlay flow data to the computing system over the network.

29. The non-transitory computer-readable medium according to claim 28, wherein the instructions that cause the processing circuit to configure each of the host devices to transmit the overlay flow data further include instructions that, when executed, further cause the processing circuit to: Configure the virtual router executed on each of the host devices to transmit the overlay flow data.

30. The non-transitory computer-readable medium according to claim 28, wherein the computing system includes a plurality of flow collector instances, and wherein the instructions that cause the processing circuit to collect the flow data further include instructions that, when executed, further cause the processing circuit to: Load balance the collection of the flow data by distributing the flow data across the plurality of flow collector instances.