Traffic Statistics Generation via Flow Agents and Analytic Controller
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenters face challenges in efficiently allocating resources due to difficulties in determining and managing datacenter traffic, leading to inefficient allocation of expensive resources and potential traffic impediments, exacerbated by the complexity of large data sets and the need for accurate traffic statistics.
Innovation Solution
Implementing a connection-based packet switching approach that establishes a common path for traffic flows, using flow agents to summarize data at the traffic-flow level, and an analytic controller to aggregate and preprocess data for distributed parallel processing, reducing the data set size and improving resource management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If packet-level data collection is implemented to obtain detailed traffic statistics, then measurement precision is improved, but device complexity and data processing burden increase significantly
Solution Approach 1:
The patent segments the data collection process into two levels: flow-level summarization (performed by flow agents at network nodes) and packet-level detail (maintained for analysis). This segmentation allows detailed traffic statistics to be collected without processing every packet at all nodes, reducing overall system complexity while maintaining measurement precision.
Solution Approach 2:
Flow agents perform preliminary summarization of packet data into flow-level records before data reaches the central analysis system. This preliminary action filters and aggregates data early in the pipeline, reducing the volume of data that requires complex processing while preserving the statistical accuracy needed for traffic analysis.
2Device complexity
If flow agents summarize data at traffic-flow level to reduce data set size, then device complexity is reduced, but measurement precision may be compromised
Solution Approach 1:
The patent merges flow-level summaries with packet-level information by maintaining a mapping between flows and their constituent packets. This allows the system to process summarized flow data for complexity reduction while retaining access to packet-level details when needed for precise statistical analysis, thus balancing both objectives.
Solution Approach 2:
Flow agents act as intermediaries between packet-level data collection and central analysis. They summarize packets into flow records to reduce immediate processing complexity, yet preserve measurement precision by maintaining complete flow information and enabling detailed queries when required.
3Productivity
If connection-based packet switching is implemented to establish common paths for traffic flows, then traffic management is improved, but device complexity increases
Solution Approach 1:
The patent implements self-service through flow agents that automatically generate and manage flow records and paths without manual configuration. Flow agents autonomously summarize traffic data, identify flows, and create path information, reducing the need for complex centralized control while improving resource allocation efficiency through automated connection-based switching.
Data Source
AI summary
Systems and methods are disclosed for generating traffic statistics for a datacenter. Distributed, parallel processing may be used to generate traffic statistics from data sets about traffic in a datacenter. To reduce data sets from which such statistics are derived to manageable sizes and relevant processing times for distributed, parallel processing, traffic agents may be provided at end hosts in the datacenter. The traffic agents may summarize data traffic over large numbers of packets in terms of the various sockets over which they are transmitted. Reports on the various sockets may be sent by the various flow agents that monitor them to an analytic controller. The analytic controller may aggregate, provide flow-path information for, further reduce, and/or provision the resultant data for distributed parallel processing.


