Distributed Edge Analytics for Telco IP Traffic
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current big data technologies, developed for centralized data centers handling user-generated content, are suboptimal for network traffic analytics in the telco domain, where machine-generated data at high speed requires distributed processing to prevent network degradation and ensure timely anomaly detection.
Innovation Solution
Implementing a distributed architecture using big data tools like Hadoop and Spark across multiple network devices to aggregate and process Internet Protocol (IP) traffic locally, forming a hierarchy where raw data is transformed into intermediate results and eventually summarized at higher levels, reducing the volume of data transmitted and enabling efficient network traffic characterization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all traffic data is uploaded to centralized data centers for processing, then comprehensive network traffic analytics can be achieved, but network resources are significantly consumed and network operation is degraded
Solution Approach 1:
The patent divides the centralized data processing architecture into distributed edge computing nodes deployed at network edges. Each edge node independently processes traffic data from its local network segment, performing analytics functions such as traffic classification, anomaly detection, and flow aggregation. This segmentation eliminates the need to upload all raw traffic data to centralized data centers, significantly reducing network resource consumption while maintaining comprehensive analytics coverage across the entire network.
2Loss of information
If all traffic data is aggregated at centralized data centers, then complete network traffic characterization is achieved, but processing lags occur that render analytics irrelevant for real-time attack detection
Solution Approach 1:
The patent implements preliminary data processing and filtering at distributed edge nodes before data is transmitted to centralized systems. Edge nodes perform initial traffic characterization, aggregate flows, filter normal traffic patterns, and pre-process data into compressed representations. This preliminary action ensures that when data reaches centralized data centers, the processing can be completed rapidly without significant lags, while still maintaining complete network traffic characterization through coordinated analysis across multiple edge nodes.
3Productivity
If distributed architecture is used for traffic processing, then network resource utilization is improved and processing speed increases, but system complexity increases
Solution Approach 1:
The patent employs standardized edge computing nodes with universal functionality that can be deployed across different network segments. Each edge node implements the same suite of analytics capabilities (traffic classification, anomaly detection, flow aggregation), allowing them to independently handle various types of network traffic. This universality simplifies the overall system architecture by using identical modular components throughout the distributed network, reducing the complexity that would otherwise arise from heterogeneous specialized processors while maintaining high processing speeds.
Data Source
AI summary
Exemplary methods for performing distributed data aggregation include receiving Internet Protocol (IP) traffic from only a first portion of the network. The methods further include utilizing a big data tool to generate a summary of the IP traffic from the first portion of the network, wherein a summary of IP traffic from a second portion of the network is generated by a second network device utilizing its local big data tool. The methods include sending the summary of the IP traffic of the first portion of the network to the third network device, causing the third network device to utilize its local big data tool to generate a summary of the IP traffic of the first and second portion of the network based on the summaries received from the first and second network devices, thereby allowing the IP traffic in the network to be characterized in a distributed manner.


