Server Grouping via Statistical IP Clustering for Encrypted Traffic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cloud administration tools and intrusion detection systems face challenges in analyzing network traffic due to the complexity of cloud environments, especially when dealing with encrypted flows to raw IP addresses, where limited information is available for statistical modeling, making it difficult to identify groups of cooperating IP addresses serving the same application or service.
Innovation Solution
A method and system that receive client-server connection data, perform statistical tests to determine related server IP addresses based on common clients, generate a graph with vertices and edges representing IP addresses, and cluster them to identify subsets of servers serving the same application, using techniques like Binomial and Bayesian statistical tests to determine significant connections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cloud administration tools analyze network traffic using exact knowledge of technical details of each software protocol and service, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent introduces an intermediary statistical model that operates between raw network traffic data and high-level cloud service understanding. This model uses probabilistic relationships and correlation analysis to bridge the gap without requiring exhaustive knowledge of every protocol and service detail, thereby maintaining measurement precision while reducing system complexity
Solution Approach 2:
The system transforms the analysis approach by changing parameters from exact protocol-specific knowledge to statistical correlation metrics. By using probabilistic models and correlation coefficients instead of deterministic protocol parsing, the system achieves effective network traffic analysis with reduced complexity
2Measurement precision
If cloud administration tools obtain complete technical knowledge about the cloud environment, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent applies preliminary action by pre-establishing statistical models and correlation frameworks that can quickly process new network traffic data. These pre-configured models enable rapid analysis without requiring time-consuming collection of complete technical knowledge about the cloud environment, as the statistical relationships are established in advance
Solution Approach 2:
The system uses partial action by focusing on collecting and analyzing only the most relevant statistical correlations needed for cloud service identification, rather than attempting to gather complete technical knowledge of all protocols and services. This selective approach achieves sufficient measurement precision with reduced time investment
3Measurement precision
If statistical tests are performed on all IP address pairs to determine related server addresses, then measurement precision is improved, but productivity decreases
Solution Approach 1:
The patent segments the analysis process into multiple stages: first performing coarse filtering using basic connection patterns, then applying statistical tests only to candidate IP address pairs that pass initial screening. This segmentation maintains measurement precision for identifying related server addresses while significantly improving processing speed by avoiding exhaustive analysis of all IP pairs
Solution Approach 2:
The system applies partial action by performing comprehensive statistical tests only on a subset of IP address pairs that are most likely to be related based on preliminary analysis criteria. This selective approach achieves sufficient measurement precision for security applications without the computational burden of analyzing all possible IP address combinations
Data Source
AI summary
In one embodiment, a method includes receiving client-server connection data for clients and servers, the data including IP addresses corresponding to the servers, for each one of a plurality of IP address pairs performing a statistical test to determine whether the IP addresses in the one IP address pair are related by common clients based on the number of the clients connecting to each of the IP addresses in the one IP address pair, generating a graph including a plurality of vertices and edges, each of the vertices corresponding to a different IP address, each edge corresponding to a different IP address pair determined to be related by common clients in the statistical test, and clustering the vertices yielding clusters, a subset of the IP addresses in one of the clusters providing an indication of the IP addresses of the servers serving a same application.


