Network Traffic Analysis via K-Means Clustering and Wavelet Transforms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network traffic analysis methods are inadequate in efficiently identifying and visualizing anomalous activity and predicting future outcomes, especially in large datasets, and fail to provide a universal taxonomy for network behavior across diverse topologies.
Innovation Solution
An unsupervised machine learning method that segments internet traffic into clusters using the K-means algorithm, determines relative activity through neuronal models, and generates continuous activation plots for real-time visualization, creating a universal taxonomy of network behavior across different organizations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional network security applications use traditional traffic analysis methods, then they can identify known attack patterns, but they fail to detect anomalous activity and cannot provide universal characterization across diverse network topologies
Solution Approach 1:
The patent transforms network traffic from traditional binary classification (malicious/benign) to continuous spectral decomposition using Fourier transforms and wavelet analysis. This parameter transformation enables the system to capture temporal patterns, frequency characteristics, and anomalous behaviors that conventional methods miss, thereby improving detection accuracy while maintaining versatility across network topologies.
Solution Approach 2:
The invention replaces conventional mechanical traffic filtering and signature matching with unsupervised machine learning algorithms that perform automatic pattern recognition and anomaly detection. This substitution allows the system to adapt to diverse network topologies without requiring topology-specific rule sets, achieving both high detection accuracy and broad adaptability.
2Quantity of substance
If the system processes large datasets of network packets, then it can analyze comprehensive traffic patterns, but the processing time and computational resources increase significantly
Solution Approach 1:
The patent segments network traffic data into discrete time windows and frequency bands using wavelet transforms. This segmentation allows the system to process large datasets by analyzing smaller, manageable segments in parallel, reducing overall processing time while maintaining comprehensive coverage of traffic patterns across the entire dataset.
Solution Approach 2:
The invention transforms high-dimensional packet data into lower-dimensional spectral representations using Fourier transforms and principal component analysis. This dimensionality reduction preserves essential traffic characteristics while significantly reducing computational complexity, enabling fast processing of large datasets without losing critical information.
3Extent of automation
If the system performs unsupervised segmentation and clustering of packets, then it can autonomously learn network behavior and identify anomalies, but the computational complexity and algorithmic difficulty increase
Solution Approach 1:
The patent employs periodic wavelet transforms and cyclic clustering algorithms that iteratively refine network behavior models. These periodic operations enable the system to autonomously learn temporal patterns and anomaly signatures through repeated cycles of data processing and model updating, achieving high automation while managing computational complexity through efficient algorithm design.
Data Source
AI summary
The present disclosure, in one embodiment, relates to a method for network traffic analysis. The method includes a step of reception of a data set associated with an internet traffic at a network traffic analyzing system with a processor. The method includes another step of segmentation of the internet traffic to create a plurality of clusters based on a pre-selected percent variation. The method includes yet another step of determination of a relative activity of a set of clusters at a point in time. The method includes yet another step of determination of the relative activity of the set of clusters during successive time intervals. The data set associated with the internet traffic comprising data in the form of packets, wherein each packet is vectorized into a sequence of n-values. Each cluster of the plurality of clusters containing similar packets assigned with a same cluster ID.


