Synthetic Traffic Data for ML Classifier Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network security systems face challenges in distinguishing between legitimate and malicious traffic flows, especially in distributed Denial of Service (DoS) attacks, due to biases in training datasets from different network environments, which leads to high false positive rates in traffic classification.
Innovation Solution
The generation of synthetic traffic data samples that match the characteristics of a targeted deployment environment, allowing for the training of machine learning-based classifiers to reduce false positives and improve classification accuracy across diverse network environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning classifiers are trained using observed traffic data from specific network environments, then the classifiers can achieve reasonable classification performance in those environments, but they exhibit high false positive rates when deployed in different network environments due to environmental biases
Solution Approach 1:
The patent creates synthetic traffic data that copies the essential characteristics and statistical properties of real network traffic from target deployment environments. By generating artificial training datasets that replicate environmental biases and traffic patterns, the classifiers can be trained offline to adapt to specific network environments without requiring actual traffic from those environments, thereby improving both reliability and environmental adaptability
Solution Approach 2:
The patent employs parameter changes by systematically varying traffic flow characteristics (such as packet size distributions, inter-arrival times, protocol patterns) in synthetic data generation to match different network environments. This allows the same base classifier to adapt to multiple environments by retraining with synthetic data that has adjusted parameters reflecting the target environment's unique characteristics
2Adaptability or versatility
If classifiers are trained with diverse training data from multiple environments, then environmental adaptability improves, but the complexity of data collection and processing increases
Solution Approach 1:
Instead of collecting actual traffic data from multiple diverse environments, the patent uses synthetic data generation to copy and simulate the necessary environmental characteristics. This approach achieves the same adaptability goal without the operational complexity of multi-environment data collection, storage, and preprocessing infrastructure
Solution Approach 2:
The patent introduces synthetic traffic data as an intermediary between the training process and the target deployment environment. This intermediary allows the classifier to be trained offline with controlled, environment-specific characteristics without requiring direct access to or processing of actual network traffic from multiple environments, thereby reducing data collection and processing complexity
Data Source
AI summary
In one embodiment, a device in a network receives traffic data regarding a plurality of observed traffic flows. The device maps one or more characteristics of the observed traffic flows from the traffic data to traffic characteristics associated with a targeted deployment environment. The device generates synthetic traffic data based on the mapped traffic characteristics associated with the targeted deployment environment. The device trains a machine learning-based traffic classifier using the synthetic traffic data.


