ML Classifier for Network Traffic Data Retention Priority
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer networks face challenges in retaining all traffic data indefinitely due to system resource constraints, making it unfeasible to store historical data for retrospective detection and network forensics effectively.
Innovation Solution
A machine learning classifier is used to determine a data retention priority for traffic data, allowing devices in the network to store traffic data for a period based on its classification as benign or malicious, thereby optimizing storage resources by retaining potentially malicious data for a longer time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all traffic data is retained indefinitely for retrospective detection and network forensics, then detection capability is improved, but system resource constraints are exceeded
Solution Approach 1:
The patent applies different retention policies to different traffic data based on their classification results. Benign traffic data is retained for a shorter period or discarded, while malicious traffic data is retained for a longer period. This local differentiation of retention quality based on data characteristics resolves the contradiction between comprehensive detection capability and limited storage resources.
Solution Approach 2:
The patent changes the retention time parameter based on the classification outcome of traffic data. By using machine learning classification to determine retention duration, the system dynamically adjusts storage parameters to optimize between detection reliability and resource consumption, retaining only potentially malicious data indefinitely while discarding benign data after a shorter period.
2Measurement precision
If traffic data is retained for a longer period to enhance forensics capability, then detection accuracy is improved, but storage costs increase
Solution Approach 1:
The patent implements differential retention strategies where malicious traffic data receives extended retention for thorough forensic analysis, while benign traffic data is retained for shorter periods or discarded. This local quality differentiation ensures high forensic accuracy for critical data while minimizing storage resource consumption for non-critical data.
Solution Approach 2:
The system dynamically changes the retention time parameter based on machine learning classification results. Traffic data identified as malicious is retained indefinitely or for extended periods to enable accurate forensic analysis, while benign data is retained for shorter durations, optimizing the balance between detection precision and storage efficiency.
3Speed
If real-time classification is performed for all traffic flows, then network response speed is improved, but processing complexity increases
Solution Approach 1:
The patent employs lightweight machine learning classification models that can be rapidly applied to traffic data without requiring complex processing infrastructure. These simplified classifiers enable real-time or near-real-time classification decisions, maintaining network response speed while reducing processing complexity through the use of computationally efficient algorithms.
Data Source
AI summary
In one embodiment, a device in a network receives traffic data regarding one or more traffic flows in the network. The device applies a machine learning classifier to the traffic data. The device determines a priority for the traffic data based in part on an output of the machine learning classifier. The output of the machine learning classifier comprises a probability of the traffic data belonging to a particular class. The device stores the traffic data for a period of time that is a function of the determined priority for the traffic data.


