Application Traffic Identification via Pseudo-Label Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for application-level traffic identification in large-scale networks are inefficient due to the inability to handle new applications and the requirement for extensive training data, particularly with rule-based technologies and supervised machine learning, which struggle with labeling and accuracy.
Innovation Solution
An identifier generation device and method that acquires flow data, calculates feature vectors, converts them into similar-type application vectors, clusters, adds pseudo-labels, generates a learning dataset, and updates the identifier settings using meta-learning to reduce the need for large training datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If rule-based technologies are used for application identification, then identification speed is maintained, but the system cannot identify newly emerging applications
Solution Approach 1:
The system performs preliminary clustering and pseudo-labeling on flow data to pre-organize application patterns before actual identification occurs. This preliminary action creates a structured knowledge base that enables fast identification of both known and new applications without requiring retraining
Solution Approach 2:
The system creates pseudo-labels by copying and adapting patterns from similar applications. When identifying new applications, the system copies feature patterns from clustered flow data and generates pseudo-labels that enable rapid identification without requiring extensive training data for each new application
2Measurement precision
If supervised machine learning is used for application identification, then identification accuracy can be improved, but a large amount of training data with application-level labels is required
Solution Approach 1:
The system performs self-service by automatically generating pseudo-labels from flow data through unsupervised clustering. The system serves its own training data generation needs without requiring external manual labeling, thereby achieving high identification accuracy with minimal training data
Solution Approach 2:
The system changes the parameter of data labeling from manually supervised labels to automatically generated pseudo-labels. This parameter change in the labeling approach enables the system to achieve supervised learning accuracy using unsupervised clustering results, dramatically reducing the quantity of training data required
3Ease of manufacture
If flow data with simple information is used for training, then data collection is easy, but adding application-level labels is difficult and accuracy remains low
Solution Approach 1:
The system introduces clustering and pseudo-labeling as intermediary processes between simple flow data collection and accurate application identification. These intermediary steps transform unlabeled flow data into structured training data with pseudo-labels, bridging the gap between easy data collection and high identification accuracy
Data Source
AI summary
An identifier generation device includes identifier generation circuitry configured to acquire flow data of an application, calculate first feature vectors from the flow data, convert the first feature vectors into second feature vectors to which feature vectors of an identical type of application are similar, cluster the second feature vectors and add a pseudo-label to the clustered second feature vectors, generate a learning data set from the second feature vectors to which the pseudo-label is added, supply the learning data set to an identifier, and update a setting of the identifier to which the learning data set is supplied.


