Data Classification Accuracy via Cumulative Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data models struggle with classifying encrypted data packets in computer networks, as they fail to provide certainty in classification and degrade over time, leading to incorrect or uncertain classification outcomes that hinder network performance management.
Innovation Solution
A method is introduced that determines statistical features of data packets, processes them through a pre-trained data model with enriched labels, and calculates a cumulative score to assess the accuracy of classification, using heuristics and conditional checks within the model to provide an explainability score for the classification output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional data models are used to classify encrypted data packets, then the classification process can be performed, but the classification accuracy and certainty degrade over time
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors classification performance and updates the data model using recent network traffic patterns. The model receives feedback about actual network conditions and adjusts its parameters accordingly, preventing performance degradation over time and maintaining high classification accuracy.
Solution Approach 2:
The patent transforms the static data model into a dynamic system that adapts to changing network conditions. The model continuously learns from new data patterns and adjusts its classification criteria in real-time, enabling it to maintain accuracy despite evolving network traffic characteristics and encrypted data formats.
2Loss of information
If conventional data models classify data packets, then classification output is obtained, but no information is provided about classification certainty
Solution Approach 1:
The patent segments the classification output into distinct components: the classified data packet information and a separate confidence score. This segmentation allows the system to provide both the classification result and the certainty level independently, adding valuable information about classification reliability without unnecessarily complicating the overall system structure.
3Adaptability or versatility
If data models are trained on historical data, then classification can be performed, but performance degrades when actual data differs from training data
Solution Approach 1:
The patent implements preliminary actions by pre-processing and normalizing data patterns before they enter the classification model. The system prepares feature representations in advance that are robust to variations in actual network traffic, enabling the model to maintain accuracy even when real-world data differs from training data distributions.
Solution Approach 2:
The system continuously monitors classification performance and uses feedback loops to detect when data distributions have shifted. When deviations are detected, the model automatically adjusts its parameters or triggers retraining processes, ensuring it remains adapted to current network conditions while maintaining high classification accuracy.
Data Source
AI summary
A system and a method of classifying data and providing an accuracy of classification are described. The method includes determining values of statistical features associated with data packets present in a data stream. The values of statistical features are provided to a data model for producing a classification output including the data packets classified into one or more categories. While producing the classification output, the data model extracts heuristics for each of the values of statistical features, compares the heuristics with one or more conditional checks defined at each node within the data model, and determines a cumulative score based on results of the comparing. The cumulative score is determined by aggregating a score assigned to successful clearance of each conditional check. The cumulative score indicates an accuracy of the classification output.


