IP Address Classifier Training Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for classifying Internet Protocol (IP) related data and metadata are not sufficiently accurate, leading to suboptimal downstream products in terms of quality and storage requirements.
Innovation Solution
A process is developed to generate a labeled training data set, which is used to train an IP address classifier. This classifier assigns classifications to input IP addresses as residential or non-residential based on a set of features and inclusion criteria, improving the accuracy and efficiency of IP address classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing classification techniques are used, then the classification process is simple, but the classification accuracy is insufficient
Solution Approach 1:
The patent applies preliminary action by generating a labeled training data set before deploying the classification system. The training data set includes IP addresses with pre-determined labels (residential or non-residential) based on multiple data sources and analysis techniques. This pre-prepared training data enables the machine learning model to learn accurate classification patterns beforehand, improving classification accuracy without increasing the complexity of the operational classification process
Solution Approach 2:
The patent introduces an intermediary element - a machine learning classifier trained on comprehensive training data that incorporates multiple data sources (ISP data, device information, network behavior). This intermediary trained model acts as a mediator between raw IP address data and classification results, enabling accurate classification while maintaining system simplicity during operation
2Reliability
If classification accuracy is improved, then downstream product quality increases, but storage requirements increase
Solution Approach 1:
The patent applies parameter changes by optimizing the training data set composition and classification model parameters to achieve high accuracy with efficient data utilization. The system adjusts parameters such as feature selection, data weighting, and classification thresholds to maximize accuracy while minimizing the storage footprint of training data and model structures
Solution Approach 2:
The patent extracts only the essential and most discriminative features from the training data set for model training and classification. By selecting key features that most strongly correlate with residential vs. non-residential classification, the system achieves high accuracy while reducing the amount of data that needs to be stored and processed, thereby reducing storage requirements
Data Source
AI summary
A set of Internet Protocol (IP) addresses is received wherein each IP address is associated with a corresponding set of features. For an IP address in the set, the IP address is evaluated based at least in part on a set of inclusion criteria. For the IP address in the set, a likelihood that the IP address is residential or non-residential is generated based at least in part on the corresponding set of features and the evaluation of the IP address based at least in part on the set of inclusion criteria. For the IP address in the set, a training sample is generated that includes the IP address, at least some of the corresponding set of features, and a label. A labeled training data set is output that includes the training sample, where an IP address classifier is trained using the labeled training data set.


