Machine Learning Account Classification Using Email Memorability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying malicious online accounts are costly, labor-intensive, and time-consuming, as they often rely on manual review of account information, and existing systems lack efficiency in distinguishing between human-generated and machine-generated email addresses.
Innovation Solution
A computer-implemented method that uses machine learning to extract features from email addresses and associated account information, applying a classification model to generate a score indicating the likelihood of an account being malicious, based on memorability and other factors such as domain and correlation with credit card information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of account information is used to identify malicious accounts, then identification accuracy can be maintained, but the process becomes costly, labor-intensive, and time-consuming
Solution Approach 1:
The patent replaces the mechanical manual review process with an automated machine learning classification system. The system extracts features from account information including email addresses, applies a classification model to generate maliciousness scores, and automatically determines whether accounts are malicious without human intervention, thereby maintaining accuracy while dramatically improving processing efficiency
Solution Approach 2:
The patent introduces a machine learning classification model as an intermediary between raw account data and final malicious account identification. This intermediary system processes account information through feature extraction and classification algorithms, serving as an automated mediator that eliminates the need for direct manual review while preserving identification accuracy
2Productivity
If existing systems use simple email address checking, then the process is fast, but they lack efficiency in distinguishing between human-generated and machine-generated email addresses
Solution Approach 1:
The patent segments the email address analysis into multiple distinct feature dimensions including memorability features (length, character composition, pattern complexity), domain features, and correlation features with other account information. This segmentation allows the system to evaluate each aspect separately and combine them for comprehensive discrimination between human-generated and machine-generated addresses
Solution Approach 2:
The patent transitions from single-dimensional email address checking to multi-dimensional analysis by incorporating memorability scoring, domain reputation, and cross-correlation with account information. This dimensional expansion enables the system to distinguish sophisticated machine-generated addresses that might pass simple checks by evaluating them across multiple concurrent dimensions
3Measurement precision
If a classification model is trained incrementally with new data, then the model accuracy improves over time, but the system complexity increases
Solution Approach 1:
The patent implements a dynamic classification model that evolves over time through incremental training. The system continuously incorporates new labeled accounts into the training dataset, allowing the classification model to adapt and improve its accuracy while maintaining operational functionality. This dynamic approach enables the system to stay current with emerging fraudulent patterns without requiring complete retraining
Solution Approach 2:
The system performs self-improvement through automatic incremental training where the classification model uses new labeled data from determined malicious and benign accounts to retrain and enhance its own performance. This self-service mechanism allows the system to automatically improve classification accuracy without external intervention, managing the complexity through automated feedback loops
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A trust level of an account is determined at least partly based on a degree of the memorability of an email address associated with the account. Additional features such as those based on the domain of the email address and those from the additional information such as name, phone number, and address associated with the account may also be used to determine the trust level of the account. A machine learning process may be used to learn a classification model based on one or more features that distinguish a malicious account from a benign account from training data. The classification model is used to determine a trust level of the account, and/or if the account is malicious or benign, and may be continuously improved by incrementally adapting or improving the model with new accounts.