Automated Malicious Traffic Detection via Metadata Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Websites face challenges in identifying and distinguishing malicious traffic from genuine traffic, which can skew analytics, erode user and advertiser trust, and damage revenue, due to the inefficiencies of manual analysis and reliance on unreliable manually labeled inputs.
Innovation Solution
The system employs a method of receiving and processing website traffic metadata, generating pairs of variables, determining visitor actions, clustering data into groups, and automatically identifying malicious visitors based on these clusters, without requiring reliable manually labeled inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual analysis methods are used to identify malicious traffic, then the system complexity is reduced, but the detection reliability and productivity are insufficient
Solution Approach 1:
The patent replaces manual analysis mechanisms with automated computational algorithms. The system automatically processes traffic metadata, generates variable combinations, performs clustering analysis, and identifies malicious visitors without human intervention, thereby improving detection reliability while managing system complexity through structured automation
Solution Approach 2:
The system performs self-service by automatically generating training data from traffic metadata and performing unsupervised learning through clustering algorithms. The automated process eliminates the need for manual labeling and continuously improves its own detection capabilities without requiring external human input
2Productivity
If automated clustering algorithms are applied to traffic data, then the detection productivity increases, but the computational resource consumption increases
Solution Approach 1:
The patent segments the traffic data into manageable components by extracting specific metadata variables (IP address, user agent, referrer, etc.) and creating pairs of variables for analysis. This segmentation allows the clustering algorithm to process data in organized units, improving productivity while managing computational resources through structured data organization
Solution Approach 2:
The system applies partial action by focusing clustering analysis on the most relevant variable pairs rather than processing all possible combinations exhaustively. The method generates pairs from a limited set of key metadata variables, achieving sufficient detection accuracy without requiring exhaustive computational analysis of all data dimensions
3Measurement precision
If manually labeled inputs are relied upon for training, then the initial setup is simpler, but the measurement precision and reliability of analytics are compromised
Solution Approach 1:
The system performs self-service by automatically generating training data from traffic metadata without requiring manual labeling. The automated process extracts variables, creates pairs, performs clustering, and identifies malicious patterns independently, eliminating dependency on manually labeled inputs and improving measurement precision while maintaining reasonable setup simplicity through automated data preparation
4Reliability
If comprehensive traffic metadata is collected and analyzed, then the detection accuracy improves, but the data processing complexity increases
Solution Approach 1:
The patent segments comprehensive traffic metadata into specific, relevant variables (IP address, user agent, referrer URL, etc.) and organizes them into pairs for analysis. This segmentation approach maintains detection accuracy by focusing on the most informative variables while reducing processing complexity through structured data organization and selective analysis of variable combinations
Data Source
AI summary
Systems and methods are disclosed for identifying malicious traffic associated with a website. One method includes receiving website traffic metadata comprising a plurality of variables, the website traffic metadata being associated with a plurality of website visitors to the website; determining a total number of occurrences associated with at least two of the plurality of variables of the website traffic metadata; generating a plurality of pairs comprising combinations of the plurality of variables of the website traffic metadata; determining a total number of occurrences associated with each pair of the plurality of pairs of combinations of the plurality of variables of the website traffic metadata; determining a plurality of visitor actions associated with the plurality of variables of the website traffic metadata; clustering each of the plurality of pairs and the plurality of visitor actions associated with the plurality of variables of the website traffic metadata into groups; and determining, based on the clustering of the plurality of pairs and the plurality of visitor actions, whether each of the plurality of website visitors are malicious visitors.


