Collusion Detection via Graph Clustering for Click Fraud
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current click fraud detection methods are inadequate in identifying collusion fraud, which involves multiple entities and is computationally expensive due to the spread of fraudulent clicks across numerous sites and IP addresses, making it difficult to detect and requiring practical and scalable solutions.
Innovation Solution
The approach models collusion detection as a clustering problem in networks or vector spaces, constructing representations that capture relevant information for click fraud, using network and vector space analysis to identify potential collusion among entities, and applying efficient clustering methods to detect botnets and click farms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional fraud detection methods are used to detect collusion fraud spread across multiple entities, then detection accuracy may be maintained for single-entity fraud, but detection capability deteriorates for multi-entity collusion fraud
Solution Approach 1:
The patent segments the complex collusion detection problem into multiple analytical components: network analysis to identify entity relationships, vector space modeling to represent click patterns, and clustering algorithms to group suspicious entities. This segmentation allows the system to maintain high detection accuracy by addressing different aspects of collusion detection separately and integrating results.
Solution Approach 2:
The patent introduces intermediary representations including network graphs that model relationships between entities, vector space embeddings that capture click behavior patterns, and clustering intermediaries that bridge raw data to detection results. These intermediaries transform complex multi-entity collusion patterns into analyzable forms, improving detection capability.
2Reliability
If comprehensive analysis of all entities and their relationships is performed to detect collusion, then detection completeness improves, but computational cost increases
Solution Approach 1:
The patent performs preliminary actions by pre-computing network representations, vector space embeddings, and similarity metrics before actual collusion detection. These pre-computed structures enable faster querying and analysis during detection operations, maintaining detection completeness while reducing real-time computational costs.
Solution Approach 2:
The system employs self-service mechanisms where the network analysis and vector space models automatically identify and focus computational resources on suspicious entity groups. The clustering algorithms self-organize entities based on their click patterns and relationships, reducing the need for exhaustive analysis of all entity combinations.
3Measurement precision
If fraud detection systems analyze numerous entities and their interrelationships to identify collusion, then detection thoroughness improves, but time consumption increases
Solution Approach 1:
The patent replaces traditional mechanical analysis methods with efficient computational approaches: using network analysis algorithms instead of manual relationship tracing, applying vector space modeling instead of sequential pattern matching, and employing clustering techniques instead of exhaustive entity pair comparisons. This substitution maintains detection thoroughness while dramatically reducing time consumption.
Solution Approach 2:
The patent transforms the detection problem by changing parameters from analyzing individual click events to analyzing aggregated patterns in vector spaces and network structures. This parameter transformation allows the system to maintain thorough detection by capturing essential collusion characteristics while reducing the dimensionality and time required for analysis.
Data Source
AI summary
Embodiments disclosed herein provide a practical solution for click fraud detection. One embodiment of a method may comprise constructing representations of entities via a graph network framework. The representations, graphs or vector spaces, may capture information pertaining to clicks by botnets/click farms. To detect click fraud, each representation may be analyzed in the context of clustering, resulting in large data sets with respect to time, frequency, or gap between clicks. Highly accurate and highly scalable heuristics may be developed/applied to identify IP addresses that indicate potential collusion. One embodiment of a system having a computer program product implementing such a click fraud detection method may operate to receive a client file containing clicks gathered at the client side, construct representations of entities utilizing the graph framework described herein, perform clustering on the representations thus constructed, identify IP addresses of interest, and return a list containing same to the client.


