Web Resource Clustering for Phishing Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying clusters of affiliated web resources are cumbersome and inefficient, often relying on specific parameters like DNS addresses, which fail to effectively respond to phishing attacks generated by affiliated web resources.
Innovation Solution
A method and system that group web resources based on attributes such as structural elements, regional settings, and ownership, using a pattern affiliation approach to detect and prevent phishing attacks by generating clusters and calculating an affiliation ratio to associate new web resources with existing clusters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If prior art methods use specific parameters like DNS addresses to identify affiliated web resources, then the identification process becomes simpler, but the effectiveness in detecting phishing attacks decreases
Solution Approach 1:
The patent transitions from using single specific parameters (DNS addresses) to using multiple diverse attributes including structural elements, regional settings, ownership information, and content characteristics. This parameter diversification maintains operational simplicity while significantly improving phishing detection effectiveness by capturing the multifaceted nature of affiliated web resources.
Solution Approach 2:
The patent creates a composite identification approach by combining multiple different attributes (structural, regional, ownership, content) into an integrated affiliation determination system. This composite methodology leverages the strengths of each attribute type to achieve both simplicity and high effectiveness in phishing detection.
2Measurement precision
If the system scans and analyzes all web resources to generate clusters, then the detection accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary clustering of web resources based on multiple attributes during a training phase, creating pre-established clusters that can be quickly referenced during runtime. This preliminary action enables fast affiliation determination for new web resources without requiring exhaustive re-analysis, thus maintaining high detection accuracy while reducing processing time.
Solution Approach 2:
The patent segments the web resource analysis process into distinct attribute categories (structural elements, regional settings, ownership, content) that can be independently evaluated and combined. This segmentation allows for efficient processing by focusing on specific attribute groups rather than analyzing all possible characteristics of each web resource comprehensively.
3Adaptability or versatility
If the system updates clusters dynamically with new web resources, then the adaptability to new phishing patterns improves, but the system complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where newly analyzed web resources are used to update and refine existing clusters. The system continuously learns from new data by incorporating affiliation ratios and attribute patterns into cluster definitions, enabling adaptation to emerging phishing patterns while maintaining manageable complexity through systematic update procedures.
Solution Approach 2:
The patent creates dynamic clusters that evolve over time as new web resources are analyzed and incorporated. The cluster structures and affiliation criteria are not static but adapt based on accumulated data, allowing the system to respond to changing phishing tactics while using structured algorithms to control the complexity of these dynamic adjustments.
Data Source
AI summary
A method and a system for determining affiliation of a web resource to a plurality of clusters are provided. The method includes: at a training stage: detecting a plurality of web resources; retrieving information associated with the plurality of web resources; generating a respective pattern based on the information; grouping the plurality of web resources into the plurality of clusters, based on the respective pattern; at a run-time stage: receiving an indication of a given web resource; retrieving the information about the given web resource; generating a new pattern of the given web resource; analyzing pattern affiliation of the new pattern with a specific one from the plurality of clusters of web resources; calculating an affiliation ratio therewith; in response to the affiliation ratio exceeding a predetermined threshold value, associating the given web resource with the specific one of the plurality of clusters.


