Phishing Detection via Visual and Structural Webpage Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting phishing webpages face challenges in efficiently utilizing overall structural and visual information, particularly in identifying newly surfaced zero-day phishing hosts and effectively differentiating between legitimate and fraudulent websites, especially in a time-efficient manner.
Innovation Solution
A hybrid system and method that captures and compares overall visual and structural information of webpages with pre-stored legitimate webpage data, using a processor-controlled method to calculate similarity measures and alert users or block phishing sites, while also maintaining a database of legitimate and phishing webpages for continuous updating and prioritization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional phishing detection methods are used, then existing phishing sites can be identified, but newly surfaced zero-day phishing hosts cannot be detected in time
Solution Approach 1:
The system performs preliminary actions by capturing and storing visual and structural information of legitimate webpages before phishing attacks occur. This allows the system to have pre-existing reference data for comparison, enabling rapid detection of new phishing sites without waiting for traditional detection methods to identify them.
Solution Approach 2:
The patent replaces traditional mechanical inspection methods with automated visual and structural information capture and comparison systems. By using image processing and data comparison techniques, the system achieves faster detection of phishing sites compared to manual or traditional automated methods.
2Reliability
If comprehensive visual and structural information comparison is performed, then phishing detection accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The system segments the webpage analysis into distinct visual information capture and structural information capture components. By dividing the complex task into separate modules that can process different aspects of webpage data independently, the system maintains comprehensive analysis while improving overall processing efficiency through parallel execution.
Solution Approach 2:
The patent applies partial action by capturing and comparing only the most critical visual and structural features of webpages rather than every possible detail. This selective approach maintains sufficient accuracy for phishing detection while significantly reducing processing time and computational resource requirements.
3Reliability
If a database of legitimate webpages is maintained for comparison, then detection accuracy for phishing sites improves, but system complexity and data management requirements increase
Solution Approach 1:
The system performs self-service by automatically capturing, storing, and updating its own database of legitimate webpage information without requiring manual intervention. The system autonomously maintains its reference data by continuously monitoring and storing visual and structural information of legitimate sites, eliminating the need for complex manual data management processes.
Solution Approach 2:
The patent uses copying by creating a database of replicated visual and structural information from legitimate webpages. Instead of storing original complex webpage data, the system stores captured representations (snapshots and structural data) that can be efficiently compared against suspected phishing sites, simplifying data management while maintaining detection accuracy.
Data Source
AI summary
A processor controlled hybrid method, an apparatus and a computer readable storage medium for identifying a phishing webpage are provided. The method comprises capturing overall visual information and overall structural information about a webpage being browsed by a user, comparing the overall visual information and overall structural information of the webpage with overall visual information and overall structural information of a legitimate webpage or a phishing webpage stored in a webpage database, calculating a measure of similarity, assessing the measure on the basis of a pre-determined threshold and concluding the measure of similarity is above the pre-determined threshold, thereby identifying a phishing webpage. The method may also provide for collecting and comparing visual information and, optionally, structural information.


