Phishing Detection via Web Page Content Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing sophistication and diversity of malware and phishing attacks make it difficult for users to securely access information channels like web browsers, email, and texts, with existing anti-phishing solutions often providing ad hoc approaches lacking comprehensive mitigation strategies.

Innovation Solution

The Malware and Phishing Detection and Mediation (MAPDAM) platform, which includes ingestion, detection, and action stages, uses a dynamically configurable sequence of detection engines employing engineered rules, machine learning, and computer vision to identify and mitigate malware and phishing threats by analyzing URLs, website certificates, and web page content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional anti-phishing solutions are used, then some phishing detection capability is provided, but the solutions are ad hoc and lack comprehensive mitigation strategies

Engineering Contradiction:
Improvephishing detection reliabilityVSAvoidcomprehensive mitigation capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system segments phishing detection into multiple specialized detection engines, each focusing on specific indicators (URL analysis, certificate validation, web content analysis, screenshot comparison). This modular segmentation allows each engine to specialize in particular detection aspects while collectively providing comprehensive mitigation coverage that ad hoc solutions cannot achieve.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The platform creates a universal anti-phishing system that performs multiple functions: URL analysis, certificate validation, web content analysis, screenshot capture and comparison, and automated mitigation. This multi-functional universal platform replaces fragmented ad hoc solutions with a single comprehensive system that adapts to various phishing vectors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Object-affected harmful factors

If sophisticated phishing attacks are used, then the deceptive capability increases, but the difficulty of identification and prevention increases

Engineering Contradiction:
Improvephishing deception capabilityVSAvoidphishing identification difficulty
Core Design Contradiction:
Object-affected harmful factorsVSDifficulty of detecting and measuring

Solution Approach 1:

The system employs nested detection layers where multiple detection engines operate in sequence, with each engine nesting within the broader detection framework. URL analysis nests within the detection phase, which nests within the broader platform that coordinates with mitigation services. This nested structure allows progressive deepening of analysis for sophisticated attacks without overwhelming complexity.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The system introduces intermediary detection mechanisms between the user and phishing content, including screenshot capture as an intermediary representation of web content, and automated analysis engines that mediate between suspicious indicators and final phishing determination. These intermediaries make detection of sophisticated attacks more feasible by breaking down complex analysis into manageable steps.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual phishing identification methods are used, then some detection capability is provided, but the process lacks automation and efficiency

Engineering Contradiction:
Improvephishing detection accuracyVSAvoiddetection automation level
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system implements self-service automation where the platform automatically performs URL analysis, certificate validation, web content analysis, screenshot capture, and comparison without requiring manual intervention. The system serves itself by automatically generating detection results and initiating mitigation actions, achieving high automation while maintaining precision through multiple automated detection engines.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces manual mechanical detection processes with automated computational engines. Instead of manual URL analysis, automated engines perform cryptographic certificate validation and algorithmic web content analysis. This substitution of mechanical manual processes with automated computational systems maintains high detection precision while dramatically increasing automation level.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If comprehensive detection analysis is performed, then detection accuracy improves, but the system complexity increases

Engineering Contradiction:
Improvephishing detection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments comprehensive detection into distinct modular engines (URL analysis engine, certificate validation engine, web content analysis engine, screenshot comparison engine), each handling specific detection tasks. This segmentation maintains high overall detection accuracy while managing system complexity through modular design, where each segment can be developed and maintained independently.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12021894B2Phishing detection based on modeling of web page content
Publication Date: 2024.06.25 PAYPAL INC
  • US12021894B2 patent drawing
  • US12021894B2 patent drawing
  • US12021894B2 patent drawing

AI summary

A method for phishing detection based on modeling of web page content is discussed. The method includes accessing suspect web page content of a suspect Uniform Resource Locator (URL). The method includes generating an exemplary model based on an exemplary configuration for an indicated domain associated with the suspect URL, where the exemplary model indicates structure and characteristics of an example web page of the indicated domain. The method includes generating a suspect web page model that indicates structure and characteristics of the suspect web page content. The method includes performing scoring functions for the potential phishing web page content based on the suspect web page model, where some of the scoring functions use the exemplary model to perform analysis to generate respective results. The method includes generating a web page content phishing score based on results from the scoring functions.