Image-Based Deep Learning for Low-False-Positive Phishing PDF Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently detecting and mitigating phishing attacks, particularly those carried out through PDF documents, as traditional methods struggle to identify evolving phishing techniques and minimize false positives.

Innovation Solution

An image-based deep learning approach is employed to detect phishing PDFs, where pages of the PDF are converted into images and used to train a deep neural network to distinguish between benign and phishing documents, thereby reducing false positives and efficiently analyzing PDFs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional detection methods are used to identify phishing PDFs, then the detection process is simple and fast, but the detection accuracy is low and false positives are high due to evolving phishing techniques

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical detection methods (rule-based filtering, keyword matching) with an image-based deep learning system. PDF pages are converted to images and processed through convolutional neural networks, substituting automated image recognition for manual analysis rules, thereby improving detection accuracy while maintaining operational efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of analysis from textual/metadata features to visual image features. By converting PDF content into image representations and analyzing visual patterns, the system adapts to evolving phishing techniques that may preserve visual characteristics while changing textual content, thus improving detection precision

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional analysis methods are used, then the processing speed is fast, but the ability to detect evolving phishing techniques is poor

Engineering Contradiction:
Improveability to detect evolving techniquesVSAvoidanalysis time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary conversion of PDF pages to images during the analysis phase, preparing visual data in advance for deep learning processing. This preliminary action enables the system to quickly compare visual patterns against trained models, improving adaptability to new phishing techniques without significant time penalty

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates image copies of PDF pages for analysis, allowing the deep learning model to process visual representations without modifying the original PDF structure. This copying approach enables rapid iteration and training on phishing samples while preserving the integrity of source documents

Inventive Principle:
Principle #26Copying

3Measurement precision

If deep learning approaches are implemented, then detection accuracy improves, but computational resources and processing time increase

Engineering Contradiction:
Improvephishing detection accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the PDF analysis task into distinct phases: conversion to images, feature extraction, and classification. By dividing the processing workflow, the system can apply computationally intensive deep learning only where necessary while using lighter processing for other aspects, thereby reducing overall energy consumption while maintaining high detection accuracy

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12348560B2Detecting phishing PDFs with an image-based deep learning approach
Publication Date: 2025.07.01 PALO ALTO NETWORKS INC
  • US12348560B2 patent drawing
  • US12348560B2 patent drawing
  • US12348560B2 patent drawing

AI summary

The detection of phishing Portable Document Format (PDF) files using an image-based deep learning approach is disclosed. A PDF document that includes a Universal Resource Locator is received. A likelihood that the received PDF document represents a phishing threat is determined, at least in part, by using an image based model. A verdict for the PDF document is provided as output based at least in part on the determined likelihood.