Phishing Detection via Linguistic Pattern Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current phishing detection methods require frequent updates and lack accuracy in distinguishing between phishing and legitimate emails without visiting potential phishing websites, leading to inefficiencies and vulnerabilities.

Innovation Solution

A comprehensive natural language processing-based scheme, PhishNet-NLP, that analyzes email headers, links, and text to classify emails as phishing or legitimate, using contextual information and machine learning classifiers to differentiate between 'actionable' and 'informational' emails, operating between the mail transfer agent and mail user agent to prevent harmful link clicks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning techniques are used for phishing detection, then detection accuracy is improved, but filters need to be updated frequently

Engineering Contradiction:
Improvephishing detection accuracyVSAvoidtime for filter updates
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional machine learning classifiers with a rule-based system that uses linguistic patterns and semantic analysis. Instead of relying on trained models that require periodic retraining with new phishing samples, the system uses manually crafted rules based on common phishing language patterns, eliminating the need for frequent filter updates while maintaining detection accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If heuristic algorithms with simple text analysis are used, then device complexity is reduced, but detection accuracy deteriorates

Engineering Contradiction:
Improveanalysis process complexityVSAvoidphishing detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer of linguistic analysis rules that bridge simple heuristic analysis and complex machine learning. The system uses pattern matching rules, semantic keyword analysis, and linguistic structure verification as intermediate steps to enhance simple text analysis without requiring full machine learning complexity, achieving improved accuracy with moderate system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If phishing detection visits potential phishing websites, then detection reliability is improved, but security vulnerabilities increase

Engineering Contradiction:
Improvephishing detection reliabilityVSAvoidsecurity vulnerabilities
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent performs preliminary analysis of email content, headers, and linguistic patterns before any website interaction occurs. By detecting phishing indicators through language analysis, URL pattern recognition, and semantic evaluation of email text, the system can identify and block phishing emails without ever visiting the embedded links, thus maintaining high detection reliability while eliminating security vulnerabilities associated with website probing

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10404745B2Automatic phishing email detection based on natural language processing techniques
Publication Date: 2019.09.03 VERMA RAKESH
  • US10404745B2 patent drawing
  • US10404745B2 patent drawing
  • US10404745B2 patent drawing

AI summary

A comprehensive scheme to detect phishing emails using features that are invariant and fundamentally characterize phishing. Multiple embodiments are described herein based on combinations of text analysis, header analysis, and link analysis, and these embodiments operate between a user's mail transfer agent (MTA) and mail user agent (MUA). The inventive embodiment, PhishNet-NLP™, utilizes natural language techniques along with all information present in an email, namely the header, links, and text in the body. The inventive embodiment, PhishSnag™, uses information extracted form the embedded links in the email and the email headers to detect phishing. The inventive embodiment, Phish-Sem™ uses natural language processing and statistical analysis on the body of labeled phishing and non-phishing emails to design four variants of an email-body-text only classifier. The inventive scheme is designed to detect phishing at the email level.