Layout Vector Spam Filtering via Text Line Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional anti-spam filtering methods are ineffective against experienced spammers who use countermeasures such as misspelling words, digital images, and unrelated text in spam messages, making it difficult to accurately classify unsolicited commercial electronic communications.

Innovation Solution

A spam filtering method that analyzes the text line structure of electronic communications to produce a layout representation, determining whether a message is spam or non-spam based on the relative positions of text line types within the communication, using a layout analysis engine that generates a layout feature vector and classifies messages by analyzing positional relationships between different text line types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional filtering methods (black-listing, white-listing, keyword filtering) are used, then implementation is simple and fast, but accuracy deteriorates due to spammer countermeasures

Engineering Contradiction:
Improvespam classification accuracyVSAvoidfiltering system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from analyzing content at the word/character level to analyzing the spatial layout dimension of email messages. By extracting positional relationships between text lines, images, and other elements, the system creates a new feature space that is independent of spammer obfuscation techniques. This dimensional shift allows the filter to identify spam based on structural patterns rather than semantic content, resolving the contradiction between maintaining simplicity and improving accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If spam filtering accuracy is improved by analyzing more message features, then spam identification becomes more accurate, but processing time increases

Engineering Contradiction:
Improvespam detection accuracyVSAvoidmessage processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts specific layout features from email messages, focusing on positional relationships between text lines, images, and other elements. Rather than analyzing all message features, the system selectively extracts spatial arrangement characteristics such as line positioning, spacing, and relative locations. This extraction approach maintains high accuracy by focusing on discriminative features while reducing processing overhead compared to comprehensive analysis methods.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If spammer uses countermeasures (misspelling, images, unrelated text), then spam evasion capability improves, but detectability by conventional filters deteriorates

Engineering Contradiction:
Improvespammer evasion capabilityVSAvoidspam detectability
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

Instead of analyzing the semantic content of messages (the traditional approach), the patent inverts the analysis by focusing on the spatial layout and structural arrangement of elements. By examining positional relationships between text lines, images, and other components rather than their textual content, the system detects spam patterns that are independent of the message's semantic meaning. This inversion makes the detection method resilient to spammer countermeasures that alter content while preserving structural patterns.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS8065379B1Line-structure-based electronic communication filtering systems and methods
Publication Date: 2011.11.22 BITDEFENDER IPR MANAGEMENT
  • US8065379B1 patent drawing
  • US8065379B1 patent drawing
  • US8065379B1 patent drawing

AI summary

In some embodiments, a layout-based electronic communication classification (e.g. spam filtering) method includes generating a layout vector characterizing a layout of a message, assigning the message to a selected cluster according to a hyperspace distance between the layout vector and a central vector of the selected cluster, and classifying the message (e.g. labeling as spam or non-spam) according to the selected cluster. The layout vector is a message representation characterizing a set of relative positions of metaword substructures of the message, as well as metaword substructure counts. Examples of metaword substructures include MIME parts and text lines. For example, a layout vector may have a first component having scalar axes defined by numerical layout feature counts (e.g. numbers of lines, blank lines, links, email addresses), and a second vector component including a line-structure list and a formatting part (e.g. MIME part) list.