Email Authorship Classification via Header Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional email classification methods require manual analysis by security experts, are inefficient, and only apply data mining techniques to email bodies, limiting their effectiveness in classifying email authors, especially in cases of small email volumes or diverse feature information.

Innovation Solution

An email authorship classification method and apparatus that analyzes email headers to extract feature information such as location, language, and time details, converting them into a feature dataset for use in a classification learning algorithm, including a bagging classification algorithm to generate a classification model for efficient author classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis by security experts is used to classify email authors, then classification accuracy can be maintained, but analysis efficiency and productivity deteriorate

Engineering Contradiction:
Improveclassification accuracyVSAvoidanalysis efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated email authorship classification through machine learning models that process email headers independently without requiring security expert intervention for each email. The classification model automatically extracts features from email headers (such as X-ClientIP, Received-SPF, authentication results) and classifies authors based on these features, making the system self-sufficient while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual analysis process with an automated computational system. Instead of security experts manually examining email headers, a machine learning classification model processes email headers automatically. The system substitutes human cognitive work with algorithmic processing, achieving both high speed and consistent accuracy in author classification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If data mining techniques are applied only to email bodies, then processing speed is maintained, but classification effectiveness and measurement precision deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoidclassification effectiveness
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts and utilizes email header information that contains authorship-specific features. Instead of relying solely on email body content, the system extracts features from headers such as X-ClientIP, Received-SPF, authentication results, and other header fields that contain metadata about the email's origin and transmission path. This extraction of header information provides additional discriminatory features for accurate author classification while maintaining processing efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent adds a new dimension to email analysis by incorporating header field information alongside or instead of body content analysis. This dimensional expansion allows the system to classify authors based on metadata characteristics (IP addresses, authentication results, header structures) that are independent of email body content, thereby improving classification effectiveness without sacrificing processing speed.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Quantity of substance

If conventional classification methods are used with small email volumes, then resource consumption is low, but classification reliability and measurement precision deteriorate

Engineering Contradiction:
Improveemail volumeVSAvoidclassification reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent changes the parameters used for classification by focusing on specific header field features rather than relying on large volumes of email body text. The classification model is trained to recognize patterns in header parameters (such as IP address formats, authentication result distributions, header structure characteristics) that remain consistent even with small datasets. This parameter transformation allows reliable classification with limited email volumes.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If manual analysis of each email header is performed, then detailed feature extraction is achieved, but device complexity and operation difficulty increase

Engineering Contradiction:
Improvefeature extraction completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system automatically identifies and extracts relevant features from email headers without requiring manual configuration or expert intervention. The classification model self-adapts to identify which header fields contain authorship information, automatically processing features such as X-ClientIP, Received-SPF, authentication results, and other header elements. This automation reduces system complexity from the user's perspective while maintaining comprehensive feature extraction.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11321630B2Method and apparatus for providing e-mail authorship classification
Publication Date: 2022.05.03 AGENCY FOR DEFENSE DEV
  • US11321630B2 patent drawing
  • US11321630B2 patent drawing
  • US11321630B2 patent drawing

AI summary

There is provided an email authorship classification apparatus. The apparatus includes an information analysis unit configured to analyze header field information in an attribute header of each of emails and extract feature field information related to an authorship of each of the emails from each of the header field information and an information conversion unit configured to convert the feature field information into a feature data set for inputting a learning model thereto. The apparatus further includes a learning model unit configured to generate a classification model for classifying the emails by author by applying a learning process to the feature data set.