Spam Email Clustering via Neural Network Feature Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spam detection methods, such as signature-based and machine learning approaches, suffer from high type one and type two errors, leading to inefficiencies in identifying spam emails and causing network congestion and financial losses.

Innovation Solution

A method using a neural network classifier that selects characteristics from email messages, calculates feature vectors, and generates clusters to improve spam detection by reducing errors, employing techniques like orthogonality preservation and batch-normalization layers to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If signature-based or heuristic approaches are used for spam detection, then type one error rate is reduced, but type two error rate increases significantly

Engineering Contradiction:
Improvetype one error rateVSAvoidtype two error rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments spam detection into multiple independent classifiers, each specializing in different aspects: neural network classifier for semantic analysis, Bayesian classifier for probabilistic assessment, and rule-based classifier for pattern matching. This segmentation allows each classifier to optimize for different error types, collectively reducing both type one and type two errors compared to single-approach systems

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite detection system combining three different classification approaches (neural network, Bayesian, rule-based) into a unified spam detection mechanism. This composite structure leverages the strengths of each approach while compensating for their individual weaknesses, achieving better overall accuracy and lower error rates than any single approach alone

Inventive Principle:
Principle #40Composite materials

2Productivity

If machine learning classifiers are used to reduce time lag, then generalization capacity increases, but false positive rate increases

Engineering Contradiction:
Improvetime lag reductionVSAvoidfalse positive rate
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary validation layer that receives outputs from the neural network classifier and applies additional rule-based filtering and Bayesian probability assessment. This intermediary mechanism validates the neural network's decisions, correcting false positives while maintaining the fast processing speed advantage of machine learning approaches

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback loops where classification results are continuously evaluated and used to refine future classifications. The system monitors false positives and adjusts classifier thresholds and parameters accordingly, allowing the system to maintain high productivity while progressively reducing false positive rates through learned experience

Inventive Principle:
Principle #23Feedback

3Measurement precision

If human analysts are used to improve heuristic approaches, then detection accuracy improves, but processing time increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service automation where the system automatically performs tasks that would traditionally require human analysts. The multi-classifier system autonomously analyzes emails, detects patterns, and makes classification decisions without human intervention, achieving high detection accuracy while eliminating the time loss associated with manual review

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of human analysis with automated computational systems. The neural network and Bayesian classifiers perform complex pattern recognition and probability calculations that previously required human cognitive processes, achieving comparable or superior accuracy while dramatically reducing processing time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20220294751A1System and method for clustering emails identified as spam
Publication Date: 2022.09.15 AO KASPERSKY LAB
  • US20220294751A1 patent drawing
  • US20220294751A1 patent drawing
  • US20220294751A1 patent drawing

AI summary

Disclosed herein are systems and methods for clustering email messages identified as spam using a trained classifier. In one aspect, an exemplary method comprises, selecting at least two characteristics from each received email message, for each received email message, using a classifier containing a neural network, determining whether or not the email message is a spam based on the at least two characteristics of the email message, for each email message determined as being a spam email, calculating a feature vector, the feature vector being calculated at a final hidden layer of the neural network, and generating one or more clusters of the email messages identified as spam based on similarities of the feature vectors calculated at the final hidden layer of the neural network.