Spam Detection Using Learning Machines and Hashing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spam detection methods are ineffective due to evolving spam patterns, as they rely on heuristic processes and lack a systematic approach to threshold optimization, leading to inefficiencies in filtering and classification of electronic data streams.

Innovation Solution

A computer method and system utilizing learning machines such as neural networks, support vector machines, and stackable hash technology to classify and cluster electronic data streams, employing uniform filters and low-collision hashes for accurate text classification and spam detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional heuristic spam detection methods are used, then the system is simple to implement, but the detection accuracy decreases as spam patterns evolve

Engineering Contradiction:
Improveease of implementationVSAvoiddetection accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent replaces traditional heuristic-based mechanical filtering systems with machine learning-based detection systems. Neural networks and support vector machines are trained on email data to automatically learn spam patterns, substituting manual rule-based approaches with adaptive computational models that improve detection accuracy while handling evolving spam techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically adjusts detection parameters and thresholds based on learned patterns from training data. Instead of fixed heuristic rules, the machine learning models adapt their decision boundaries and sensitivity parameters to match current spam characteristics, allowing the system to maintain high accuracy as spam patterns change over time.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple filters and thresholds are used to improve detection accuracy, then the reliability increases, but the device complexity and processing time increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple filtering approaches into unified machine learning models. Instead of separately implementing multiple independent filters with individual thresholds, the system integrates detection logic into trained neural networks and support vector machines that process emails through a single cohesive framework, reducing system complexity while maintaining or improving accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The machine learning models serve multiple detection functions simultaneously. A single trained model can detect various types of spam patterns, phishing attempts, and malicious content through universal feature extraction and classification mechanisms, eliminating the need for separate specialized filters for each threat type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If traditional clustering methods are used, then the algorithm is simple, but it cannot effectively handle evolving spam patterns and regime changes

Engineering Contradiction:
Improvealgorithm simplicityVSAvoidadaptability to evolving patterns
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic clustering algorithms that adapt to changing data distributions over time. Unlike static traditional clustering methods, the system continuously re-trains machine learning models on new email data, allowing cluster boundaries and groupings to evolve dynamically as spam patterns and regime changes occur, maintaining effectiveness in detecting emerging threats.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where detection results and new spam samples are fed back into the training process. Machine learning models are continuously refined using feedback from actual spam encounters and detection performance metrics, enabling the clustering and classification algorithms to adapt and improve their ability to handle evolving patterns without requiring complete re-engineering.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS7574409B2Method, apparatus, and system for clustering and classification
Publication Date: 2009.08.11 SYSXNET
  • US7574409B2 patent drawing
  • US7574409B2 patent drawing
  • US7574409B2 patent drawing

AI summary

The invention provides a method, apparatus and system for classification and clustering electronic data streams such as email, images and sound files for identification, sorting and efficient storage. The inventive systems disclose labeling a document as belonging to a predefined class though computer methods that comprise the steps of identifying an electronic data stream using one or more learning machines and comparing the outputs from the machines to determine the label to associate with the data. The method further utilizes learning machines in combination with hashing schemes to cluster and classify documents. In one embodiment hash apparatuses and methods taxonomize clusters. In yet another embodiment, clusters of documents utilize geometric hash to contain the documents in a data corpus without the overhead of search and storage.