Threat Actor Identification Through Domain Data Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems struggle to automate the identification of threat actors targeting specific brands or multiple brands with similar attack techniques, as threat actors can obscure their identities using various domain names and tactics, leading to unreliable and subjective manual identification.

Innovation Solution

A threat actor identification system that analyzes domain data, including web page content, domain registration information, and infrastructure data to generate domain clusters, determining similarities and associating clusters with threat actors, and providing indications to brand owners.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual identification of threat actors is performed, then flexibility in analysis can be maintained, but reliability and objectivity of identification deteriorate due to subjectivity

Engineering Contradiction:
Improveflexibility in analysisVSAvoidobjectivity of identification
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent replaces manual mechanical analysis with automated computational analysis using machine learning models and algorithms. The system automatically processes domain data, performs clustering, and identifies threat actors through objective computational methods rather than subjective human judgment, thereby improving reliability while maintaining operational flexibility through programmable parameters.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an automated analysis system as an intermediary between raw domain data and threat actor identification. This intermediary system uses machine learning models, clustering algorithms, and structured data processing to objectively transform unstructured domain information into reliable threat assessments, eliminating direct human subjectivity while preserving analytical flexibility through configurable system parameters.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If automated identification systems are implemented, then objectivity and reliability improve, but complexity of the system increases due to multiple data sources and analysis methods

Engineering Contradiction:
Improveobjectivity of identificationVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex automated identification system into distinct functional modules: data collection module, data processing module, clustering module, and identification module. Each module handles specific tasks independently, processing different types of domain data through specialized algorithms before integrating results. This segmentation manages system complexity by creating manageable, independent components while maintaining overall reliability through systematic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal automated analysis platform that handles multiple data sources (domain registration data, web page content, infrastructure data), various analysis methods (clustering, machine learning, pattern recognition), and different threat actor types through a single integrated system. This multi-functional approach manages complexity by consolidating diverse capabilities into one cohesive system rather than requiring separate systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If comprehensive domain data analysis is performed across multiple sources, then identification accuracy improves, but time and computational resources required increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and structuring domain data before main analysis operations. The system collects and organizes domain registration data, web page content, and infrastructure data in advance, creating structured datasets ready for clustering and identification. This preliminary data preparation reduces the time required during actual threat actor identification while maintaining high accuracy through comprehensive pre-analyzed information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates simplified representations or copies of complex domain data through structured data models and feature extractions. Instead of analyzing raw, unstructured domain information directly, the system generates condensed data representations that capture essential characteristics for threat identification. This copying approach maintains identification accuracy by preserving critical information while reducing computational complexity and analysis time.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3937465B1Threat actor identification systems and methods
Publication Date: 2025.09.03 PROOFPOINT INC
  • EP3937465B1 patent drawingFigure 1
  • EP3937465B1 patent drawingFigure 2
  • EP3937465B1 patent drawingFigure 3

AI summary

A threat actor identification system that obtains (405) domain data for a set of domains, generates (415) domain clusters, determines (420) whether the domain clusters are associated with threat actors, and presents (440) domain data for the clusters that are associated with threat actors to brand owners that are associated with the threat actors. The clusters may be generated based on similarities in web page content, domain registration information, and/or domain infrastructure information. For each cluster, a clustering engine determines (425) whether the cluster is associated with a threat actor, and for clusters that are associated with threat actors, corresponding domain information is stored (430) for presentation to brand owners to whom the threat actor poses a threat.