Threat Actor Identification Through Domain Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems struggle to automate the identification of threat actors targeting specific brands or multiple brands with similar attack techniques, as threat actors can obscure their identities using various domain names and tactics, leading to unreliable and subjective manual identification.
Innovation Solution
A threat actor identification system that analyzes domain data, including web page content, domain registration information, and infrastructure data to generate domain clusters, determining similarities and associating clusters with threat actors, and providing indications to brand owners.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual identification of threat actors is performed, then flexibility in analysis can be maintained, but reliability and objectivity of identification deteriorate due to subjectivity
Solution Approach 1:
The patent replaces manual mechanical analysis with automated computational analysis using machine learning models and algorithms. The system automatically processes domain data, performs clustering, and identifies threat actors through objective computational methods rather than subjective human judgment, thereby improving reliability while maintaining operational flexibility through programmable parameters.
Solution Approach 2:
The patent introduces an automated analysis system as an intermediary between raw domain data and threat actor identification. This intermediary system uses machine learning models, clustering algorithms, and structured data processing to objectively transform unstructured domain information into reliable threat assessments, eliminating direct human subjectivity while preserving analytical flexibility through configurable system parameters.
2Reliability
If automated identification systems are implemented, then objectivity and reliability improve, but complexity of the system increases due to multiple data sources and analysis methods
Solution Approach 1:
The patent segments the complex automated identification system into distinct functional modules: data collection module, data processing module, clustering module, and identification module. Each module handles specific tasks independently, processing different types of domain data through specialized algorithms before integrating results. This segmentation manages system complexity by creating manageable, independent components while maintaining overall reliability through systematic processing.
Solution Approach 2:
The patent implements a universal automated analysis platform that handles multiple data sources (domain registration data, web page content, infrastructure data), various analysis methods (clustering, machine learning, pattern recognition), and different threat actor types through a single integrated system. This multi-functional approach manages complexity by consolidating diverse capabilities into one cohesive system rather than requiring separate systems for each function.
3Measurement precision
If comprehensive domain data analysis is performed across multiple sources, then identification accuracy improves, but time and computational resources required increase
Solution Approach 1:
The patent performs preliminary actions by pre-processing and structuring domain data before main analysis operations. The system collects and organizes domain registration data, web page content, and infrastructure data in advance, creating structured datasets ready for clustering and identification. This preliminary data preparation reduces the time required during actual threat actor identification while maintaining high accuracy through comprehensive pre-analyzed information.
Solution Approach 2:
The patent creates simplified representations or copies of complex domain data through structured data models and feature extractions. Instead of analyzing raw, unstructured domain information directly, the system generates condensed data representations that capture essential characteristics for threat identification. This copying approach maintains identification accuracy by preserving critical information while reducing computational complexity and analysis time.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A threat actor identification system that obtains (405) domain data for a set of domains, generates (415) domain clusters, determines (420) whether the domain clusters are associated with threat actors, and presents (440) domain data for the clusters that are associated with threat actors to brand owners that are associated with the threat actors. The clusters may be generated based on similarities in web page content, domain registration information, and/or domain infrastructure information. For each cluster, a clustering engine determines (425) whether the cluster is associated with a threat actor, and for clusters that are associated with threat actors, corresponding domain information is stored (430) for presentation to brand owners to whom the threat actor poses a threat.