Threat URL Forensics Clustering for Cross-Space Similarity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cybersecurity systems struggle to efficiently cluster malicious URLs based on their behaviors, relying on URL rollups or manual analysis, which are time-consuming and inconsistent.

Innovation Solution

A forensics-based clustering method that analyzes the behaviors of URLs in a sandboxed environment, generating similarity scores based on shared forensic elements to group similar threats using graph-based techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If URL rollups are used to cluster threats, then URLs that are similar or part of the same URL hierarchy can be grouped together, but threat URLs across different URL spaces cannot be identified as related

Engineering Contradiction:
Improveability to identify related threats across URL spacesVSAvoidclustering method limitations
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transitions from clustering URLs based on their hierarchical structure (one dimension) to clustering based on forensic behaviors observed in sandboxed environments (another dimension). This allows threats across different URL spaces to be identified as related by comparing their actual malicious behaviors, files distributed, and techniques used, rather than being constrained by URL structure similarities.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If manual analysis is used to identify groups of related threats, then accurate threat clustering can be achieved, but the process becomes costly and time-consuming

Engineering Contradiction:
Improveaccuracy of threat clusteringVSAvoidtime required for manual threat analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service threat clustering by having threats analyze themselves through sandboxed execution. Each URL is executed in an isolated environment and automatically generates forensic data about its behaviors. The system then automatically compares these forensic elements across multiple URLs and performs clustering without human intervention, achieving both accuracy and efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual analyst review with an automated computational system. Instead of human analysts examining and grouping threats manually, the system uses automated sandboxed execution, data extraction, and algorithmic clustering to perform the same function with greater speed and consistency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If individual threat URLs are addressed separately, then each threat can be analyzed in detail, but user experience deteriorates and response efficiency decreases

Engineering Contradiction:
Improvethreat response efficiencyVSAvoiduser experience with threat handling
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent merges multiple related threat URLs into clusters based on their forensic similarities. Instead of presenting users with individual threat URLs that may be variants of the same malicious campaign, the system groups them together and allows for consolidated response actions. This improves user experience by reducing the number of individual threats users must review and increases productivity by enabling batch processing of related threats.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12513192B2Identifying threat similarity using forensics clustering
Publication Date: 2025.12.30 GOLDMAN SACHS BANK USA
  • US12513192B2 patent drawing
  • US12513192B2 patent drawing
  • US12513192B2 patent drawing

AI summary

Systems, methods and products for identifying “similar” threats by clustering the threats based on corresponding forensics. A corpus of forensic data for a plurality of threat URLs is obtained by a threat protection system, the data including forensic elements corresponding to each threat URLs. For each pair of threat URLs, the corresponding forensic elements are examined to identify shared forensic elements. A similarity score is then generated for the pair of threat URLs based on the comparison of the corresponding forensic elements, including both malicious and non-malicious elements. Based on the similarity score generated for each pair of threat URLs, clusters of the threat URLs are identified, with each cluster including a subset of the plurality of threat URLs. Clusters of URLs similar to a selected URL may be identified by accessing the threat cluster information using a similar-threat search interface or through internal APIs of the threat protection system.