Threat URL Forensics Clustering for Cross-Space Similarity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cybersecurity systems struggle to efficiently cluster malicious URLs based on their behaviors, relying on URL rollups or manual analysis, which are time-consuming and inconsistent.
Innovation Solution
A forensics-based clustering method that analyzes the behaviors of URLs in a sandboxed environment, generating similarity scores based on shared forensic elements to group similar threats using graph-based techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If URL rollups are used to cluster threats, then URLs that are similar or part of the same URL hierarchy can be grouped together, but threat URLs across different URL spaces cannot be identified as related
Solution Approach 1:
The patent transitions from clustering URLs based on their hierarchical structure (one dimension) to clustering based on forensic behaviors observed in sandboxed environments (another dimension). This allows threats across different URL spaces to be identified as related by comparing their actual malicious behaviors, files distributed, and techniques used, rather than being constrained by URL structure similarities.
2Measurement precision
If manual analysis is used to identify groups of related threats, then accurate threat clustering can be achieved, but the process becomes costly and time-consuming
Solution Approach 1:
The system enables automated self-service threat clustering by having threats analyze themselves through sandboxed execution. Each URL is executed in an isolated environment and automatically generates forensic data about its behaviors. The system then automatically compares these forensic elements across multiple URLs and performs clustering without human intervention, achieving both accuracy and efficiency.
Solution Approach 2:
The patent replaces the mechanical process of manual analyst review with an automated computational system. Instead of human analysts examining and grouping threats manually, the system uses automated sandboxed execution, data extraction, and algorithmic clustering to perform the same function with greater speed and consistency.
3Productivity
If individual threat URLs are addressed separately, then each threat can be analyzed in detail, but user experience deteriorates and response efficiency decreases
Solution Approach 1:
The patent merges multiple related threat URLs into clusters based on their forensic similarities. Instead of presenting users with individual threat URLs that may be variants of the same malicious campaign, the system groups them together and allows for consolidated response actions. This improves user experience by reducing the number of individual threats users must review and increases productivity by enabling batch processing of related threats.
Data Source
AI summary
Systems, methods and products for identifying “similar” threats by clustering the threats based on corresponding forensics. A corpus of forensic data for a plurality of threat URLs is obtained by a threat protection system, the data including forensic elements corresponding to each threat URLs. For each pair of threat URLs, the corresponding forensic elements are examined to identify shared forensic elements. A similarity score is then generated for the pair of threat URLs based on the comparison of the corresponding forensic elements, including both malicious and non-malicious elements. Based on the similarity score generated for each pair of threat URLs, clusters of the threat URLs are identified, with each cluster including a subset of the plurality of threat URLs. Clusters of URLs similar to a selected URL may be identified by accessing the threat cluster information using a similar-threat search interface or through internal APIs of the threat protection system.


