Guided Crawling of Toxic Network Neighborhoods for Malicious Domain Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cybersecurity measures struggle to detect newly registered malicious domains in real-time due to their reactive nature, leading to inefficiencies and vulnerabilities, while proactive methods lack precision and speed, resulting in either over-blocking legitimate domains or under-detecting malicious ones.
Innovation Solution
A proactive approach using guided crawling of attack infrastructure through unsupervised machine learning to identify toxic network neighborhoods, expand network graphs, and classify potentially malicious domains before any traffic reaches a firewall, integrating advanced data analytics and machine learning techniques to score newly registered domains based on comprehensive features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If proactive methods are used to discover malicious domains, then detection coverage is improved, but precision deteriorates leading to over-blocking legitimate domains
Solution Approach 1:
The system segments the domain discovery process into multiple stages: initial broad crawling to maximize coverage, followed by progressive filtering and scoring stages that progressively refine results. This multi-stage segmentation allows the system to maintain high detection coverage while progressively eliminating false positives through layered validation.
Solution Approach 2:
The system dynamically adjusts classification parameters and scoring thresholds based on the analysis stage and domain characteristics. By changing parameters adaptively across different processing stages, the system optimizes the balance between coverage and precision, reducing over-blocking while maintaining detection effectiveness.
2Measurement precision
If reactive cybersecurity measures are used, then precision is maintained, but detection speed deteriorates due to delayed response
Solution Approach 1:
The system performs preliminary analysis and classification of domains before they are actively used for attacks. By conducting guided crawling and machine learning-based classification in advance, the system prepares threat intelligence proactively, enabling faster response when threats are detected without sacrificing precision.
Solution Approach 2:
The system implements accelerated processing pathways for high-risk domains identified during guided crawling. By skipping unnecessary validation steps for domains showing clear malicious indicators, the system rushes through critical analysis phases to maintain detection speed while preserving precision through targeted deep analysis.
3Measurement precision
If comprehensive feature analysis is performed on newly registered domains, then detection accuracy is improved, but computational complexity increases
Solution Approach 1:
The system applies different levels of feature analysis to different domains based on their risk profiles and characteristics. High-risk domains receive comprehensive multi-feature analysis, while low-risk domains undergo simplified checking. This local differentiation of analysis quality maintains detection accuracy for critical threats while reducing overall computational complexity.
Solution Approach 2:
The system performs partial feature analysis on the majority of domains and reserves comprehensive analysis for a smaller subset of high-priority targets. By applying excessive (comprehensive) action only where necessary and partial action elsewhere, the system achieves high detection accuracy for critical threats while managing computational complexity through selective resource allocation.
Data Source
AI summary
The present application discloses a method, system, and computer system for proactively discovering malicious domains through a guided crawling of attack infrastructure. The method includes (i) determining a set of toxic network neighborhoods on the internet, (ii) expanding one or more network graphs for the set of toxic network neighborhoods; (iii) determining a set of domains expected to be malicious from the set of toxic network neighborhoods, and (iv) performing an action based at least in part on the set of domains expected to be malicious. A particular toxic network neighborhood shares a plurality of hosting environments.


