Dynamic IT Asset Discovery With ML Domain Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing security measures and tools are inadequate in providing a comprehensive and up-to-date overview of an entity's IT systems due to the increasing complexity and interconnectedness of online services, rapid growth in vulnerabilities, and decreasing exploitation time, allowing malicious users with low expertise to hack large organizations.
Innovation Solution
A method involving target name searching, main domain query, subdomain enumeration, domain validation using machine learning, IP address and netblock discovery, and validation, to dynamically and automatically determine an entity's IT systems, utilizing data sources like merge and acquisition databases, web certificate databases, and machine learning algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing security measures and tools are used to monitor IT systems, then basic security coverage is maintained, but comprehensive and up-to-date overview of IT systems cannot be provided due to increasing complexity and interconnectedness
Solution Approach 1:
The patent segments the IT system discovery process into multiple specialized modules: domain name discovery module, IP address discovery module, vulnerability analysis module, and machine learning validation module. Each module handles a specific aspect of the complex task, making the overall system manageable and effective despite increasing IT system complexity
Solution Approach 2:
The patent introduces machine learning models as intermediary components that validate and prioritize discovered domains and vulnerabilities. These ML intermediaries process the large volume of discovered IT assets and provide refined security intelligence, bridging the gap between raw data collection and actionable security insights
2Productivity
If manual security assessment methods are used, then detailed analysis can be performed, but real-time discovery and evaluation of attack surface cannot be achieved
Solution Approach 1:
The patent performs preliminary domain name discovery and IP address enumeration continuously in the background before actual security assessments are needed. This preliminary action maintains an up-to-date inventory of IT assets, enabling rapid vulnerability analysis and response when security events occur, effectively reducing the exploitation time window
Solution Approach 2:
The patent implements continuous automated discovery and monitoring of IT systems, maintaining constant surveillance of the attack surface. This continuous action ensures that new domains, IP addresses, and vulnerabilities are detected and analyzed in real-time, keeping pace with rapidly evolving threats
3Measurement precision
If comprehensive IT system discovery is performed to identify all assets, then complete attack surface visibility is achieved, but false positives increase making validation difficult
Solution Approach 1:
The patent implements feedback loops where machine learning models validate discovered domains and IP addresses by analyzing patterns from previously confirmed assets. The system continuously learns from validation results, adjusting its discovery and prioritization strategies to reduce false positives while maintaining comprehensive asset detection
Solution Approach 2:
The patent dynamically changes validation parameters and thresholds based on the specific context of discovered assets. Different domains and IP ranges are validated using tailored criteria derived from machine learning analysis, allowing the system to maintain high detection completeness while adapting to reduce false positives in different network environments
Data Source
Figure 1
Figure 2
Figure 3a
AI summary
Method for automatic and dynamic determination of Information Technology systems belonging to an entity, comprising the following steps performed by a computer system (1): - a target name searching step (200), wherein a list of additional target names is generated by using an initial target name as input to a at least one first data source (2), in particular a merge and acquisition data base or a web search engine; - a main domain query step (300), wherein a list of main domains, in particular second-level domains, is generated by using the initial target name and/or the additional target names as input to at least one second data source (3), in particular a web certificate database; - a subdomain enumeration step (400), wherein a list of subdomains is generated by using the main domain names as input for at least one third data source (4), in particular a subdomain enumeration tool; - a domain validation step (500), wherein validation features (51) extracted from domains of the list of main domains and/or the list of subdomains are compared to at least one reference feature in order to determine the validity of the domains; wherein supportive features (52) extracted from the domains of the list of main domains and/or the list of subdomains are evaluated by a machine learning unit (53) in order to determine the validity of the domains; and wherein supportive features (52) extracted form valid domains for which validity was determined on basis of the validation features (51) are used to train the machine learning unit (53).