Data Enrichment Engine for Abbreviated Domain Name Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Abbreviated domain names provide limited contextual information, making it difficult to identify relevant or spoofed domains, and existing technologies struggle to efficiently process and classify these names for digital risk analysis.
Innovation Solution
The data enrichment system processes abbreviated domain names by extracting textual content from corresponding web pages, determining sets of words with initial characters matching the domain name, and querying WHOIS servers to identify candidate domain names owned by the same entity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Length of moving object
If abbreviated domain names are used to represent brands or entities, then domain name brevity and ease of memorization are improved, but the ability to accurately identify and classify domains for digital risk analysis deteriorates
Solution Approach 1:
The patent introduces web content source code as an intermediary between the abbreviated domain name and the full contextual meaning. By extracting textual content from the web page associated with the abbreviated domain name, the system resolves the ambiguity of multi-word abbreviations and establishes accurate relationships between domain names and their intended meanings, thereby improving classification accuracy without increasing domain name length.
2Adaptability or versatility
If multiple word combinations are associated with abbreviated domain names, then the representational flexibility of domain names is improved, but the difficulty of determining relevant or spoofed domains increases
Solution Approach 1:
The patent extracts textual content from the web content source code of the abbreviated domain name and uses this extracted information to determine the actual meaning and ownership context. By taking out the contextual information from the web page, the system can accurately identify whether a domain is relevant or spoofed, resolving the ambiguity introduced by multiple possible word combinations.
Solution Approach 2:
The system uses feedback from web content analysis to refine domain classification. By extracting and analyzing the actual textual content from the web page, the system receives feedback about the true meaning and ownership of the abbreviated domain name, which then feeds back into the classification process to improve accuracy in identifying relevant or spoofed domains.
3Ease of operation
If limited characters are used in abbreviated domain names, then the simplicity and memorability of domain names are improved, but the amount of contextual information available for analysis is reduced
Solution Approach 1:
The patent moves the contextual information from the domain name itself (one dimension) to the web content source code (another dimension). By accessing and analyzing the web page content associated with the abbreviated domain name, the system recovers the lost contextual information without compromising the simplicity of the abbreviated domain name structure.
Data Source
AI summary
To find enriching contextual information for an abbreviated domain name, a data enrichment engine can comb through web content source code corresponding to the abbreviated domain name. From textual content in the web content source code, the data enrichment engine can identify words with initial characters that match characters of the abbreviated domain name to thereby establish a relationship there-between. This relationship can facilitate more accurate and efficient domain name classification. The data enrichment engine can query a WHOIS server to find out if candidate domains having initial characters that match the characters of the abbreviated domain name are registered to the same entity. If so, keywords can be extracted from the candidate domains and used to find more relevant domains for domain risk analysis and detection. Candidate domains determined by the data enrichment engine can be provided to a downstream computing facility such as a domain filter.


