Data Enrichment Engine for Abbreviated Domain Name Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Abbreviated domain names provide limited contextual information, making it difficult to identify relevant or spoofed domains, and existing technologies struggle to efficiently process and classify these names for digital risk analysis.

Innovation Solution

The data enrichment system processes abbreviated domain names by extracting textual content from corresponding web pages, determining sets of words with initial characters matching the domain name, and querying WHOIS servers to identify candidate domain names owned by the same entity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If abbreviated domain names are used to represent brands or entities, then domain name brevity and ease of memorization are improved, but the ability to accurately identify and classify domains for digital risk analysis deteriorates

Engineering Contradiction:
Improvedomain name lengthVSAvoiddomain classification accuracy
Core Design Contradiction:
Length of moving objectVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces web content source code as an intermediary between the abbreviated domain name and the full contextual meaning. By extracting textual content from the web page associated with the abbreviated domain name, the system resolves the ambiguity of multi-word abbreviations and establishes accurate relationships between domain names and their intended meanings, thereby improving classification accuracy without increasing domain name length.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple word combinations are associated with abbreviated domain names, then the representational flexibility of domain names is improved, but the difficulty of determining relevant or spoofed domains increases

Engineering Contradiction:
Improvedomain name representational flexibilityVSAvoiddomain relevance determination
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent extracts textual content from the web content source code of the abbreviated domain name and uses this extracted information to determine the actual meaning and ownership context. By taking out the contextual information from the web page, the system can accurately identify whether a domain is relevant or spoofed, resolving the ambiguity introduced by multiple possible word combinations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses feedback from web content analysis to refine domain classification. By extracting and analyzing the actual textual content from the web page, the system receives feedback about the true meaning and ownership of the abbreviated domain name, which then feeds back into the classification process to improve accuracy in identifying relevant or spoofed domains.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If limited characters are used in abbreviated domain names, then the simplicity and memorability of domain names are improved, but the amount of contextual information available for analysis is reduced

Engineering Contradiction:
Improvedomain name simplicityVSAvoidcontextual information quantity
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent moves the contextual information from the domain name itself (one dimension) to the web content source code (another dimension). By accessing and analyzing the web page content associated with the abbreviated domain name, the system recovers the lost contextual information without compromising the simplicity of the abbreviated domain name structure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12242548B2Data enrichment systems and methods for abbreviated domain name classification
Publication Date: 2025.03.04 GOLDMAN SACHS BANK USA
  • US12242548B2 patent drawing
  • US12242548B2 patent drawing
  • US12242548B2 patent drawing

AI summary

To find enriching contextual information for an abbreviated domain name, a data enrichment engine can comb through web content source code corresponding to the abbreviated domain name. From textual content in the web content source code, the data enrichment engine can identify words with initial characters that match characters of the abbreviated domain name to thereby establish a relationship there-between. This relationship can facilitate more accurate and efficient domain name classification. The data enrichment engine can query a WHOIS server to find out if candidate domains having initial characters that match the characters of the abbreviated domain name are registered to the same entity. If so, keywords can be extracted from the candidate domains and used to find more relevant domains for domain risk analysis and detection. Candidate domains determined by the data enrichment engine can be provided to a downstream computing facility such as a domain filter.