Guided Word Association for Combo-Squatted Domain Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Popular brands face challenges in identifying and mitigating combo-squatted domain names that divert internet traffic and potentially host malicious content, as existing methods are inefficient in detecting these domains within the vast internet space.

Innovation Solution

A computer program product that constructs a feature space from a corpus of text to represent words as vectors, detects domain name registrations by combining original domain names with seed words, and iteratively selects nearest neighbor candidate words based on vector distance, using a feedback mechanism to guide the detection of potentially combo-squatted domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional domain name detection methods are used to search the vast internet space, then the search coverage can be extensive, but the detection efficiency and accuracy are insufficient to identify combo-squatted domains effectively

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system pre-constructs a feature space from a corpus of text and pre-identifies seed words that are likely to be used in combo-squatted domains. This preliminary preparation allows the detection process to start with a focused set of candidate domains rather than searching the entire internet space, significantly improving both accuracy and efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses detected combo-squatted domains as feedback to iteratively refine the feature space and identify new seed words. This feedback mechanism continuously improves detection accuracy by learning from actual squatted domains and expanding the search vocabulary to include newly discovered patterns

Inventive Principle:
Principle #23Feedback

2Reliability

If the feature space construction includes more seed words to increase detection coverage, then more combo-squatted domains can be detected, but the computational complexity and processing time increase

Engineering Contradiction:
Improvedetection completenessVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically selects seed words from the corpus of text based on their relevance to the original domain name, rather than requiring manual curation of extensive word lists. This self-service approach to seed word selection reduces system complexity while maintaining comprehensive detection coverage

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts the feature space parameters based on the original domain name being protected. By changing the parameters (seed words, weightings) according to the specific domain context, the system achieves high detection completeness without requiring a universally complex configuration

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10938779B2Guided word association based domain name detection
Publication Date: 2021.03.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10938779B2 patent drawing
  • US10938779B2 patent drawing
  • US10938779B2 patent drawing

AI summary

Guided word association based domain name detection may be performed by obtaining an original domain name, constructing a feature space from a corpus of text, wherein each word appearing in the corpus is represented as a vector in the feature space, detecting whether a domain name registration exists for each combination of the original domain name and each of a plurality of seed words from the feature space, determining, for each seed word included in an existing domain name registration, a plurality of nearest neighbor candidate words, based on vector distance in the feature space, and repeating, for one or more repetitions, the detecting and the determining, wherein the plurality of nearest neighbor candidate words are utilized as the plurality of seed words.