Malware Domain Identification via Lexical Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for identifying potential malware domain names are limited, time-consuming, and resource-intensive, often requiring extensive analysis of traffic patterns and file downloads, which is not efficient enough to quickly detect and respond to malicious sites that may have short lifespans.

Innovation Solution

A system and method that employs lexical and linguistic analysis, including phonetic analysis and N-gram statistical analysis, to identify potential malware domain names by analyzing features such as hyphen count, keyword count, character length, top-level domain origination, and suspicious alphanumeric sequences, and uses machine learning to build a classifier for fine-tuned identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional computer security techniques are used to detect potential malware domain names, then detection accuracy may be maintained, but the process becomes time-consuming and resource-intensive

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the domain name analysis into multiple independent lexical features (hyphen count, keyword count, character length, word-to-digit ratio, top level domain analysis, suspicious alphanumeric sequences). Each feature is analyzed separately and contributes to the overall malware probability score, enabling parallel processing and reducing detection time while maintaining accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the detection approach by changing parameters from traditional security analysis to lexical and linguistic feature analysis. By evaluating domain names based on configurable thresholds for each lexical feature and combining them into a probability score, the system achieves faster detection with reduced resource consumption while maintaining reliable identification of malware domains

Inventive Principle:
Principle #35Parameter changes

2Reliability

If conventional computer security techniques are used to detect potential malware domain names, then comprehensive analysis may be performed, but computational resources are excessively consumed

Engineering Contradiction:
Improvedetection reliabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent divides the detection process into discrete lexical feature evaluations (hyphen count, keyword count, character length, word-to-digit ratio, top level domain analysis, suspicious sequences). Each feature can be independently computed with minimal computational overhead, and results are combined through a probability scoring mechanism, significantly reducing overall resource consumption compared to comprehensive conventional analysis

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a configurable threshold system where detection can be performed with partial analysis. By allowing administrators to set thresholds for individual lexical features and combine them into an overall probability score, the system can achieve reliable detection with less computational effort than exhaustive conventional methods, enabling resource-constrained environments to effectively detect malware domains

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If traditional malware detection methods are used, then thorough analysis of traffic patterns and file downloads may be conducted, but the response speed to emerging threats is insufficient

Engineering Contradiction:
Improvedetection thoroughnessVSAvoidresponse speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary lexical and linguistic analysis on domain names before actual malware execution or traffic generation occurs. By evaluating lexical features (hyphen count, keyword count, character length, suspicious sequences) and computing malware probability scores in advance, the system identifies potential threats at the domain registration or first-access stage, enabling rapid response to emerging threats without waiting for traffic pattern analysis or file download detection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments domain name analysis into independent lexical features that can be evaluated quickly and in parallel. This segmentation enables the system to perform thorough analysis of multiple dimensions (lexical structure, linguistic patterns, phonetic properties, N-gram statistics) simultaneously, achieving both detection thoroughness and high response speed by avoiding sequential processing of traditional traffic analysis methods

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8631498B1Techniques for identifying potential malware domain names
Publication Date: 2014.01.14 CA TECH INC
  • US8631498B1 patent drawing
  • US8631498B1 patent drawing
  • US8631498B1 patent drawing

AI summary

Techniques for identifying potential malware domain names are disclosed. In one particular exemplary embodiment, the techniques may be realized as a system for identifying potential malware domain names. The system may comprise one or more processors communicatively coupled to a network. The one or more processors may be configured to receive a request for network data, where the request for network data may comprise a domain name. The one or more processors may also be configured to apply a lexical and linguistic analysis to the domain name. The one or more processors may also be configured to identify whether the domain name is a potential malware domain name based on the lexical and linguistic analysis.