Threat sensing system based on active domain name generation and real-time malicious domain name detection
By generating similar domain names based on Transformer and combining real-time monitoring, the rules dependence and character set adaptability of domain name counterfeiting attacks in the existing technology are solved, and efficient identification and evaluation of phishing websites are realized, adapting to different character sets and semantic changes, and improving the accuracy and real-time detection.
Patent Information
- Application Number
- CN202510626191.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing technology has problems such as strong rule dependence, low adaptability to complex attack methods, ignoring character sets and semantic information, and lacking brand context understanding when identifying domain name counterfeiting attacks, resulting in the inability to effectively identify new domain name variants and malicious domain names under different character sets.
The autoregressive model based on Transformer is used to generate similar domain names, combine historical threat perception modules and real-time threat perception modules, and generate similar domain names lists through autoregressive models, and use WHOIS query and DNS resolution to determine the domain name registration status, and monitor the newly registered domain names in combination with real-time transparency logs to perform phishing detection.
It realizes rapid identification and evaluation of potential domain name abuse, and can efficiently detect phishing websites in large-scale network environments in real time, reduce the security threat of phishing and brand abuse, adapt to different character sets and semantic changes, and improve the accuracy and real-time detection.
Smart Images

Figure CN120498763A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security detection, and in particular relates to a threat perception system based on active domain name generation and real-time detection of malicious domain names. Background Art
[0002] In the field of Internet security, attackers often use domain names that are similar to the target domain name to carry out phishing or other malicious activities.
[0003] In the field of internet security, attackers register domain names that are similar to their target domain names to carry out malicious activities such as phishing and data theft. This attack method is called domain typosquatting. Attackers select domain names that are extremely similar to legitimate websites in terms of characters. They exploit users' carelessness or spelling errors when entering the URL to lure users to malicious websites, thereby stealing their sensitive information.
[0004] The current method for generating domain name squatting is fixed, and the main methods are as follows:
[0005] Domain name distortion: Attackers create a domain name that looks similar to the target website by making minor changes to the spelling, such as replacing letters, deleting, inserting, or swapping characters. For example, attackers might register a domain name similar to g00gle.com or amaz0n.com. This method relies on user inattention or typos when entering a domain name. Typically, attackers choose domain names that are highly similar to well-known brand names in terms of characters and spelling.
[0006] Exploiting similar characters: Some attackers exploit similarities between different character sets, such as using Latin and Cyrillic characters, for example, using "а" instead of "a," to create phishing domain names. This method can evade detection by character-level similarity detection tools, making the phishing domain names more covert.
[0007] Phonetic / phonetic variations of brand names: Some attackers exploit phonetic similarity by registering domain names that sound similar to the target brand or app name. This method uses similar-sounding words to lure users into the phishing site.
[0008] Leveraging different top-level domains (TLDs): Attackers may register domains that are similar to the target domain, but change the top-level domain (e.g., .com, .net, .org, etc.). For example, attackers may register facebook.net instead of facebook.com, or use some uncommon TLD (e.g., .xyz, .club) instead.
[0009] Problems with existing methods:
[0010] 1. Strong rule dependency leads to ineffectiveness against unknown attack patterns: Rule-based domain name generation methods often rely on manually written or statically defined rules. This means that the system can only identify new domain name types when the rule base is updated with new misspellings or common variations. If new attack methods are not incorporated into the rule base, the detection system may miss related malicious domain names. Attackers simply need to constantly change their strategies and tactics, and rule-based methods are often unable to cope with this ever-evolving attack. For example, attackers may create new domain name variations that do not conform to the preset rules, and these domain names may go undetected.
[0011] 2. Low adaptability to complex attack methods: Extremely low character set and language diversity: Existing rules often assume that the target domain name is based on a standard character set (such as ASCII), ignoring the various character set variants that attackers may use (for example, Latin characters, non-Latin characters, emoticons, etc.). When attackers use different character sets (such as Cyrillic instead of Latin characters) to generate phishing domain names, rule-based methods often fail to identify these domain names.
[0012] 3. Lack of context and semantic understanding: Rule-based domain name detection methods typically focus solely on character similarity in domain names, ignoring the semantic information that may exist behind the domain names. For example, some malicious domain names may replace characters to resemble well-known brand names, but semantically, they remain unrelated domains. Without semantic understanding, rule-based methods cannot identify these situations.
[0013] 4. Ignoring brand context: Many rules only consider character-level changes but fail to consider contextual information about the target brand (such as the brand's product types, business scope, etc.). As a result, some legitimate domain names may be improperly matched with the target brand. Summary of the Invention
[0014] In response to the problems existing in the above methods, the present invention proposes a threat perception system based on active domain name generation and real-time detection of malicious domain names. It actively generates a list of similar domain names to the target domain name through an autoregressive model, which is then used for phishing page risk assessment and detection. At the same time, by real-time monitoring of newly registered domain names, phishing detection is performed on domain names with high similarity.
[0015] The present invention provides a threat perception system based on active domain name generation and real-time detection of malicious domain names, comprising a historical threat perception module, a real-time threat perception module, and a malicious detection module. The historical threat perception module is provided with a domain name generation module and a risk assessment module.
[0016] The domain name generation module uses a Transformer-based autoregressive model to build a similar domain name generation model, collects the identification information of the target enterprise and all registered and used subdomains, identifies keywords and phrases in the subdomains, and generates similar domain names of the target enterprise based on the keywords or phrases of the subdomains, the target enterprise logo, and the top-level domain name generation model input. The similar domain name generation model generates similar domain names of the target enterprise and stores them in a similar domain name list.
[0017] The risk assessment module determines whether the domain names in the list have been registered through WHOIS query and DNS resolution, and adds the registered and accessible domain names to the list to be detected.
[0018] The real-time threat perception module monitors newly registered domain names in real time and detects whether the similarity between the monitored domain names and the target enterprise domain names exceeds the set similarity threshold. If so, the monitored domain names are added to the suspicious domain name list.
[0019] The malicious detection module detects domain names in the to-be-detected list and the suspicious domain name list to identify phishing websites.
[0020] The domain name generation module generates the input of the autoregressive model for the target enterprise, and the implementation includes:
[0021] (1) Collect the target company's logo, including the company's official name, pinyin, English name and abbreviation;
[0022] (2) Collect all subdomains registered and used by the target enterprise;
[0023] (3) performing semantic analysis and splitting on the collected identifiers and subdomains, extracting keywords and phrases, performing clustering to eliminate redundancy, and then generating the input of the autoregressive model according to the set pattern;
[0024] The pattern for generating autoregressive model input is: brand phrase constraint + subdomain phrase constraint + top-level domain constraint; among them, the top-level domain constraint is to select a top-level domain from the top-level domain set, which includes .com and .cn; the brand phrase constraint is to select a phrase from the brand cluster, and the company's official name, pinyin, English name and abbreviation are all added to the brand cluster as a phrase; the subdomain phrase constraint is to select a phrase from at least one subdomain cluster set, split and cluster the subdomains, and divide them into different clusters. The keywords or phrases in each cluster are similar items.
[0025] The domain name generation module is a similar domain name generation model constructed based on the Transformer autoregressive model, including an embedding processing module, K decoder blocks, and a linear projection and output module, where K is an integer greater than 1. The embedding processing module splits the model input into a character sequence, obtains a character embedding representation through a table lookup, obtains a character position code based on the sequence, adds the character embedding representation and the character position code element-wise to form an initial representation X0, and processes X0 sequentially through K decoder blocks, gradually integrating context information. Finally, the linear projection and output module maps the output of the last decoder block to the vocabulary dimension and obtains the next character probability through a softmax activation function. The output character is fed back to the input end through an autoregressive loop, and iterative generation is performed until a termination symbol is encountered to generate a similar domain name. The similar domain name generation model generates similar domain names based on beam search, outputting all hits for keywords or fields in the initial input.
[0026] The advantages and positive effects of the present invention are:
[0027] (1) The system of the present invention can quickly identify and assess potential domain name abuse risks. In practical applications, it can be connected with more security analysis and intelligence platforms. After discovering high-risk domain names, it can take blocking measures or notify relevant parties as soon as possible, ultimately effectively reducing phishing, brand abuse, and other domain name-based security threats.
[0028] (2) Compared to the rule-based, one-by-one comparison method, the similar domain name generation model of the present invention pre-learns the underlying patterns of domain names from tens of thousands of domain names. It can generate a large number of similar domain names of corporate domain names at once, and then detect them to determine whether there are phishing websites, thus achieving efficient real-time detection in a large-scale network environment. At the same time, the similar domain name generation model is based on the autoregressive characteristics of the transformer model, enabling it to learn from the latest data in real time without the need for manual rule updates.
[0029] (3) The similar domain name generation model of the system of the present invention learns the statistical characteristics of legitimate domain names and malicious domain names based on historical domain names. When generating, the learned specific keywords can be added to generate counterfeit websites; at the same time, the similar domain name generation model considers information such as brand logos and existing subdomain characteristics to ensure that the generated domain names can maintain a high degree of similarity with the company's real domain names in terms of visual and semantics, thereby more effectively simulating potential phishing attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a block diagram of the threat perception system based on active domain name generation and real-time detection of malicious domain names of the present invention;
[0031] Figure 2This is a schematic diagram of the system of the present invention performing historical threat perception on the target enterprise domain name;
[0032] Figure 3 This is a diagram of an implementation framework for generating similar domain names according to an embodiment of the present invention;
[0033] Figure 4 This is a flow chart of threat detection performed by the malicious detection module of the present invention. DETAILED DESCRIPTION
[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0035] The threat perception system, based on proactive domain name generation and real-time detection of malicious domains, detects both historical and real-time threats to target enterprise domain names. Historical threat perception primarily utilizes an autoregressive domain name generation algorithm to proactively generate potential threat domains. Real-time threat perception utilizes transparency logs to monitor newly registered domain names in real time. Using a similarity algorithm, domains with high similarity are compared with the target domain name for phishing detection.
[0036] Specifically, if Figure 1 As shown, the threat perception system of the present invention includes a historical threat perception module, a real-time threat perception module and a malicious detection module. The historical threat perception module performs historical threat perception on the target enterprise domain name. The implementation of the historical threat perception of the present invention relies on the domain name generation algorithm. First, the autoregressive model is used to generate and simulate threat domain names similar to the target enterprise that the attacker may use. The model is pre-trained on one million benign and malicious domain names, and the statistical characteristics of brand domain names are learned. A list of domain names that may be used for phishing or malicious distribution is automatically constructed, and then the registered and accessible domain names are evaluated for phishing page risks. The overall Figure 2 The real-time threat perception module monitors newly registered domain names in real time and performs a phishing page risk assessment if the domain name has a high degree of similarity to the target domain name.
[0037] The historical threat perception module of the present invention includes a domain name generation module and a risk assessment module. To effectively assess the potential threat level faced by a target domain name, the domain name generation module first uses an autoregressive model to proactively generate a list of domain names similar to the target domain name. The risk assessment module then uses WHOIS queries and DNS resolution to determine whether the generated similar domain names have been registered. Registered and accessible domain names are then added to a pending detection list to conduct a phishing risk assessment.
[0038] The domain name generation module uses a diverse domain name generation algorithm to construct a domain name list for individual companies, as follows:
[0039] Step 11. Collecting corporate identification information: The system collects as much identification information as possible about the target company, including its official name, common abbreviation, pinyin, English name, etc. This step ensures that the generated domain name maintains a high degree of visual and semantic similarity to the company's real domain name, thereby more effectively simulating potential phishing attacks.
[0040] Step 12. Subdomain Information Collection: The system collects all subdomains registered and used by the enterprise through web crawlers and enterprise IT asset management records. For example, for Baidu, this might include subdomains such as baidu.com and hao123.com. This step includes not only public primary domains but also uncommon or internally used subdomains.
[0041] Step 13. Perform semantic analysis on the collected subdomains to extract keywords or phrases, then cluster and remove redundancy for the keywords or phrases of all identifiers and subdomains, and generate the input of the model according to the set pattern.
[0042] The embodiment of the present invention splits the subdomain name by delimiters to extract keywords or phrases, for example, extracting "hao123" from "hao123.com". The formal name, English name, pinyin, abbreviation, etc. of the enterprise are all regarded as a keyword or phrase.
[0043] In this embodiment of the present invention, all collected target enterprise logos and subdomain phrases are clustered and de-redundant, including: processing subdomain keywords or phrases: first, using a length threshold and global frequency of occurrence, remove subdomain keywords that are too long, such as those greater than 15 characters, and rarely appear. Second, the keywords extracted from the subdomains are clustered and merged using edit distance or Jaro-Winkler distance to form different clusters. The keywords within each cluster are all similar terms. Then, based on frequency of occurrence or TF-IDF (term frequency-inverse document frequency), two or three keywords are selected from each cluster to retain, thereby filtering out redundant similar words in the subdomain. For logos, the company's official name, pinyin, English, and abbreviation are all treated as phrases and grouped into a synonym cluster, namely the brand cluster.
[0044] Based on the obtained subdomain cluster set and brand cluster, the input of the autoregressive model is constructed. In the embodiment of the present invention, the disjunctive constraint DisjunctiveConstraint is used to represent each constraint. DisjunctiveConstraint means "any one of the clusters must appear", such as {baidu|bd}. In the embodiment of the present invention, a top-level domain name (TLD) set is also set, which includes .com and .cn. Each time input is made, a phrase is arbitrarily selected from the brand cluster as a brand phrase constraint, a domain name is randomly selected from the top-level domain name set as a top-level domain name constraint, and a phrase is randomly selected from at least one subdomain cluster set as a subdomain phrase constraint. Then, the input is composed together according to "brand phrase constraint + subdomain phrase constraint + TLD constraint".
[0045] For example, after clustering and removing redundancy from a target company's subdomains, distinct clusters are obtained. The retained phrases belonging to these clusters include zhidao, image, hao123, cloud, and pan. When generating input, if batch generation is performed for a single subdomain, for example, with the brand phrase constraint set to bd, the subdomain phrase constraint set to hao123, and the top-level domain constraint set to .com, then bd-hao123.com is formed as the input. If multiple subdomains are generated simultaneously, for example, with the brand phrase constraint set to bd, the subdomain phrase constraints set to zhidao and image, and the top-level domain constraint set to .com, then bd-zhidao-image.com is formed as the input.
[0046] The autoregressive model of the present invention generates similar domain names based on Constrained Beam Search (CBS). Constrained Beam Search ensures that the sequence with the highest language model probability is selected while enforcing that the generated text must contain pre-specified words or phrases. This means that the autoregressive model of the present invention must include the input fields when generating the first-kill domain name. By using this input generation method, the present invention can reduce the number of hard constraints in each CBS round, even if there are hundreds of subdomains, brand information, etc., thus avoiding search space explosion.
[0047] The autoregressive model generates similar domain names based on CBS. When traversing the path with the highest probability, the model forces all inputs to hit, and the remaining positions are freely spliced by the autoregressive model, maintaining natural hyphen and number insertion. This ensures that the model can generate registered domain names that contain target subdomains and brand information.
[0048] Step 14. Generate input to the autoregressive model based on the Transformer architecture according to the method provided in step 13, and the autoregressive model outputs similar domain names.
[0049] like Figure 3As shown, an embodiment of the present invention implements a similar domain name generation model based on an autoregressive model of a Transformer architecture. The model includes an embedding processing module, several decoder blocks, and a linear projection and output module. The embedding processing module splits the input containing subdomains and brand information into character sequences, and obtains a token embedding for each character by looking up the table. At the same time, the corresponding position code is found according to the sequence position, and the two are added element by element to form the initial representation X0. X0 then flows through multiple decoder blocks in sequence, gradually refining and integrating context information. Each decoder block contains "multi-head self-attention with lower triangular mask" and "feedforward fully connected network", and is equipped with residual connection and normalization layer. For specific implementation, please refer to the decoder in the autoregressive model based on the Transformer architecture, which will not be described in detail in the present invention. The linear projection and output module maps the output projection of the last decoder block to the vocabulary dimension and obtains the probability of the next character through the Softmax activation function. The output character is fed back to the input end through an autoregressive loop, and it is iteratively generated until the termination symbol is encountered to generate a similar domain name.
[0050] During the generation phase of this embodiment of the present invention, the system first performs "constraint pool compression" on the subdomain segments and brand aliases according to the method of step 13. Each round retains fewer than eight phrase constraints to generate input, which is then packaged into several "brand × subdomain" batches. CBS decoding is then applied to each input: Constraint completion is dynamically tracked during beam-search, forcing the candidate sequence to hit all key words. The language model then freely splices in other positions to maintain the natural "domain flavor," such as hyphens and number insertions. Finally, the model output undergoes a linear projection onto the vocabulary dimension and a Softmax operation to determine the next character probability. The sampled characters are then fed back to the input through an autoregressive loop, and the process continues until a termination symbol is encountered, generating a complete similar domain name. After all batches are processed, regularization filtering, WHOIS registrability verification, and similarity ranking are performed uniformly to produce a comprehensive and highly realistic list of similar domain names.
[0051] When training the autoregressive model, the present invention collects one million existing domain names, including both legitimate and phishing webpage domain names. The model uses these domain names to learn the statistical characteristics of domain names. For example, attackers often add deceptive phrases such as "login," "community," "support," and "account" to phishing domain names. These phrases are often closely related to brand-related login pages, social features, or user accounts. These phrases are learned during model training and added to the vocabulary. Through training, the model learns how to generate phishing websites that lure users by adding specific keywords such as "login" and "community."
[0052] The Transformer-based autoregressive model of the present invention can learn contextual information about the domain name structure from existing domain name data. By training on a large number of legitimate domain names, the model can understand the common patterns, spelling rules and brand characteristics of domain names. The Transformer model can automatically learn and identify new domain name variants from data. The Transformer-architecture autoregressive model can accurately capture the multi-level features of similarity: learning the deep dependencies between characters, words and their context in the domain name. This enables the model to more accurately identify subtle differences between legitimate domain names and counterfeit domain names, rather than just simple matching at the character level. In this way, the model can reduce missed detections, especially when dealing with different languages or non-standard character sets.
[0053] Transformer can handle input in different character sets and capture the underlying similarities between characters. It can not only identify spelling errors, but also understand similarities between different languages and character sets (such as using Russian letters instead of Latin letters). This makes it possible to identify domain name variants composed of different character sets and languages.
[0054] In addition to character similarity, the Transformer model can also capture the semantic relevance of domain names through semantic embedding. For example, amazon.com may have subtle character differences from amazon.com, but semantically, they still refer to the same brand. Transformer-based autoregressive models can be trained to understand brand semantics and better identify this type of counterfeiting. Semantic learning of brand features: By incorporating brand-related background information (such as company name, product name, trademark, etc.) into the model training process, the Transformer can better understand the uniqueness of a specific brand. For example, goog1e.com is not only similar to google.com in terms of characters, it is also semantically a variant of google.com.
[0055] When analyzing cybersquatting threats, a specific domain name structure—a second-level domain structure similar to baidu.com.example.com—was found to be unusually frequently used in malicious webpages. This domain name pattern is often used by cyber attackers to mislead users into believing they are accessing a legitimate subdomain. This type of second-level domain structure is difficult to obtain through simple domain name construction methods because the owner of the top-level domain example.com has not registered all possible subdomain combinations. Therefore, malicious attackers can register seemingly legitimate second-level domains under it. While this is technically legal, it can pose a security threat to users.
[0056] Due to the unique structure and difficulty of constructing this type of domain name, simple historical data analysis and pre-generated domain name methods cannot effectively cover all potential risk domain names. Therefore, this system integrates a detection module for similar domain names into the real-time threat perception module.
[0057] After obtaining a list of similar domains to the target domain, the risk assessment module uses DNS resolution and WHOIS queries to determine whether the generated similar domains have been registered. It then conducts a phishing page risk assessment for registered and accessible domains. Combining URL characteristics, page text similarity, certificate information, suspicious redirection behavior, and security vendor intelligence databases, the module provides a comprehensive score or judgment on the detected pages, automatically identifying potential phishing risks. The risk assessment module also prioritizes the risks of registered and accessible domains. It generates a risk grading report based on registration status, maliciousness of web content, and proximity to the target domain / website, helping users prioritize high-risk domains to mitigate potential harm. Regulators or businesses can also use the risk grading report to investigate the attacker organizations behind malicious domains based on similarities between registration information and server IP addresses, improving their awareness of ongoing attack activity.
[0058] The real-time threat awareness module of an embodiment of the present invention utilizes transparency logs to monitor newly registered domain names in real time, providing real-time threat awareness for target enterprise domain names. Transparency logs provide a public, auditable system that records all issued SSL / TLS certificates, including those used for HTTPS websites. By analyzing these logs, the system can capture newly registered domain names that have applied for certificates in real time, focusing particularly on new domain names that are similar to the target enterprise domain name. SSL stands for Secure Sockets Layer, and TLS stands for Transport Layer Security.
[0059] The system's real-time threat awareness module utilizes CertStream API to monitor reports to the Certificate Transparency Log (CTL), enabling near-real-time identification of suspicious TLS certificate issuances, often associated with phishing campaigns. Once a new domain with a high degree of similarity to a known enterprise domain is detected obtaining a TLS certificate, this "suspicious" issuance activity is identified as having a domain score exceeding a specific threshold based on a profile. CertStream API is a real-time Certificate Transparency Log update stream that provides real-time updates from the Certificate Transparency Log network.
[0060] This embodiment of the present invention uses the Jaro-Winkler distance, a similarity detection algorithm, to compare the similarity between new domain names and target enterprise domain names. The Jaro-Winkler distance is specifically optimized for comparing the similarity of short text strings and is more robust to spelling differences and common typos. By integrating similarity algorithms into real-time data stream analysis, the present invention can effectively screen and flag suspicious domain names that have only minor variations from the official enterprise domain name or contain common typos. This mechanism enables enterprises to quickly identify potentially malicious domain names as soon as TLS certificates are issued.
[0061] In addition, the solution proposed by historical threat perception for detecting malicious domain names constructed based on the target enterprise's second-level domain name and the real-time threat perception module includes the following steps:
[0062] Step 21. Domain Segmentation: Process the captured full domain name Y, using dots (.) as separators. For example, if the captured domain name Y is baidu.com.example.com, segment Y into the word array ['baidu', 'com', 'example', 'com']. This step is to extract the individual components from the composite domain name for further analysis and processing.
[0063] Step 22. Recombination: Recombine each word in the resulting word array after segmentation. This step primarily simulates domain name variation strategies that attackers might employ, such as recombining the original subdomain and top-level domain to form a new domain name. For example, from the array ['baidu', 'com', 'example', 'com'] above, multiple possible domain names can be formed, such as baidu.com, example.com, com.example.com, and so on.
[0064] Step 23. Domain Matching: Each newly reorganized domain is matched against the known official domain of the target company. This step aims to identify domains that appear to be official but are actually potentially malicious. For example, if the company's official domain is baidu.com, and the combination above yields baidu.com, the two would match, indicating that the domain may be a malicious domain attempting to mimic the official website.
[0065] Step 24. Suspect Domain Name Classification: If the combined domain name is found to be identical or extremely similar to the target company's domain name, it will be classified as highly suspicious and added to the list of suspicious domain names. These domain names will undergo further security review and verification to determine their authenticity and security.
[0066] The historical threat perception module of the system of the present invention outputs a list of domain names to be detected, the real-time threat perception module outputs a list of suspicious domain names, and the malicious detection module executes a phishing threat detection method to perform a phishing page risk assessment on the domain names in the two lists. In order to more finely describe the risks that these domain names may bring, the malicious detection module uses a fuzzy discrimination method to divide the threat level of similar domain names into four levels: "benign", "partially suspicious", "suspicious" and "very suspicious". The whole process is as follows Figure 4 As shown, the domain name to be detected is first redirected. If it redirects to an official company page, it can be determined that the domain name is likely an official cybersquatting operation and is immediately deemed benign, removing it from the list of pending domains. Next, a list of domain names generated and retrieved in real time from major security vendors or community-maintained blacklists, such as Google Safe Browsing, PhishTank, VirusTotal, and Openphish, is checked to see if they are already on the blacklist of phishing websites. If the domain name is already included and marked as a phishing or malicious website, it can be quickly identified and blocked, and moved to the list of suspicious domains. For domains that are not on the blacklist and cannot be properly redirected, suspicious features in the URL are used to assess potential phishing risks. A classification model pre-trained using public phishing data is used to determine URL features. The domain names are analyzed based on multiple dimensions, such as length, spelling, keywords, and registration information, to determine their similarity to known phishing domains and output a corresponding risk rating. For example, abnormal URL length, excessive subdomains or suspicious parameters, and keywords such as "login," "verify," and "secure" that do not match the target brand name are often obfuscated configurations made by phishing attackers to imitate legitimate websites. In addition, the detection method for malicious domain names constructed using the second-level domain name mentioned above is used to determine domain name anomalies, and then the abnormal domain name is moved to the list of suspicious domain names. At the same time, the digital certificate information used by the suspicious website is checked, including the issuing agency of the certificate, the validity period, the domain name issued to, etc. If the certificate does not match the actual domain name, the domain name is moved to the list of suspicious domain names. For domain names that have not been marked as phishing or malicious websites, and have also passed the detection of the second-level domain name word segmentation and reorganization by the real-time threat perception module, as well as the detection of URL features, the domain name is considered to be a benign domain name and is removed from the list to be detected.
[0067] To further reduce false positives and identify phishing pages, the Malicious Detection Module uses the Web Content Detection Module to detect domains on the suspicious domain list. The system automatically takes screenshots of suspicious websites and uses computer vision and optical character recognition (OCR) to assess their page structure and image and text content for similarity to legitimate websites. The following details the Web Content Detection Module's functions and workflow, along with a case study to illustrate its role in identifying phishing websites.
[0068] The web content detection module is primarily used to identify phishing features in suspicious website pages and determine whether they are impersonating trusted sites and conducting attacks. Specifically, this module focuses on the following aspects:
[0069] Counterfeit Logo Check: Detects whether suspicious pages contain the target website's logo or brand identity. If the target website's logo or other logo is found in web content or screenshots from a suspicious domain, the website may be impersonating the target site and is suspected of phishing.
[0070] Sensitive Information Requests: Checks whether the page requires users to enter credentials, such as usernames, passwords, and bank account numbers. If the webpage contains login forms or password input boxes, be vigilant, as phishing websites often trick users into entering sensitive information.
[0071] Suspicious file downloads: Check whether the page attempts to direct users to download unknown software or compressed files. Phishing websites often trick users into downloading malware by disguising themselves as updates or security patches. Therefore, if you detect a download link and the file's source is suspicious, pay close attention.
[0072] Through these assessments, the system can initially determine whether a suspicious website is a phishing threat. For example, if a fake bank website's homepage displays the bank's official logo and requires login credentials, or offers to download "security upgrade" software, these are obvious red flags.
[0073] Specifically, the webpage content detection module of the embodiment of the present invention performs the following series of steps on each suspicious domain name to comprehensively analyze whether its webpage content contains phishing characteristics. The entire process is as follows:
[0074] Data Acquisition: First, the system crawls the HTML source code of the suspicious domain's website homepage and generates a screenshot of the page. The HTML source code provides text and structural information about the page, while the screenshot facilitates subsequent image analysis, such as logo recognition.
[0075] Phishing signature identification: Based on the extracted page elements and screenshots, the system further checks for phishing signatures:
[0076] Logo matching: We analyze page screenshots using image recognition technology to determine if they contain the target website's logo or name. If a well-known brand / website logo is found on an unrelated domain, this is generally unnatural and suggests a website imitating the brand's appearance.
[0077] Sensitive information input: Check whether there are login forms or similar components in the page DOM elements that require users to enter credentials, such as account and password input boxes, submit buttons, etc.
[0078] Download link scanning: Scan links and button text on the page to see if there are any prompts to download software or attachments. For example, buttons or links with the words "Download Update" or "Security Patch." If so, extract the file information (file name, type) and determine whether the file source is trustworthy.
[0079] Hidden link and jump detection: For some more subtle phishing techniques, the system uses a headless browser to simulate user operations to detect multiple jumps. Some phishing websites may not directly present malicious content on the homepage, but hide links to steal information or download malicious files in buttons or page interactions. After triggering a click, the real destination is exposed through multiple URL jumps. For example, when a user clicks a seemingly normal button, the website may redirect several times and finally jump to a page that requires a password or directly pops up a download dialog box. In response to this, the detection module will automatically simulate clicking on suspicious buttons or links on the page and track its jump chain. If it is found that the final jump is to an address that does not match normal expectations, for example, if it should go to the official website but jumps to another suspicious domain name or download site, it can be determined that the website has malicious redirection behavior and is phishing in nature.
[0080] Result Verification and Recording: Based on the above inspection results, the system determines whether the suspicious domain name carries a phishing threat. If multiple key detection criteria indicate a problem, such as a counterfeit logo, a password input field, or unusual redirects, the domain name is flagged as a suspected phishing website. The system then saves the full HTML source code and screenshots of the suspicious page as evidence for subsequent manual review and further analysis by security personnel. Detection results are also recorded for all healthy, unaffected domain names, completing the final content scanning process. Researchers can use this information to quickly identify truly threatening phishing sites and take timely blocking measures or alert relevant authorities. This maximizes the automated processing capabilities of fuzzy classification and machine learning models while also leveraging manual review to prevent missed or false positives during high-risk periods, effectively improving overall detection accuracy and stability. This approach enhances the system's automated detection capabilities while maintaining interpretability and accuracy for complex scenarios, thereby maximizing protection for businesses and users against phishing attacks.
[0081] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0082] Except for the technical features described in the specification, all other technical features are known to those skilled in the art. The present invention omits descriptions of well-known components and well-known technologies to avoid redundancy and unnecessary limitation of the present invention. The implementation methods described in the above embodiments do not represent all implementation methods consistent with the present application. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the present invention.
Claims
1. A threat perception system based on active domain name generation and real-time detection of malicious domain names, characterized by: Includes historical threat perception module, real-time threat perception module and malicious detection module; A domain name generation module and a risk assessment module are set up in the historical threat perception module; The domain name generation module uses a Transformer-based autoregressive model to build a similar domain name generation model. It first collects the target company's identification information and all registered and used subdomains, identifies keywords and phrases in the subdomains, and then uses the subdomain keywords or phrases, the target company's identification, and the top-level domain name generation model as input to generate similar domain names for the target company and store them in a similar domain name list. The risk assessment module determines whether the domain names in the list have been registered through WHOIS query and DNS resolution, and adds the registered and accessible domain names to the list for detection; The real-time threat perception module monitors newly registered domain names in real time and detects whether the similarity between the monitored domain names and the target enterprise domain names exceeds the set similarity threshold. If so, the monitored domain names are added to the list of suspicious domain names. The malicious detection module detects domain names in the to-be-detected list and the suspicious domain name list, identifies phishing websites, and performs page risk assessment.
2. The system according to claim 1, wherein: The domain name generation module generates the input of the autoregressive model for the target enterprise, and the implementation includes: (11) Collect the target company's logo, including its official name, pinyin, English name, and abbreviation; (12) Collect all subdomains registered and used by the target enterprise; (13) performing semantic analysis and splitting on the collected identifiers and subdomains, extracting keywords and phrases, performing clustering to eliminate redundancy, and then generating the input of the autoregressive model according to the set pattern; The pattern for generating autoregressive model input is: brand phrase constraint + subdomain phrase constraint + top-level domain constraint; among them, the top-level domain constraint is to select a top-level domain from the top-level domain set, which includes .com and .cn; the brand phrase constraint is to select a phrase from the brand cluster, and the company's official name, pinyin, English name and abbreviation are all added to the brand cluster as a phrase; the subdomain phrase constraint is to select a phrase from at least one subdomain cluster set, split and cluster the subdomains, and divide them into different clusters. The keywords or phrases in each cluster are similar items.
3. The system according to claim 2, characterized in that The domain name generation module clusters and removes redundancy from keywords and phrases extracted from all subdomains registered and used by the target enterprise, including: first, using a length threshold and global occurrence frequency to eliminate keywords and phrases that do not meet the threshold length and are less than the global occurrence frequency; second, clustering the remaining keywords and phrases using edit distance or Jaro-Winkler distance to form different clusters, wherein the keywords in each cluster are similar terms, and keywords with high occurrence frequency or TF-IDF are retained for each cluster, where TF-IDF stands for term frequency minus inverse document frequency.
4. The system according to claim 1 or 2, characterized in that The domain name generation module is a similar domain name generation model constructed based on the Transformer autoregressive model, including an embedding processing module, N decoder blocks and a linear projection and output module, where N is an integer greater than 1; the embedding processing module splits the model input into a character sequence, obtains a character embedding representation through a table lookup, obtains a character position code based on the sequence, adds the character embedding representation and the character position code element by element to form an initial representation X0, processes X0 in sequence through K decoder blocks, gradually integrates context information, and finally, the linear projection and output module maps the output projection of the last decoder block to the vocabulary dimension and obtains the next character probability through a Softmax activation function. The output character is fed back to the input end through an autoregressive loop, and iterative generation is performed until a termination symbol is encountered to generate a similar domain name; the similar domain name generation model generates similar domain names based on beam search, and outputs all hit keywords or fields in the initial input.
5. The system according to claim 1 or 2, characterized in that The domain name generation module pre-trains a similar domain name generation model, collects existing domain names, including legitimate domain names and malicious website domain names, learns the statistical features of domain names, and adds corresponding feature phrases to the vocabulary.
6. The system according to claim 1, wherein: The real-time threat perception module also detects the target enterprise's second-level domain name to determine whether it is a malicious domain name. The detection includes: (21) Split the captured full domain name Y using ".." as a delimiter to obtain a word array; (22) combining the words in the word array to form a new domain name; (23) Each newly reorganized domain name is matched with the official domain name of the target enterprise. If the domain names match or are highly similar, domain name Y is considered to be a malicious domain name that imitates the official domain name of the target enterprise and is added to the list of suspicious domain names.
7. The system according to claim 1 or 6, characterized in that The malicious detection module performs detection, including: (31) Redirect the domain names in the list to be detected. If the redirection leads to the official page of the enterprise, the domain name is determined to be an official preemptive registration and the domain name is removed from the list to be detected. For the domain names that cannot be redirected, check whether they are in the blacklist. If the domain name has been marked as a phishing or malicious website, identify the domain name as a malicious website and block it, and add the domain name to the list of suspicious domain names. If the domain name has not been marked as a phishing or malicious website, perform word segmentation and word reorganization on the domain name, calculate the matching degree between the reorganized domain name and the target enterprise domain name, if the matching degree is high, add the domain name to the list of suspicious domain names, otherwise, continue to detect the URL features of the domain name to determine whether it is abnormal. If it is abnormal, add it to the list of suspicious domain names, otherwise, consider the domain name to be a benign domain name and remove it from the list to be detected. (32) Performing web content detection on the domain names in the suspicious domain name list, including: crawling the HTML source code of the website homepage of the domain name and generating a screenshot of the page; parsing the page elements and extracting the page components; detecting whether the following phishing features exist from the page screenshots and page components: The page contains the target company's website logo or branding; There are components in the page components that require users to enter credentials; There are links or buttons on the page that prompt you to download software or files; Simulate user operations to detect multiple jumps. If the jump to an address that does not match the expected address; Domain name threat levels are classified based on the presence of phishing characteristics.
Citation Information
Patent Citations
Adaptive security threat analysis method and system for counterfeit domain name
CN110855716A
Adversarial domain name generation model with high detection resistance
CN113709152A
Malicious domain name intelligent detection method, system and device and storage medium
CN117579335A
Phishing security verification method, system and device and storage medium
CN119094186A
Malicious homoglyphic domain name detection and associated cyber security applications
US20230083949A1
Cited By
Cross-message HTTP domain name matching method and system on network shunting equipment
CN121887772A