Malignant Domain Hosting Type Classification System and Method
Patent Information
- Application Number
- JP2022562116
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-13
- Filing Date
- 2021-04-13
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-04-13
AI Technical Summary
Existing methods for detecting malicious website hosting types are inadequate, as attackers increasingly use infrastructure they do not own, evading detection by current rating systems, and conventional approaches based on domain popularity and longevity are not accurate.
A software-based classifier using machine learning models to distinguish between public and private malicious URL hosting apex domains, enabling precise identification of compromised or attacker-owned domains for tailored mitigation actions.
The classifier achieves high accuracy in distinguishing between public and private apex domains, with 97.2% accuracy, 97.7% precision, and 95.6% recall for public/private classification, and 96.4% accuracy, 99.1% precision, and 92.6% recall for compromised/attacker-owned domains, facilitating effective security measures.
Smart Images

Figure 00000015_0000 
Figure 00000016_0000 
Figure 00000017_0000
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the priority and benefit of U.S. Provisional Application No. 63 / 009,151, filed on April 13, 2020, which is hereby incorporated by reference in its entirety.
[0002] This application generally relates to domain classification. More specifically, this application provides a software - based classifier built on a machine - learning model that differentiates between public and private malicious URL - hosting apex domains.
Background Art
[0003] Every week, millions of users are deceived into accessing malicious websites, from which malicious actors initiate various attacks including phishing, spam, and malware. Despite recent advancements in technologies and tools for detecting malicious websites, many malicious websites are not detected or are detected only well after damage has occurred. One important reason for this negative trend is that attackers increasingly host their websites on infrastructure they do not own instead of registering their own domains, thus avoiding detection by current evaluation systems. The detection of malicious websites, especially phishing and malware websites registered by attackers, has been widely studied, but little has been done to analyze how these malicious websites are hosted. Knowing early which hosting type a malicious URL comes from helps security operators take appropriate actions.
[0004] Appropriate mitigation actions against a malicious website can vary significantly depending on how the site is hosted. If it is hosted under a private apex domain, and all its subdomains and pages are under the direct control of the apex domain owner, the malicious website may be blocked at the apex domain level. If it is hosted under a public apex domain (e.g., a web hosting service provider), blocking at the subdomain level is more appropriate. Furthermore, in the former case, the private apex domain may be legitimate but vulnerable, or it may be generated by an attacker, which also ensures different mitigation behavior. An apex domain owned by an attacker may be permanently blocked, while a vulnerable domain may only be blocked temporarily.
[0005] The hosting type of malicious URLs is traditionally detected manually through domain reputation systems and blacklists. For example, the Anti-Phishing Working Group can identify them. While lists of public apex domains exist from multiple sources, they are not complete even when combined. Furthermore, these lists are often outdated due to the highly dynamic nature of public web hosting and cloud businesses. Therefore, given a malicious URL, simply examining such lists is not sufficient to determine whether it is hosted on a public apex domain.
[0006] Furthermore, traditional methods for classifying malicious websites as those hosted on compromised or attacker-owned apex domains are not as effective as desired. One traditional method is to consider domain popularity, such as Alexa ranking. Generally, compromised domains are understood to have some lingering reputation and longevity, while attacker-owned domains have low reputation and short lifespan. However, the inventors' analysis of malicious websites in VT shows that such observations are not always true. While there are compromised domains with high Alexa rankings and long lifespans (e.g., linode.com, cleverreach.com), the inventors have observed that there are many other domains that are likely to be compromised by attackers to launch attacks, having low or no Alexa rankings (e.g., gemtown88.com, vanemery.com), being abandoned, or being barely maintained. Moreover, newly created benign domains do not possess any of the above characteristics and may be mislabeled as attacker-owned when actually compromised. On the one hand, it is true that many domains created by attackers have very low Alexa rankings and are short-lived, but sophisticated attackers today increasingly utilize long-lived domains by creating and parking them for a period of time (e.g., crackarea.com, estilo.com.ec) to evade detection. Furthermore, attackers can artificially increase the popularity of their domains, at least in the short term, without requiring significant resources. Therefore, relying solely on popularity and / or lifespan does not result in accurate labeling of these malicious domains.
[0007] Therefore, to improve security, it is necessary to detect the hosting type of malicious URLs more quickly and efficiently. [Overview of the project] [Problems that the invention aims to solve]
[0008] This application provides a software-based classifier built on a machine learning model that distinguishes between two types of malicious URL hosting apex domains: public and private. This classification helps security professionals specify which domain levels to block, i.e., the entire apex domain in the case of a private apex, or a specific subdomain in the case of a public apex. In at least some aspects, the classifier is built on a machine learning model that distinguishes hosting domains owned by attackers from compromised hosting domains. This distinction is important to help security operators take appropriate mitigation measures. For example, domains owned by attackers may be permanently blocked, while compromised domains may be temporarily blocked. [Means for solving the problem]
[0009] In a first aspect of the disclosure of this application, which may be combined with any other aspect unless otherwise specified, in light of the technical features described herein, the system includes a display and memory that communicates with a processor. The processor may be configured to identify a malicious domain from a set of received domains; use a model to determine whether the identified malicious domain is a public domain or a private domain; if the identified malicious domain is a private domain, use a model to determine whether the private domain is a compromised domain or an attacker-owned domain; display the determined malicious domain hosting type on the display, such that the determined malicious hosting type is a public domain, a compromised private domain, or an attacker-owned private domain.
[0010] In a second aspect of the disclosure of this application, which can be combined with any other aspect unless otherwise specified, the method includes the step of identifying a malicious domain from a set of received domains. A model may be used to determine whether the identified malicious domain is a public domain or a private domain. If the identified malicious domain is a private domain, a model may be used to determine whether the private domain is a compromised domain or an attacker-owned domain. The determined malicious domain hosting type may be displayed. In this aspect, the determined malicious hosting type is a public domain, a compromised private domain, or an attacker-owned private domain.
[0011] Additional features and advantages of the disclosed methods and apparatus will be described and made apparent from the following detailed description and drawings. The features and advantages described herein are not exhaustive, and many additional features and advantages will be apparent to those skilled in the art in consideration of the drawings and description. Furthermore, it should be noted that the language used herein has been chosen primarily for readability and explanatory purposes and does not limit the scope of the subject matter of the invention. [Brief explanation of the drawing]
[0012] [Figure 1] This is a block diagram of an exemplary system for classifying malicious domain hosting types according to one aspect of the present disclosure.
[0013] [Figure 2] This is a flowchart illustrating an exemplary method for classifying malicious domain hosting types according to one aspect of the present disclosure.
[0014] [Figure 3] This graph compares VT URL intelligence with SA and GSB.
[0015] [Figure 4]A graph showing that the AUC of the ROC curve is 96% for GT1.
[0016] [Figure 5] A graph showing that the AUC of the ROC curve is 99% for GT2.
[0017] [Figure 6] A table showing various features of five feature groups that can be considered by a private domain classifier according to one aspect of the present disclosure.
[0018] [Figure 7] A diagram showing a correlation matrix of class labels, domain duration, scanner count, and Alexa rank according to one aspect of the present disclosure.
[0019] [Figure 8] A graph showing the ROC curve and feature importance in an example where the private domain classifier is a random forest classifier according to one aspect of the present disclosure.
[0020] [Figure 9] A graph showing the ROC curve for an example where the private domain classifier is a random forest classifier according to one aspect of the present disclosure.
[0021] [Figure 10] A graph showing the CDF of the number of FQDNs per apex during the period for putative benign and malignant domains.
[0022] [Figure 11A] A graph showing the number of FQDNs per apex for apex domains in two categories, public and private.
[0023] [Figure 11B]This graph shows the average Alexa ranking distribution for public and private apex domains.
[0024] [Figure 11C] This graph shows the domain lifetime distribution of public and private apex domains.
[0025] [Figure 12A] This graph shows the #FQDNs for each apex of compromised domains and domains owned by attackers.
[0026] [Figure 12B] This graph shows the average Alexa rank distribution of compromised apex domains and apex domains owned by attackers.
[0027] [Figure 12C] This graph shows the domain lifetime distribution of compromised apex domains and apex domains owned by attackers.
[0028] [Figure 13] This figure shows the feature correlation matrix of features used in a public domain classifier according to one aspect of this disclosure.
[0029] [Figure 14A] This graph shows the feature importance of a random forest-based public domain classifier for the GT1 dataset.
[0030] [Figure 14B] This graph shows the feature importance of a random forest-based public domain classifier for the GT2 dataset.
[0031] [Figure 15A] This graph shows the t-SNE of a random forest-based public domain classifier for the GT1 dataset.
[0032] [Figure 15B] This graph shows the t-SNE of a random forest-based public domain classifier for the GT2 dataset.
[0033] [Figure 16A] This graph shows the precision-recall ratio of a random forest-based public domain classifier for the GT1 dataset.
[0034] [Figure 16B] This graph shows the precision-recall ratio of a random forest-based public domain classifier for the GT2 dataset.
[0035] [Figure 17A] This graph shows the feature importance of 140 random forest-based private domain classifiers for the GT1 dataset.
[0036] [Figure 17B] This graph shows the feature importance of 140 random forest-based private domain classifiers for the GT2 dataset.
[0037] [Figure 18A] This graph shows the t-SNE of 140 random forest-based private domain classifiers for the GT1 dataset.
[0038] [Figure 18B] This graph shows the t-SNE of 140 random forest-based private domain classifiers on the GT2 dataset.
[0039] [Figure 19A] This graph shows the precision-recall ratio of 140 random forest-based private domain classifiers for the GT1 dataset.
[0040] [Figure 19B] This graph shows the precision-recall ratio of 140 random forest-based private domain classifiers for the GT2 dataset. [Modes for carrying out the invention]
[0041] This application provides a novel and innovative malicious domain hosting type classification system and method. Knowing early on which hosting type a malicious URL originates from helps security operators take appropriate action. The distinction between public and private apex domains has a significant impact on the inference and prediction of malicious domains, especially when it relies on the association of subdomains belonging to the same apex domain. Furthermore, when a malicious website is detected, the action taken against the hosting apex domains will differ depending on whether they are public or private. The provided classification system identifies public and private apex domains based on the key observation that subdomains of private apex domains exhibit more consistent behavior and characteristics compared to subdomains of public apex domains.
[0042] In at least some aspects, a classification system can determine whether a hosting domain marked as malicious is compromised or owned by an attacker. For example, if a provided system identifies a malicious website as being hosted on a private apex domain, the system may further classify the apex domain based on its owner. A malicious website may be created by an attacker on their own registered domain (e.g., getbinance.org) or a compromised benign domain (e.g., questionpro.com). In the latter case, the legitimate domain used for the malicious activity is the sacrificial domain. The takedown strategy and who to contact will differ depending on the type of apex domain. Early detection of compromised domains helps owners identify the root cause, take corrective action, and control the assessed damage, although a Security Operations Center (SOC) team may temporarily block such sacrificial domains to protect users. Domains owned by attackers, on the other hand, require entirely different actions. They are typically blacklisted first to mitigate immediate damage. If these are involved in cybersquatting, they can be further shut down through third-party takedown services, domain registration removal, or transfer of ownership.
[0043] The inventors found that the provided classifier achieves an accuracy of 97.2% with a precision of 97.7% and a recall of 95.6% in identifying public and private apex domains. Furthermore, the inventors found that the provided classifier achieves an accuracy of 96.4% with a precision of 99.1% and a recall of 92.6% in determining whether a malicious hosting domain is compromised or owned by an attacker.
[0044] As used herein, an apex domain is defined as a public apex domain if no subdomains (e.g., alice.000webhostapp.com) or pages (e.g., sites.google.com / alice) are created and are not under the control of the owner of the apex domain (e.g., 000webhostapp.com). As used herein, an apex domain is defined as a private apex domain if its subdomains (e.g., careers.nsa.gov) are created and managed by the owner of the apex domain (e.g., nsa.gov).
[0045] Figure 1 shows a block diagram of an exemplary system 100. In other examples, the components of system 100 may be combined, rearranged, removed, or served on separate devices or servers. The exemplary system 100 may include an exemplary classification system 110 for classifying the hosting type of malicious domains. For example, the exemplary classification system 110 may automatically label a malicious website (i.e., a URL) as an attacker-owned public domain (e.g., 000webhostapp.com), a compromised (private) domain (questionpro.com), or an attacker-owned (private) domain (getbinanace.org). In various embodiments, the classification system 110 may communicate with at least one evaluation system 160 via a network 150. The network 150 may include, for example, the Internet or any other data network, including any suitable wide area network or local area network, but is not limited to.
[0046] The rating system 160 may be any suitable blacklist or rating system that provides ratings of websites or URLs (e.g., whether they are malicious). In some embodiments, the rating system 160 is the VirusTotal (VT) system. VirusTotal (VT) is a known rating service that provides aggregated intelligence on any URL by examining third-party antivirus tools and URL / domain rating services. Each of these tools is referred to herein as a scanner. VT aggregates query results per second and makes them available as a feed to subscribers. In other examples, the rating system 160 may be generated / maintained by Google Safe Browsing (GSB), Phishtank, Anti-Phishing Working Group (APWG), McAfee Site Advisor (SA), or other suitable blacklist or rating system. In some embodiments, the classification system 110 may communicate with multiple blacklists or rating systems.
[0047] In various embodiments, the classification system 110 may include a processor that communicates with memory 114. The processor may be a CPU 112, an ASIC, or any other similar device. In some examples, the classification system 110 may include a display 116. The display 116 may be any suitable display for displaying information. In various embodiments, the classification system 110 may include a malicious domain identifier 120. The malicious domain identifier 120 may identify a malicious domain based on information received from the evaluation system 160. In various embodiments, the classification system 110 may include a public domain classifier 130. The public domain classifier 130 may determine whether a malicious domain is a public apex domain or a private apex domain. In various embodiments, the classification system 110 may include a private domain classifier 140. The private domain classifier 140 may determine whether a private apex domain is compromised or owned by an attacker. Each of the malicious domain identifiers 120, the public domain classifier 130, and the private domain classifier 140 may be implemented by software executed by the CPU 112. In other examples, the components of the classification system 110 may be combined, rearranged, removed, or provided on separate devices or servers.
[0048] In some examples, the public domain classifier 130 may be a random forest classifier. In other examples, the public domain classifier may be a support vector classification (SV), extra tree (ET), logistic regression (LR), decision tree (DT), gradient boosting (GB), Ada boosting (AB), or K-neighbors (KN) classifier. In some examples, the private domain classifier 140 may be a random forest classifier or an extra tree (ET) classifier. In other examples, the public domain classifier may be a support vector classification (SV), logistic regression (LR), decision tree (DT), gradient boosting (GB), Ada boosting (AB), or K-neighbors (KN) classifier.
[0049] Figure 2 shows a flowchart of an exemplary method 200 for classifying the hosting type of a malicious domain. While exemplary method 200 is described with reference to the flowchart shown in Figure 2, it will be understood that many other methods may be used to perform operations related to method 200. For example, the order of some blocks may be changed, certain blocks may be combined with others, and some of the described blocks are optional. Method 200 may be performed by processing logic that may include hardware (circuitry, dedicated logic, etc.), software, or a combination of both.
[0050] In some embodiments, the exemplary method 200 may begin with a step of identifying malicious domains (block 202). For example, a malicious domain identifier 120 may identify a malicious domain. The malicious domain identifier 120 may identify a malicious domain from a set of received URLs (e.g., from an evaluation system 160). Domains that are likely to be malicious may be identified from all URLs marked by at least one scanner of the evaluation system 160. In some embodiments, a domain may be identified as likely to be malicious if a threshold number of scanners mark it. For example, a provided classification system may utilize historical VirusTotal (VT) URL feed information when identifying malicious domains. A basic measure of malice from VT results is the number of scanners that mark a URL as malicious. The higher this value for a given URL, the more likely it is to be malicious. In one example, a URL marked by five or more scanners may be identified as malicious. In other examples, URLs may be identified as malicious by utilizing different threshold numbers of scanners that mark URLs.
[0051] Figure 3 shows a graph comparing VT URL intelligence with SA and GSB, where "ext_malicious" corresponds to the percentage of URLs from each class of VT marked as malicious by either SA or GSB. When the number of scanners is less than 5, the majority of VTs marked as malicious URLs are not identified as malicious by either SA or GSB. On the other hand, for URLs marked as malicious by 5 or more scanners in a VT (i.e., #scanner5), the majority (over 70%) are consistent with external intelligence from SA and GSB.
[0052] Returning to Figure 2, in at least one example, the malicious domain identifier 120 sequentially profiles the domains observed in the VT URL feed. In such an example, the malicious domain identifier 120 incrementally constructs aggregated records for each fully qualified domain name (FQDN). The profile record for a given FQDN may include the time it was first seen, the time it was last seen, the number of times it was scanned, the number of times it was marked as malicious, and / or the corresponding URL and VT scan summary.
[0053] It is understood that the malicious domain identifier 120 can identify a malicious domain from a URL received from the evaluation system 160 when the evaluation system 160 is a blacklist or evaluation system other than a VT. In some embodiments, the malicious domain identifier 120 can identify a malicious domain from URLs received from multiple evaluation systems 160. For example, the malicious domain identifier 120 can cross-check results from one evaluation system 160 with results from another evaluation system 160.
[0054] Next, it may be determined whether the identified malicious domain is a public domain or a private domain (block 204). For example, a public domain classifier 130 may determine whether the identified malicious domain is a public domain or a private domain. Publicly available lists such as browser public suffix lists, CDN lists, dynamic DNS lists, popular web hosting domains, or proxy services may be useful, but they can also be very limited because they take time to keep up-to-date and therefore tend to include many non-existent domains and miss newly emerging public domains.
[0055] Ground truth datasets for public domain classifier 130 can be collected as follows: Publicly available lists, including public suffix lists, popular web hosting providers and CDN lists, and dynamic DNS lists, can be aggregated, and crossovers with apex domains in datasets DS1 and DS2 can be obtained. Potential public domains can be identified by searching the dataset for keywords likely to be used by public apex domains such as hosting domains, free domains, web domains, shared domains, upload domains, drop domains, cdn domains, file domains, photo domains, and proxy domains. A random sample of 500 apex domains can be taken from DS1 and DS2, respectively.
[0056] A provisional private domain ground truth data set can be collected by randomly selecting 1000 apex domains from each mutually exclusive dataset (DS1 and DS2) from a provisional public dataset. Manual validation may be performed from these provisional ground truth sets to create a final ground truth set. For each apex domain, a confidence score between 50 and 100 can be assigned to indicate the confidence level of the label, with 100 being the most confident and 50 being undecided. To improve the quality of labeling, two domain experts performed labeling for all domains and excluded domains with conflicting labels.
[0057] The public domain classifier 130 may take into account at least some of the features detailed in Table 1 below in order to determine whether a malicious domain is a public domain or a private domain. [Table 1]
[0058] Compared to private apex domains, public domains tend to host more subdomains, and these are also scanned more frequently by VTs. The features #subdomains and #scans capture these observations. Because subdomains are not under the control of the public apex domain owner, in practice, some subdomains are malicious and others are benign, while subdomains under private apex tend to be mostly benign or malicious. #malicious_scans and #malicious_scan_ratio capture volume and this difference. Most public apex, especially CDNs and proxy services, utilize the FQDN of the domain they serve (e.g., www.superwhys.com.akamai.com), while private apex primarily use descriptive popular keywords in the subdomain parts such as www, mail, ns, and m. By profiling all domains seen in PDNS during the study period, the inventors identified the top 100 subdomains as popular keywords. These differences are captured using the features #popular_keywords, #ratio_popular_keywords, and #mean_depth. The inventors observed that there is more variation among subdomain names under public apex domains than under private apex domains. Average sub-entropy measures the average entropy across all subdomains to capture this observation.
[0059] Figures 4 and 5 show graphs illustrating that the AUCs of the two ROC curves are 96% and 99% for GT1 and GT2, respectively, demonstrating a high degree of separability between the two classes.
[0060] Such FQDNs associated with the public domain can be created by attackers, and the number of such FQDNs can be used to calculate the public domain's value.
[0061] In some aspects, the public domain may be classified into one of seven groups: dynamic DNS, web proxy services, CDNs, web hosting, blogging and content hosting, content sharing services, and abbreviation tools and forms.
[0062] Returning to Figure 2, if the public domain classifier 130 determines that an identified malicious domain is a private domain, it can then determine whether the private domain is a compromised domain or an attacker-owned domain (block 206). For example, the private domain classifier 130 can determine whether the private domain is a compromised domain or an attacker-owned domain. To identify compromised domains, deviations in visual and auxiliary information between the apex domain and the domain under consideration are relied upon. The inventors observed that compromised domains have content that is very different from the main website, and auxiliary information such as hosting IP differs between the main website (evaluated hosting provider) and the domain under consideration (bulletproof hosting). On the other hand, attacker-owned domains have relatively recent registration information, are more likely to utilize high-speed flux networks, are short-lived (likely NX domains), and are blacklisted.
[0063] Two ground truth sets, AC-GT1 (AC stands for Attack-Owned / Attacked) and AC-GT2, for compromised apex domains and attacker-owned apex domains, can be manually constructed from private domains identified from DS1 and DS2, respectively, using public / private classifiers. A random sample of 2500 domains may be selected from each of DS1 and DS2. As with public / private ground truth collection, manual inspection of each sample may be performed, providing a confidence score to indicate how confident domain experts are about the label. The following information and sources are manually inspected to determine whether a malicious apex is compromised or attacker-owned: In addition to website verification, supplementary information such as registration information including historical WHOIS records, hosting information, and PDNS information was checked. Detailed reports from two threat intelligence platforms, riskiq.com and otx.alienvault.com, were also checked. Furthermore, detailed reports were inspected from two assessment services, GSB and SA.
[0064] To identify compromised domains, deviations in visual and auxiliary information between the apex domain and the domain under consideration were relied upon. The inventors observed that compromised domains had significantly different content compared to the main website, and that auxiliary information such as hosting IP differed between the main website (evaluated hosting provider) and the domain under consideration (bulletproof hosting). On the other hand, domains owned by attackers had relatively recent registration information, were likely to utilize high-speed flux networks, were short-lived (likely NX domains), and were blacklisted. After manual verification, a high-confidence label was selected.
[0065] In at least one example, the private domain classifier 140 considers at least five feature groups: vocabulary, VT report, VT profile, PDNS, and Alexa features. Vocabulary features capture the properties of the URL under consideration. VT report features include attributes directly available from the VT report, VT profile features are extracted from the VT NOD system, and PDNS features are extracted from the Farsight Passive DNS DB. The majority of the vocabulary, Alexa, and PDNS features are known from previous research in detecting malicious domains or URLs. The table shown in Figure 6 illustrates the various features of the five feature groups that the private domain classifier 140 can consider. Compared to conventional methods, novel features considered by the private domain classifier include VT_duration, positive_count, domain_malignant, #total_scans, #benign_scans, sibling_malignant, SOA_domain_number, and SOA_domain.
[0066] VT report features are extracted directly from the VT report. The inventors observed that the VT_Duration feature of compromised domains tended to be higher than that of attacker-owned domains. One reason for this is that compromised domains are generally difficult to detect by existing systems because attackers leverage the reputation of legitimate domains. For the same reason, the inventors observed that fewer scanners marked compromised sites as malicious than attacker-owned sites. The Positive_Count captures this observation. Compared to attacker-owned domains, attackers were observed to frequently use compromised domains as redirect sites to evade detection.
[0067] The VT profile feature captures the intuition that scans of almost all subdomains and attacker-owned domains are malicious, but only scans of some subdomains and compromised domains are malicious.
[0068] From PDNS features, the number of trusted name servers and SOA domains captures the observation that domains owned by attackers change hosting providers more frequently than benign domains to evade detection or takedown. Furthermore, attackers drop caught domains and leverage the residual trust within them, which also results in domains associated with multiple name servers. A comparison of apex domains with name server domains and SOA features captures the observation that benign domains are more likely to be hosted on their own servers compared to those owned by attackers.
[0069] This disclosure improves upon several lexical features presented in previous research. Specifically, the inventors observed that attacker-owned domains are more likely to impersonate brands using these squatting methods compared to compromised domains. This disclosure profiles the top 1 million Alexa domains over a year to identify the top 1000 Alexa brands for detecting CoboSquat, LevelSquat, and Targeted Embedding Domains, which have been shown to be hundreds of times more common than more traditional squatting types. Brand, similar, and popular_keyword features capture new squatting tactics used by attackers.
[0070] In addition to the VT features shown in the table in Figure 6, the private domain classifier 140 considers three new classes of features—PDNS, Alexa, and lexical features—to improve classification performance. This actually improves the performance matrix, and as shown in Figure 7, several classifiers, including GB, ET, and RF, perform very well, resulting in an accuracy slightly above 90% with 10x cross-validation of AC-GT1. Figure 8 shows a graph of the ROC curve and feature importance in an example where the private domain classifier 140 is a random forest classifier. The private domain classifier 140 achieves an accuracy of 90.6% with a precision of 94.7% and a recall of 86.1%. A key consideration when building robust machine learning models is that the model should generalize to different ground truth datasets. For this purpose, a new model is trained using AC-GT2. Using the RF classifier, the inventors achieved an accuracy of 96.8% with a precision of 99.1% and a recall of 93.4%. Figure 9 shows a graph illustrating the ROC curve for an example where the private domain classifier 140 is a random forest classifier.
[0071] The inventors made various insights into the VT URL Feed dataset to help determine the features used by the public domain classifier 130 and the private domain classifier 140. The VT URL Feed dataset contains 814,678,956 unique URLs for the period from August 1, 2019 to November 18, 2019. Note that the same URL may be scanned multiple times in a day. Each new scan is considered a different scan. However, if VT is queried multiple times simply to retrieve an existing report instead of triggering a new scan, it does not change the scan ID. Therefore, multiple such reports with the same scan ID are considered a single record. The average daily number of potentially benign scans observed (i.e., #scanner=0) was observed to be 89.3% of the total number of scans, approximately 4.8M. The inventors observed that, on average, malicious URLs were scanned 6 times, while benign URLs were scanned only 2 times. This follows general user behavior where the more suspicious a URL is, the more times they are checked. Another observation was that the average daily scan count was approximately twice the average URL count.
[0072] The inventors also compared the coverage of malicious websites in their dataset with that of typical blacklists and rating services. While many VT reports have one or two #scanners, an average of 45.7% of malicious scans have five or more #scanners (i.e., the top two areas in the figure). The inventors noted that scans with five or more #scanners correspond to an average of 1659K malicious reports per week. This corresponds to an average of 276K malicious websites per week. In contrast, Google Transparency Report and Phishtank report approximately 50K and 4K per week, respectively. This indicates that classification system 110 is trained on a much larger set of malicious URLs compared to typical blacklists and therefore has a greater impact.
[0073] The VT scanner assigns each malicious URL one of the following class labels: malicious, malware, phishing, mining, and suspicious site. In most cases, the VT scanner assigns conflicting class labels, so a simple majority-vote heuristic may be used to derive the final class label for a malicious website. For example, we took a random sample of 100 websites for each class type and manually cross-checked them against several publicly available blacklists or APIs, including fish tanks, GSB, and SA. Our manual inspection validated our heuristic by showing that over 98% of the labels using majority vote were consistent with external intelligence. While malware and phishing sites dominate the reported malicious websites, the dataset contains only a small number of malicious mining sites and suspicious sites.
[0074] Figure 10 shows a graph illustrating the CDF of the number of FQDNs per apex over a given period for presumed benign domains (i.e., #scanner=0) and malicious domains (i.e., #scanner=5). Frequencies less than 5 are excluded, and long tails with frequencies greater than 500 are excluded. It can be seen that 90.2% of apex in the benign category have only one FQDN, compared to only 12.3% of apex in the malicious category. Furthermore, approximately 40% of malicious apex domains have more than 40 FQDNs, compared to only 5% of benign apex domains. These observations suggest that attackers create many subdomains to initiate attacks in a manner similar to that of fast flux networks.
[0075] Another observation is the existence of a long tail of apex domains with more than 500 FQDNs, some of which have millions. For example, blogspot.com (blogging), coop.it (URL shortening tool), mcafee.com (mcafee endpoint host), and opendns.com (Cisco Open DNS) all have more than 1 million FQDNs. The number of observed FQDNs is used as a feature in the Public Domain Classifier 130 because a higher number indicates a higher likelihood that the domain is public.
[0076] Returning to Figure 2, the determined malicious domain hosting type may then be displayed (block 208). For example, the classification system 110 may display the determined malicious domain hosting type on display 116. The displayed malicious domain hosting type may be displayed along with the malicious domain URL. The determined malicious domain hosting type may be a public domain (e.g., a public domain owned by the attacker), a compromised private domain, or a private domain owned by the attacker. A security operator can view the determined malicious domain hosting type and the URL of the malicious domain on display 116 to determine and take appropriate action.
[0077] Experimental verification Our analysis identified 6,675 malicious public apex domains and 725,325 malicious private apex domains in both datasets. In other words, only 0.91% of apex domains in the VT URL feed are public. However, we observed a high proportion of URLs and scans belonging to these public apex domains. Of all reports, 46.5% of URLs were hosted on public apex domains. This observation is consistent with the fact that public apex domains host many subdomains, while private apex domains generally host only a few.
[0078] Figure 11A shows a graph illustrating the number of FQDNs per apex for apex domains in two categories: public and private. Over 80% of public apex domains have more than 20 FQDNs, while 95% of private apex domains have fewer than 10 FQDNs. While many public domains have a large number of subdomains, there is a long tail of public domains with a massive number of subdomains (over 200K). These observations suggest that attackers prefer to create many subdomains under public apex domains because they are freely available and can leverage the reputation of public apex domains, including their TLS certificates, hosting, and registration information, potentially making them less easily detected by traditional blacklists and reputation systems.
[0079] Figure 11B shows a graph illustrating the average Alexa ranking distribution for public and private apex domains. For unranked domains, one million non-significant ranks were assigned for better visualization. It is not surprising that public domains have a higher average Alexa ranking compared to private domains, as they are accessed more frequently by users. An interesting result is that half of the public domains are unpopular (unranked), indicating that attackers also create subdomains on less popular public domains to launch attacks. Since public apex domains host many benign domains, current registration and domain reputation-based and inference-based systems may inadvertently blacklist public apex domains, disrupting benign sites.
[0080] Domain lifetime can be estimated by taking the lifetime of the PDNS footprint of each apex domain. Figure 11C is a graph showing the domain lifetime distribution of public and private apex domains. Despite the vast majority of public domains being unranked sites that provide attackers with a free platform to launch attacks over extended periods, the inventors observed that public domains have a longer lifetime compared to private domains. Furthermore, approximately 10% of private domains are very short-lived, suggesting that they are likely to be domains owned by attackers.
[0081] The Private Domain Classifier 140 detected that 65.6% of apex domains in the VT URL feed were compromised, indicating that some websites were more compromised than those owned by attackers. This observation is consistent with previous research conducted on phishing websites and public threat intelligence reports.
[0082] Figure 12A shows a graph displaying #FQDNs per apex for compromised domains and domains owned by attackers. An interesting observation is that most compromised domains host slightly more malicious subdomains than those owned by attackers. If mitigation measures are taken, it is crucial that domain owners first identify and clean up all malicious subdomains, which could number exceed 500.
[0083] Figure 12B shows a graph illustrating the average Alexa rank distribution of compromised apex domains and attacker-owned apex domains. As expected, most attacker-owned domains have either a low Alexa rank or no rank at all. However, it is interesting to note that there are several attacker-owned domains with an Alexa ranking of less than 100K. Another interesting observation is that there are approximately 10% of compromised domains that are unranked, indicating that attackers will also launch attacks from less popular, benign websites that can be used to launch attacks such as DDoS attacks that do not require a ranked site.
[0084] Figure 12C shows a graph illustrating the domain lifetime distribution of compromised apex domains and attacker-owned apex domains. Generally, it is not surprising that compromised domains survive longer than those owned by attackers. However, approximately 40% of attacker-owned domains remain active for over 200 days, indicating the need to develop better techniques for early detection of these malicious domains and taking appropriate action. One reason for their long duration is that attackers register domains and park them for a period of time as evasion techniques.
[0085] Figure 13 shows the feature correlation matrix of the features used in the public domain classifier 130. Figures 14A and 14B show graphs illustrating the feature importance of the random forest-based public domain classifier 130 for two datasets, GT1 and GT2. The feature importance graphs show which features are important when building the model. Figures 15A and 15B show graphs illustrating the t-SNE of the random forest-based public domain classifier 130 for the two datasets, GT1 and GT2. The t-SNE graphs utilize a nonlinear dimensionality reduction technique to embed feature vectors into 2D spatial data for visualization. These show how the two classes are clustered based on the collected features. One reason for the better performance on the second ground truth set is that, as shown in Figures 15A and 15B, the two classes in the ground truth data have better separation in the second set, resulting in a better decision boundary. Furthermore, the feature importance graphs show that almost all features are involved in the label determination, which reduces bias towards adversarial operations and makes it less susceptible to what is important. Figures 16A and 16B are graphs showing the precision-recall of 130 random forest-based public domain classifiers on two datasets, GT1 and GT2.
[0086] Figures 17A and 17B show graphs illustrating the feature importance of a random forest-based private domain classifier 140 for two datasets, GT1 and GT2. Figures 18A and 18B show graphs illustrating the t-SNE of the random forest-based private domain classifier 140 for two datasets, GT1 and GT2. Figures 19A and 19B show graphs illustrating the precision-recall of the random forest-based private domain classifier 140 for two datasets, GT1 and GT2.
[0087] Without further detail, a person skilled in the art will likely find the above description useful for making the most of the claimed invention. The examples and embodiments disclosed herein should be construed as illustrative only and not in any way limit the scope of this disclosure. It will be apparent to a person skilled in the art that modifications can be made to the details of the above examples without departing from the basic principles described. In other words, various modifications and improvements to the examples specifically disclosed above are within the scope of the appended claims. For example, any suitable combination of features of the various examples described is conceivable.
Claims
1. 1. A system for classifying malicious domain hosting types, the system comprising: The display and Memory and a processor in communication with the memory; Equipped with The processor: Identifying malicious domains from the set of received domains; Using a model to determine whether the identified malicious domain is a public domain or a private domain; If the identified malicious domain is a private domain, using a model to determine whether the private domain is a compromised domain or an attacker-owned domain; Displaying the determined malicious domain hosting type on the display, the determined malicious hosting type being a public domain, a compromised private domain, or a private domain owned by an attacker. It is configured as follows: system.
2. 1. A method for classifying malicious domain hosting types, comprising: identifying malicious domains from the received set of domains; using a model to determine whether the identified malicious domain is a public domain or a private domain; If the identified malicious domain is a private domain, using a model to determine whether the private domain is a compromised domain or an attacker-owned domain; displaying the determined malicious domain hosting type, wherein the determined malicious hosting type is a public domain, a compromised private domain, or an attacker-owned private domain; A method comprising: