Illegal website detection method and device based on third-party service ID
By constructing ID matching rules and models to identify illegal communities, the limitations and disguise problems of illegal website detection in existing technologies are solved, and automated detection and rapid identification of illegal websites are achieved.
Patent Information
- Application Number
- CN202310019128.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-01-06
AI Technical Summary
Existing technologies are limited to specific third-party services in detecting illegal websites, lack automated detection methods, and have difficulty discovering illegal domain names disguised under camouflage technology.
By constructing ID matching rules, whitelist IDs are extracted from legitimate domain names. The whitelist IDs are used to filter the websites to be detected, clustered to form communities, and the semantic features of community domain names, website ID features, and community statistical features are extracted. The model is used to identify illegal communities.
It realizes the automatic detection of illegal websites, can find illegal domain names under camouflage technology, and quickly discover new illegal domain names by identifying the IDs of illegal communities.
Smart Images

Figure CN116055155B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of security detection, and in particular to a method and device for detecting illegal websites based on a third-party service ID. Background Art
[0002] Illegal websites have been restricted by government agencies and app markets due to their negative social impact. However, easily deployable third-party services allow cyberattacks to quickly deploy websites to bypass censorship, making the detection of rapidly changing websites essential. To provide differentiated services, third-party service providers typically require website request URLs to contain identity credentials (hereinafter referred to as IDs). These URLs typically appear on websites as JavaScript code or links. Third-party service IDs are typically unique, allowing them to link different websites belonging to the same website administrator. Starov et al. (Starov, Oleksii, et al. "Betrayed by your dashboard: Discovering malicious campaigns via web analytics." Proceedings of the 2018 World Wide Web Conference. 2018.) extracted 18 different analytics service IDs from phishing websites and used the collected IDs as a blacklist. They found that these blacklisted IDs could be used to discover new phishing websites. Yang et al. (Yang, Hao, et al. "Casino royale: a deep exploration of illegal online gambling." Proceedings of the 35th Annual Computer Security Applications Conference. 2019.) measured illegal websites and found that many illegal websites shared third-party analysis service IDs and third-party customer service service IDs. For the convenience of description, the present invention refers to all websites associated with the same ID as a community. Examples of communities are as follows: Figure 1 shown.
[0003] The above research provides a new approach to detecting illegal websites: extracting IDs from illegal websites and using these IDs to discover new illegal domain names. However, these studies have the following shortcomings: 1) They focus only on specific third-party services, such as third-party analytics and customer service. In reality, many third-party services include IDs, such as third-party gaming and website building services. 2) Most of these studies focus on measuring and analyzing IDs, lacking automated methods for detecting illegal websites using IDs. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the present invention discloses a method for detecting illegal websites based on third-party service IDs. This method expands the research scope of third-party services and is not limited to specific third-party services.
[0005] The technical contents of the present invention include:
[0006] A method for detecting illegal websites based on a third-party service ID, the method comprising:
[0007] Use whitelist IDs to filter multiple websites to be detected, and construct a community to be detected based on the filtering results of the websites;
[0008] Extracting the domain name semantic features, website ID features and community statistical features of the community to be detected;
[0009] Based on the semantic features of community domain names, website ID features and community statistical features, illegal detection results of multiple websites to be detected are obtained.
[0010] Furthermore, the method of filtering multiple websites to be detected by using the whitelist ID and constructing a community to be detected based on the filtering results of the websites includes:
[0011] Establish ID matching rules;
[0012] Based on the ID matching rules, extract the website ID from the legal domain name to obtain the whitelist ID;
[0013] Use the whitelist ID to filter multiple websites to be detected and obtain suspicious websites;
[0014] Use website IDs to cluster websites to obtain several communities;
[0015] Communities with more than 2 domain names are considered as communities to be detected.
[0016] Furthermore, the ID matching rules include: the website ID contains at least one number, the website ID appears after the '?' character or the '=' character, the website ID length is within a specified length range, and there are no more than two website IDs in a URL.
[0017] Furthermore, the step of extracting semantic features of the community domain name of the community to be detected includes:
[0018] Preprocessing the domain name of the community to be detected; the preprocessing includes: mapping different characters in the domain name into numbers and aligning the length of the domain name;
[0019] Based on the preprocessed domain names, a domain name semantic matrix is constructed;
[0020] Utilize the kernel to capture the semantic similarity in the domain name semantic matrix, obtaining a number of semantic vectors;
[0021] Horizontally merge the semantic vectors to obtain the community domain name semantic features of the to-be-detected community.
[0022] Furthermore, the extraction of the website ID features of the to-be-detected community includes:
[0023] Utilize a heterogeneous graph to capture the connection relationships in the to-be-detected community; wherein, the edges of the heterogeneous graph include: <ID,Domain>, <Domain,IP>, <Domain,Cname>, <Domain,Whois registrar>, <Domain,CA registrar>, and <IP,AS>; input the connection relationships into the HAN model to obtain the website ID features of the to-be-detected community.
[0024] Furthermore, the community statistical features include: the average length of all domain names, the average number of digits in all domain names, the average number of letters in all domain names, the average of the longest consecutive digits in all domain names, the average number of meaningful words in all domain names, the New TLD diversity of all domain names, the diversity of all domain name TLDs, the Levenshtein distance between all domain names, the IP deviation value of all domain names, the AS deviation value of all domain names, the difference in registration time between all domain names, the average number of Whois registrars of all domain names, and the average survival time of all domain names.
[0025] Furthermore, the New TLD diversity W(c) of all domain names = -∑P i (x)log2(P i (x)); wherein, P i (x) represents the occurrence probability of the i-th New TLD.
[0026] Furthermore, the IP deviation value of all domain names where m represents the number of IPs corresponding to domain names in the community, and IP k and IP l respectively represent the IP of the k-th domain name and the IP of the l-th domain name,
[0027] Furthermore, the difference in registration time between all domain names where time i represents the registration time of the i-th domain name, and n represents the total number of domain names in the to-be-detected community.
[0028] An illegal website detection device based on a third-party service ID, the device includes:
[0029] The community construction module is used to filter multiple websites to be detected using the whitelist ID and construct a community to be detected based on the filtering results of the websites;
[0030] Feature extraction module, used to extract the community domain semantic features, website ID features and community statistical features of the community to be detected;
[0031] The illegal detection module is used to obtain illegal detection results of multiple websites to be detected based on community domain name semantic features, website ID features and community statistical features.
[0032] The present invention has the following beneficial effects compared to the prior art:
[0033] 1) The present invention extracts community-level features, which are more difficult for attackers to bypass and can detect domain names that use camouflage techniques (page redirection, page mixing with legitimate content).
[0034] 2) After using the present invention to identify illegal communities, the IDs corresponding to the illegal communities can be used as blacklists. Based on existing observations, IDs will appear on multiple websites and will not change frequently, so illegal domain names can be quickly discovered using IDs. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Example graph of a community.
[0036] Figure 2 Flowchart of illegal website detection method.
[0037] Figure 3 An example diagram of a domain name semantic matrix. DETAILED DESCRIPTION
[0038] To further illustrate the technical solution of the present invention, the above steps are described in detail below through drawings and specific examples, but the examples given are not intended to limit the present invention.
[0039] The illegal website detection method of the present invention first extracts the ID from the website to be detected. If the ID hits the blacklist, the website is marked as an illegal domain name and output. The blacklist ID of the first run comes from the positive sample in the model training process. If the ID does not hit the blacklist, the ID is filtered by the whitelist. For the filtered ID, multiple websites are clustered using the ID to form a community. Communities with more than 2 websites are sent to the model for identification. After judgment, the model outputs illegal communities. The domain names contained in the illegal communities are illegal domain names. In addition, the model will also output the IDs corresponding to the illegal communities, and use these IDs as blacklist IDs for subsequent identification.
[0040] Specifically, if Figure 2 As shown, the present invention includes the following steps:
[0041] Step 1: Use the whitelist ID to filter multiple websites to be detected, and construct a community to be detected based on the filtering results of the websites.
[0042] In order to broaden the research scope of third-party services, the present invention first needs to construct a universal ID matching rule, which is as follows:
[0043] 1) The ID must contain at least one number;
[0044] 2) ID must appear after the '?' or '=' character;
[0045] 3) The ID length must be at least 5 and less than 32 characters;
[0046] 4) A URL can contain no more than 2 IDs.
[0047] This method uses this ID matching rule to first extract IDs from legitimate domain names (e.g., Tranco's top 1 million domain names), which are then used as a whitelist. For multiple websites to be tested, IDs are first extracted from the websites and filtered using the whitelisted IDs. The websites are then clustered using the IDs to create several communities. Communities with more than two domain names are selected as the communities to be tested, and then classified using the model.
[0048] Step 2: Extract the domain name semantic features, ID features and community statistical features of the community to be detected.
[0049] In order to automatically detect illegal websites using IDs, this paper designs a detection model. The model involves three feature extractors and one classifier, which are described in detail as follows:
[0050] Feature extractor 1 is designed to capture the semantic similarity of community domain names. The present invention maps all domain names in the same community into a domain name semantic matrix. The specific steps include: 1) mapping different characters in the domain name into numbers; 2) setting the maximum length of the domain name to 13, truncating it if it exceeds 13, and padding it if it is less than 13; 3) forming the processed domain name into a domain name semantic matrix. The present invention designs 3 kernels to capture semantic similarity. Assuming that the number of domain names in the community is n, the sizes of the 3 kernels are 1×n, 2×n, and 3×n respectively. Each kernel is used to calculate the variance of all numbers in the current window, for example Figure 3As shown, use each kernel to scan the domain name semantic matrix from left to right to obtain semantic vectors. Since the maximum length of the domain name is 13, the sizes of the semantic vectors obtained by scanning with kernels of 1×n, 2×n, and 3×n are 13×1, 12×1, and 11×1 respectively. Horizontally merge the three semantic vectors to finally obtain 36-dimensional features. Then use a 1D CNN to reduce the dimensions of the features, and finally obtain X-dimensional features.
[0051] Feature extractor 2 aims to capture the connection relationship between the community domain name and the infrastructure. The present invention uses a heterogeneous graph to capture the connection relationship. The edges of the heterogeneous graph include <ID,Domain>, <Domain,IP>, <Domain,Cname>, <Domain,Whois registrar>, <Domain,CA registrar>, <IP,AS>. Then use the HAN model to extract the features of ID, and the dimension of the extracted features is X-dimensional.
[0052] Feature extractor 3 aims to capture the statistical features of the community. The specific features include the average length of all domain names in the community (F1), the average number of digits in all domain names (F2), the average number of letters in all domain names (F3), the average of the longest consecutive digits in all domain names (F4), the average number of meaningful words in all domain names (F5), the New TLD diversity of all domain names (F6), the diversity of all domain name TLDs (F7), the Levenshtein distance between all domain names (F8), the deviation value of the IP of all domain names (F9), the ASN deviation value of all domain names (F10), the difference in registration time between all domain names (F11), the average number of whois registrars of all domain names (F12), and the average survival time of all domain names (F13).
[0053] The calculation formulas for F6 and F7 are as follows: W(c) = -∑P i (x)log2(P i (x)), P i (x) represents the occurrence probability of the i-th TLD / NewTLD. The calculation formulas for F9 and F10 are as follows: T k represents the IP or ASN of the k-th domain name, T l represents the IP or ASN of the l-th domain name, m represents the number of IPs or ASNs corresponding to the domain names in the community (since multiple domain names may correspond to the same IP / ASN, so the m IPs / ASs are repeated), if T k = T l ,, then h(T k ,T l ) = 1, otherwise h(T k ,T l)=0. The calculation formula of F11 is as follows: where time i Indicates the registration time of the i-th domain name.
[0054] Step 3: Based on the semantic features, ID features and community statistical features of the community domain name, illegal detection results of multiple websites to be detected are obtained.
[0055] The present invention horizontally concatenates the features obtained by the three classifiers to obtain a 33-dimensional feature set. This classifier uses a linear fully connected layer to classify communities. Domain names in the same community belong to the same developer, so all domain names in illegal communities are illegal domain names.
[0056] Finally, for the illegal communities discovered by the classifier, the present invention extracts the corresponding IDs and uses them as a blacklist to quickly discover illegal domain names.
[0057] The above-described embodiments are merely specific ways of presenting the present invention. Any technical solutions that can be easily obtained by simple transformation or equivalent replacement of the above embodiments fall within the protection scope of the present invention.
Claims
1. A method for detecting illegal websites based on third-party service IDs, characterized in that: The method includes: Filtering multiple websites to be detected using a whitelist ID, and constructing a community to be detected based on the filtering results of the websites. The process of constructing the community to be detected includes: Establishing an ID matching rule, which includes: at least one digit in the website ID, the website ID appears after the '?' character or the '=' character, the length of the website ID is within a specified length range, and there are at most 2 website IDs in one URL; Extracting the website ID from the legal domain names based on the ID matching rule to obtain the whitelist ID; Filtering multiple websites to be detected using the whitelist ID to obtain suspicious websites; Clustering the suspicious websites using the website ID to obtain several communities; Regarding the community with more than 2 domain names as the community to be detected; Extracting the semantic features of the domain names, website ID features, and community statistical features of the community to be detected. Among them, The extraction of the semantic features of the domain names of the community to be detected includes: Preprocessing the domain names of the community to be detected. The preprocessing includes: mapping different characters in the domain name to numbers and aligning the lengths of the domain names; Constructing a domain name semantic matrix based on the preprocessed domain names; Using a kernel to capture the semantic similarity in the domain name semantic matrix to obtain several semantic vectors; Horizontally merging the semantic vectors to obtain the semantic features of the domain names of the community to be detected; The extraction of the website ID features of the community to be detected includes: Using a heterogeneous graph to capture the connection relationships in the community to be detected. Among them, the edges of the heterogeneous graph include: <ID,Domain>, <Domain,IP>, <Domain,Cname>, <Domain,Whois registrar>, <Domain,CA registrar>, and <IP,AS>; Inputting the connection relationships into a HAN model to obtain the website ID features of the community to be detected; The community statistical features include: the average length of all domain names, the average number of digits in all domain names, the average number of letters in all domain names, the average of the longest consecutive digits in all domain names, the average number of meaningful words in all domain names, the New TLD diversity of all domain names, the TLD diversity of all domain names, the Levenshtein distance between all domain names, the IP deviation value of all domain names, the AS deviation value of all domain names, the difference in registration time between all domain names, the average number of whois registrars of all domain names, and the average survival time of all domain names; Obtaining the illegal detection results of multiple websites to be detected based on the semantic features of the domain names, website ID features, and community statistical features of the community to be detected.
2. The method according to claim 1, wherein New TLD diversity of all domain names W(c)=-∑P i (x)log2(P i (x)); where P i (x) represents the occurrence probability of the i-th NewTLD.
3. The method according to claim 1, wherein The deviation value of the IP of all domain names Among them, m represents the number of IP addresses corresponding to domain names in the community. k 、IP l Represent the IP of the k-th domain name and the IP of the l-th domain name, 4. The method according to claim 1, wherein The difference in registration time between all domain names Among them, time i represents the registration time of the i-th domain name, and n represents the total number of domain names in the community to be detected.
5. An illegal website detection device based on a third-party service ID, characterized in that: The device includes: A community construction module for filtering multiple websites to be detected using a whitelist ID and constructing a community to be detected based on the filtering results of the websites. The construction of the community to be detected includes: Establish ID matching rules, which include: at least one digit is included in the website ID, the website ID appears after the '?' character or the '=' character, the length of the website ID is within a specified length range, and there are at most 2 website IDs in one URL; Based on the ID matching rules, extract the website ID from legal domain names to obtain the whitelist ID; Use the whitelist ID to filter multiple websites to be detected to obtain suspicious websites; Cluster websites using the website ID to obtain several communities; Regard the communities with more than 2 domain names as communities to be detected; A feature extraction module for extracting the semantic features of the domain names of the communities to be detected, the website ID features, and the community statistical features; where, The extraction of the semantic features of the domain names of the communities to be detected includes: Preprocess the domain names of the communities to be detected; the preprocessing includes: mapping different characters in the domain names to numbers and aligning the lengths of the domain names; Based on the preprocessed domain names, construct a domain name semantic matrix; Use a kernel to capture the semantic similarity in the domain name semantic matrix to obtain several semantic vectors; Horizontally merge the semantic vectors to obtain the semantic features of the domain names of the communities to be detected; The extraction of the website ID features of the communities to be detected includes: Use a heterogeneous graph to capture the connection relationships in the communities to be detected; where, the edges of the heterogeneous graph include: <ID,Domain>, <Domain,IP>, <Domain,Cname>, <Domain,Whois registrar>, <Domain,CA registrar> and <IP,AS>; Input the connection relationships into the HAN model to obtain the website ID features of the communities to be detected; The community statistical features include: the average length of all domain names, the average number of digits in all domain names, the average number of letters in all domain names, the average of the longest consecutive digits in all domain names, the average number of meaningful words in all domain names, the New TLD diversity of all domain names, the TLD diversity of all domain names, the Levenshtein distance between all domain names, the IP deviation value of all domain names, the AS deviation value of all domain names, the difference in registration time between all domain names, the average number of whois registrars of all domain names, and the average survival time of all domain names; An illegal detection module for obtaining the illegal detection results of multiple websites to be detected based on the semantic features of the domain names of the communities, the website ID features, and the community statistical features.
Citation Information
Patent Citations
Pirate video website detection method and system based on third-party service
CN113163234A
Phishing website detection method and system based on heterogeneous graph feature extraction
CN115065518A