A phishing website detection method based on URL multi-angle features

By performing multi-angle feature extraction and lightweight machine learning on URLs, the shortcomings of existing phishing website detection methods in terms of real-time performance and resource usage are solved, and efficient detection of unseen phishing websites is achieved, adapting to the rapidly changing network environment.

CN115766212BActive Publication Date: 2025-09-05SOUTHEAST UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202211422976.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-09-05
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Existing phishing website detection methods are inefficient and have poor real-time performance when facing zero-day phishing attacks. They also consume a lot of computing resources and are difficult to adapt to the rapidly changing network environment.

Method used

By extracting multi-angle features of the URL, including component features and language features, and combining them with machine learning algorithms to determine the legitimacy of the URL, the complete URL and component features are extracted, and a lightweight model is used for real-time detection.

Benefits of technology

It realizes the detection of never-before-seen phishing websites, reduces the demand for computing resources, improves the real-time and generalization capabilities of detection, and adapts to complex and changing network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766212B_ABST
    Figure CN115766212B_ABST
Patent Text Reader

Abstract

The present invention discloses a phishing website detection method based on multi-angle features of URLs, which belongs to the field of information security technology. The method first captures the actual URL corresponding to the web page as the URL to be tested, and then decomposes the URL to be tested to obtain information of each component; then pre-processes the URL and components, including using a text decomposition algorithm to obtain a token and a word list, and using a text readability detection algorithm to calculate the token readability weight; extracts component features and language features from the complete URL, each component and the pre-processing result as features of the URL to be tested; finally, inputs the URL features into a trained machine learning classifier for legitimacy judgment. Compared with the list-based method, the method can detect URLs that have not appeared before; compared with the method based on visual similarity and content, the method does not need to wait for the page to load and has higher real-time performance; compared with similar methods based on URLs, the method occupies fewer resources, has richer features and stronger generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a phishing website detection method based on URL multi-angle features, belonging to the field of information security technology. Background Art

[0002] Phishing is a type of cyberattack that uses social engineering and technical means. Criminals typically disguise themselves as legitimate entities to gain users' trust and send messages via email or social networks to deceive, pressure, or manipulate users into revealing personal data (such as bank account numbers, credit card numbers, login credentials, etc.) or downloading malware.

[0003] Most anti-phishing tools typically use a list-based approach to improve detection efficiency. For example, Google SafeBrowsing maintains a list of URLs containing malware and phishing. Many browsers, including Google and Firefox, subscribe to this service to check for malicious links. McAfee SiteAdvisor, accessible as a browser plug-in, displays website safety ratings. eBay Toolbar, Egress Defend, Netcraft Toolbar, ZoneAlarm, and many other tools are also widely used for phishing website detection. List-based approaches are simple and effective, but require frequent updates and maintenance and are difficult to combat against zero-day phishing attacks.

[0004] Phishing attacks are becoming increasingly serious, and researchers are trying to extract more effective features from web pages to detect phishing websites. Methods based on visual similarity detect phishing websites by comparing the similarity of visual features such as logos, pictures, and screenshots. Such methods have complex models, high computational complexity, and high time costs. Content-based methods detect phishing attacks by extracting web page features. Feature sources include URLs, web page source code, whois information, web page rankings, etc. Detection types include URL analysis, image analysis, text analysis, email analysis, spelling and grammar checking, DNS checking, etc. Methods based on visual similarity and content (such as the existing patents "A method for detecting phishing websites based on multi-feature fusion CN201810373630.1" and "Phishing website attack detection method, device, electronic device and storage medium CN202210553089.9") need to wait for the page to load, which seriously affects the real-time performance of phishing detection.

[0005] Extracting features from URLs to detect phishing websites can detect zero-day phishing attacks while also being highly real-time. Existing patents, "A Phishing Website URL Detection Method Based on Deep Learning (CN201810750707.2)" and "A Phishing Website Detection Method for URLs (CN202011361704.3), both use neural networks (such as CNN and LSTM) to extract features from URLs for classification. While deep learning-based phishing URL detection can avoid feature engineering, its models are complex, costly, and time-consuming to train. It consumes a lot of computing resources and is not easily usable on resource-constrained devices. Among the machine learning-based methods, the existing patent "Phishing website URL detection method and system based on machine learning CN202110231656.4" lacks analysis of the various components of the URL; the features used in "A phishing website detection method based on URL string random rate feature extraction CN202110359991.2" and "Method and device for identifying phishing websites CN201510885473.9" include third-party features, which will increase the time of feature extraction and have poor applicability in real-time scenarios. Summary of the Invention

[0006] In order to solve the above problems, the present invention discloses a phishing website detection method based on multi-angle features of URL, which belongs to the field of information security technology. The method first captures the actual URL corresponding to the web page as the URL to be tested, and then decomposes the URL to be tested to obtain information of each component; then preprocesses the URL and components, including using a text decomposition algorithm to obtain a token and a word list, and using a text readability detection algorithm to calculate the token readability weight; extracts component features and language features from the complete URL, each component and the preprocessing results as features of the URL to be tested; finally, inputs the URL features into a trained machine learning classifier for legitimacy judgment. Compared with the list-based method, the method can detect URLs that have not appeared before; compared with the method based on visual similarity and content, the method does not need to wait for the page to load and has higher real-time performance; compared with similar methods based on URL, the method occupies fewer resources, has richer features and stronger generalization ability.

[0007] To achieve the purpose of the present invention, the specific technical steps of this solution are as follows: A phishing website detection method based on URL multi-angle features, the method comprising the following steps:

[0008] Step (1) capture the actual URL of the target website as the URL to be tested;

[0009] Step (2) decomposes the URL to be tested to obtain the URL's scheme, authority, path, parameters, query, anchor and other components;

[0010] Step (3) extracting component features and language features from the URL and its components to obtain URL features;

[0011] Step (4) inputs the URL features into the machine learning detection model to determine the legitimacy of the URL.

[0012] Furthermore, in step (1), due to reasons such as URL shortening service and redirection, the URL accessed by the user is inconsistent with the URL of the final target website. In this method, the final URL of the target website is used as the URL to be tested.

[0013] Furthermore, in step (2), a complete URL consists of several components, among which the scheme is usually called the protocol part, and the authority contains four subcomponents: username, password, host name, and port number. Among all the components, the protocol and host name are required components, the path and query are optional components, and the remaining components are classified as uncommon components.

[0014] Furthermore, in step (3), the component features extracted from the URL and each component include:

[0015] (3.1.1) URL component characteristics: URL length, use of HTTPS, number of uncommon components, number of suspicious symbols in the URL, number of numbers in the URL, use of double slashes, and number of URL tokens;

[0016] (3.1.2) Host name (domain) component characteristics: domain name length, number of subdomains, use of IP addresses, number of "-", use of top-level domain names in subdomains, use of shortening services, and number of digits in the domain name;

[0017] (3.1.3) Path component characteristics: path length, path depth, number of suspicious symbols in the path, use of top-level domain names in the path, maximum length of tokens in the path, use of file extensions;

[0018] (3.1.4) Component characteristics of the query: length of the query, number of numbers in the query, number of queries.

[0019] Furthermore, in step (3), extracting language features specifically includes the following sub-steps:

[0020] (3.2.1) Build a relative word frequency database. Based on the word frequency database and dynamic programming principles, write a continuous text decomposition algorithm to break the text into several words and ensure the correctness of the words as much as possible;

[0021] (3.2.2) Based on the N-Gram language model and the Markov chain principle, a text readability detection algorithm is constructed to calculate the ease or difficulty of understanding the text;

[0022] (3.2.3) Split the URL components by the primary delimiter to obtain a list of intermediate tokens. The primary delimiter in the domain name is ".", the primary delimiter in the path is " / ", and the primary delimiters in the query are "&" and "=".

[0023] (3.2.4) Use the continuous text decomposition algorithm to decompose the intermediate token list of the component, take the logarithm of the number of word segments of all tokens and sum them to obtain the text dispersion, and count all word segments with a length less than 2 to obtain the number of short tokens in the text.

[0024] (3.2.5) Use the text readability detection algorithm to calculate the readability of each token of the component. If the readability is greater than a pre-set threshold, it is considered unreadable, otherwise it is considered readable.

[0025] (3.2.6) Among the linguistic features extracted from URL components, hostname linguistic features include: the number of sensitive words (variations of "www" and the use of "https"), the number of vowels, the number of unique characters, the dispersion of subdomains, the number of short tokens in a domain, the readability of second-level domains, and the maximum readability of tokens in a subdomain; path linguistic features include: the dispersion of the path, the number of short tokens in the path, and the number of unreadable tokens in the path. Compositional and linguistic features fully reflect the differences in URL components and languages, making URL features more diverse and improving the generalization ability of the model.

[0026] Furthermore, in step (4), the machine learning algorithms trained include random forest, XGBOOST, artificial neural network, decision tree, gradient boosting tree, AdaBoost, K nearest neighbor, support vector machine, logistic regression, naive Bayes, etc.; the URL features are input into the machine learning classifier to determine the legitimacy of the URL. If it is a phishing URL, a warning is issued to the user, otherwise the user activity proceeds normally.

[0027] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects.

[0028] (1) The phishing website detection method proposed in this invention combines heuristics with machine learning. Phishing websites have a short life cycle and are updated and iterated quickly. Compared with list-based methods, this method does not require frequent updates and maintenance of blacklists or whitelists, and can detect phishing websites that have not yet appeared.

[0029] (2) The phishing website detection method proposed in the present invention extracts features from the URL without waiting for the page to load and without obtaining features from a third party. Compared with the method based on visual similarity and content, the method requires less time and can be applied in a real-time environment.

[0030] (3) The present invention extracts component features and language features from URLs. Component features include not only an overall analysis of the URL but also an analysis of each component. Language features can further distinguish between legitimate and phishing URLs based on the analysis of URL components, allowing the features used to more comprehensively reflect the differences between URL components and languages. Compared with traditional URL-based methods, the model of this method is more lightweight, does not rely on third-party services, and has more diverse features, greatly improving the model's generalization ability and better meeting the complex and changing characteristics of real-world network environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Detection flow chart for phishing websites;

[0032] Figure 2 An infographic of URL structure;

[0033] Figure 3 Break down the flow chart into a text. DETAILED DESCRIPTION

[0034] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0035] Example: The present invention provides a phishing website detection method based on URL multi-angle features, and its phishing website detection process is as follows: Figure 1 As shown, the following steps are included:

[0036] Step (1) capture the actual URL of the target website as the URL to be tested;

[0037] In one embodiment of the present invention, due to URL shortening services and redirection, the URL accessed by the user may be inconsistent with the URL of the final target website. In this method, the final URL of the target website is used as the URL to be tested. Table 1 shows the initial and final states of the URL of an example webpage, which uses the shortening service.

[0038] Table 1: URL Examples

[0039] URL Initial URL http: / / bit.ly / SHPee_COins URL to be tested https: / / bonussh0peecoin.blogspot.com / 2021 / 08 / shopee-fortune-box.html?m=1

[0040] Step (2) decomposes the URL to be tested to obtain the URL's scheme, authority, path, parameters, query, anchor and other components. Figure 2 It is the structural information of the URL;

[0041] In one embodiment of the present invention, a complete URL consists of several components, of which the scheme is usually referred to as the protocol portion, and the authority consists of four subcomponents: username, password, host name, and port number. Among all the components, the protocol and host name are required, the path and query are optional, and the remaining components are classified as uncommon components. The decomposition results of the URL are shown in Table 2:

[0042] Table 2: URL and component information

[0043] result URL https: / / bonussh0peecoin.blogspot.com / 2021 / 08 / shopee-fortune-box.html?m=1 plan https username None password None Hostname bonussh0peecoin.blogspot.com Port number None path / 2021 / 08 / shopee-fortune-box.html parameter None Query m=1 Anchor None

[0044] Step (3) extracting component features and language features from the URL and its components to obtain URL features;

[0045] In one embodiment of the present invention, the component features extracted from the URL and each component are as follows. The specific features are shown in Table 3:

[0046] (3.1.1) URL component characteristics: URL length, use of HTTPS, number of uncommon components, number of suspicious symbols in the URL, number of numbers in the URL, use of double slashes, and number of URL tokens;

[0047] (3.1.2) Host name (domain) component characteristics: domain name length, number of subdomains, use of IP addresses, number of "-", use of top-level domain names in subdomains, use of shortening services, and number of digits in the domain name;

[0048] (3.1.3) Path component characteristics: path length, path depth, number of suspicious symbols in the path, use of top-level domain names in the path, maximum length of tokens in the path, use of file extensions;

[0049] (3.1.4) Component characteristics of the query: length of the query, number of numbers in the query, number of queries.

[0050] Table 3: URL component characteristics

[0051] feature value feature value url_length 72 shortening_service -1 no_https -1 digit_in_domain 1 suspicious_components_url 0 path_length 32 suspicious_symbols_url 13 path_depth 3 digit_in_url 8 suspicious_symbols_path 3 double_slash_redirecting -1 tld_in_path 1 url_token_count 7 longest_path_token 23 domain_length 28 'having_file_ext 1 subdomain_count 1 query_length 3 having_IP_Address -1 digit_in_query 1 prefix_suffix_in_domain -1 query_count 1 tld_in_subdomain 0

[0052] In one embodiment of the present invention, extracting language features specifically includes the following sub-steps:

[0053] (3.2.1) Build a relative word frequency database. Based on the word frequency database and dynamic programming principles, write a continuous text decomposition algorithm. The purpose is to break the text into several words and ensure the correctness of the words as much as possible. The text decomposition process is as follows: Figure 3 As shown;

[0054] (3.2.2) Based on the N-Gram language model and the Markov chain principle, a text readability detection algorithm is constructed to calculate the ease or difficulty of understanding the text;

[0055] (3.2.3) Split the URL components by the primary delimiter to obtain the intermediate token list of the components. The token list of the components in this embodiment is shown in Table 4. The "." in the domain name is the primary delimiter, the " / " in the path is the primary delimiter, and the "&" and "=" in the query are the primary delimiters;

[0056] Table 4: URL component token list

[0057] Components Token List Hostname ['bonussh0peecoin','blogspot','com'] path ['2021','08','shopee-fortune-box.html'] Query ['m','1']

[0058] (3.2.4) Use the continuous text decomposition algorithm to decompose the intermediate token list of the component. Take the logarithm of the number of segmented words of all tokens and sum them to obtain the text dispersion. Count all segmented words with a length less than 2 to obtain the number of short tokens in the text. The segmented word list of some components is shown in Table 5.

[0059] Table 5: List of URL component segmentation

[0060] Components Token List Hostname ['bonus','sh','0','pee','coin','blogspot','com'] path ['2021','08','shop','ee','fortune','box','html'] Query ['m','1']

[0061] (3.2.5) Use the text readability detection algorithm to calculate the readability of each token of the component. If the readability is greater than a pre-set threshold, it is considered unreadable, otherwise it is considered readable.

[0062] (3.2.6) Among the language features extracted from URL components, the host name's language features include: the number of sensitive words (variations of "www", use of "https"), the number of vowels, the number of unique characters, the discreteness of subdomains, the number of short tokens in the domain name, the readability of the second-level domain, and the maximum readability of tokens in the subdomain. The path's language features include: the discreteness of the path, the number of short tokens in the path, and the number of unreadable tokens in the path. The URL language features for this example are shown in Table 6:

[0063] Table 6: Example URL language characteristics

[0064] feature value sensitive_vocabulary 0 vowel_domain 8 unique_char_domain 7 dispersion_subdomain 2.322 short_token_domain 3 elusive_domain 0 most_elusive_subdomain 2.865 dispersion_path 2.322 short_token_path 2 elusive_token_path 0

[0065] Step (4) inputting the URL features into the machine learning detection model to determine the legitimacy of the URL;

[0066] In one embodiment of the present invention, the trained machine learning algorithms include random forest, XGBOOST, artificial neural network, decision tree, gradient boosting tree, AdaBoost, K nearest neighbor, support vector machine, logistic regression, naive Bayes, etc. Random forest is used as the main classifier, URL features are input and the legitimacy of the URL is judged. If the URL to be tested is finally judged as a phishing URL, the user will receive a warning message of the web page.

[0067] The technical means disclosed in the solutions of the present invention are not limited to those disclosed in the above-mentioned embodiments, but also include technical solutions composed of any combination of the above-mentioned technical features. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A phishing website detection method based on URL multi-angle features, characterized in that: The method comprises the following steps: Step (1) Capture the actual URL of the target website as the URL to be tested; Step (2) decomposes the URL to be tested to obtain the URL's scheme, authority, path, parameters, query, and anchor components; Step (3) extracting component features and language features from the URL and its components to obtain URL features; Step (4) inputting URL features into the machine learning detection model to determine the legitimacy of the URL; The component features extracted from the URL and each component in step (3) include: (3.1.1) URL component characteristics: URL length, use of HTTPS, number of uncommon components, number of suspicious symbols in the URL, number of numbers in the URL, use of double slashes, and number of URL tokens; (3.1.2) Hostname component characteristics: domain name length, number of subdomains, use of IP addresses, number of "-" characters, use of top-level domain names in subdomains, use of shortening services, and number of digits in the domain name; (3.1.3) Path component characteristics: path length, path depth, number of suspicious symbols in the path, use of top-level domain names in the path, maximum length of tokens in the path, use of file extensions; (3.1.4) Query component characteristics: query length, number of digits in the query, number of queries; Among them, language features include: the number of sensitive words, the number of vowels, the number of unique characters, the discreteness of subdomains, the number of short tokens in the domain name, the readability of the second-level domain name, and the maximum readability of tokens in the subdomain; the language features of the path include: the discreteness of the path, the number of short tokens in the path, and the number of unreadable tokens in the path.

2. A phishing website detection method based on URL multi-angle features according to claim 1, characterized in that: In the step (1), due to the URL shortening service and redirection, the URL accessed by the user is inconsistent with the URL of the final target website. This method uses the final URL of the target website as the URL to be tested.

3. The method for detecting phishing websites based on multi-angle features of URLs according to claim 1, characterized in that: In step (2), a complete URL contains several components, among which the scheme is usually called the protocol part, and the authority contains four sub-components: username, password, host name, and port number. Among all the components, the protocol and host name are required components, the path and query are optional components, and the remaining components are classified as uncommon components.

4. The method for detecting phishing websites based on multi-angle features of URLs according to claim 1, characterized in that: The step (3) of extracting language features specifically includes the following sub-steps: (3.2.1) Build a relative word frequency database. Based on the word frequency database and dynamic programming principles, develop a continuous text decomposition algorithm to break the text into several words, ensuring the correctness of the words as much as possible. (3.2.2) Based on the N-Gram language model and the Markov chain principle, a text readability detection algorithm is constructed to calculate the ease of understanding of the text; (3.2.3) Split the URL components by the primary delimiter to obtain a list of intermediate tokens. The primary delimiter in domain names is ".", the primary delimiter in paths is " / ", and the primary delimiters in queries are "&" and "=". (3.2.4) Use the continuous text decomposition algorithm to decompose the intermediate token list of the component. Take the logarithm of the number of segmented words of all tokens and sum them to obtain the text dispersion. Count all segmented words with a length less than 2 to obtain the number of short tokens in the text. (3.2.5) Use a text readability detection algorithm to calculate the readability of each token in the component. If the readability is greater than a pre-set threshold, it is considered unreadable; otherwise, it is considered readable. (3.2.6) In the language features extracted from URL components.

5. The method for detecting phishing websites based on multi-angle features of URLs according to claim 1, characterized in that: In step (4), the machine learning algorithms trained include random forest, XGBOOST, artificial neural network, decision tree, gradient boosting tree, AdaBoost, K nearest neighbor, support vector machine, logistic regression, naive Bayes, etc.; the URL features are input into the machine learning classifier to determine the legitimacy of the URL. If it is a phishing URL, a warning is issued to the user, otherwise the user activity proceeds normally.

Citation Information

Patent Citations

  • Method and apparatus for identifying phishing website

    CN105357221A

  • Phishing website detection method based on multi-feature fusion

    CN108777674A

  • A phishing website URL detection method based on depth learning

    CN109101552A

  • A URL-based phishing website detection method

    CN112468501B

  • Phishing website URL detection method and system based on machine learning

    CN112948725A