Anti-phishing webpage detection system and method thereof

By extracting multi-dimensional web page data in the anti-phishing web page detection system and performing risk score analysis, the problems of low identification accuracy and high computational complexity in the prior art are solved, and efficient identification and real-time detection of phishing web pages are achieved.

CN120200832AInactive Publication Date: 2025-06-24HANGZHOU YIJIS DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510541025.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has low recognition accuracy when facing phishing web pages that are rigorously disguised, and the detection methods based on deep learning or complex algorithms have high computational complexity, making it difficult to support real-time detection of large-scale web pages.

Method used

A phishing-proof web detection system is proposed, including web page data crawling module, risk analysis module, malicious web page judgment module and response decision-making module. By extracting URL information, domain name information and web page content information, conducting multi-dimensional data analysis and behavioral judgment, calculating various risk scores and conducting batch analysis, determining the web page risk level and taking corresponding response measures.

Benefits of technology

It realizes high-precision identification of phishing web pages with camouflaged strict phishing, supports real-time detection of large-scale web pages, reduces the possibility of user victimization, and improves the timeliness and effectiveness of defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200832A_ABST
    Figure CN120200832A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-phishing webpage detection system and method. The anti-phishing webpage detection system comprises a webpage data capture module, a risk analysis module, a malicious webpage judgment module and a response decision module, relates to the technical field of anti-phishing webpage detection. According to the method, multi-dimensional data, including URL information, domain name information and webpage content information, of a webpage are extracted through the webpage data capturing module, a comprehensive original data basis is provided, and accurate support is provided for follow-up risk analysis; according to the method, comprehensive risk analysis is carried out by utilizing various webpage features including URL features, domain name features and content features, and potential phishing webpages can be effectively identified by calculating risk scores of the URL, the domain name and the webpage content; according to the method, the web pages are classified according to the comprehensive risk scores, the high-risk web pages are intercepted, the emergency response logs are generated, the low-risk web pages are continuously monitored, and safe browsing of a user is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of phishing web page detection, and particularly to a phishing web page detection system and method thereof. Background Art

[0002] With the popularization of the Internet and the rapid development of e-commerce, phishing attacks have become a serious security threat. Phishing is a malicious attack behavior that deceives users into entering sensitive information by disguising as a legitimate website. Phishing attacks not only threaten the personal information security of users, but also cause huge economic losses to enterprises and organizations.

[0003] Currently, the detection of phishing web pages mainly includes several methods such as blacklist, whitelist, and behavior analysis. The blacklist method relies on a list of known malicious websites, but due to the variability and concealment of phishing web pages, the accuracy and timeliness of this method are limited. For some static web pages or phishing web pages with very strict disguises, effective analysis may not be possible.

[0004] First of all, when facing some phishing web pages with strict disguises, the existing methods have low recognition accuracy and may misjudge or miss phishing web pages.

[0005] Secondly, although some existing detection methods based on deep learning or complex algorithms have high accuracy, their computational complexity is large, they call more network resources for calculation, and there is a lag in the detection and defense of phishing websites, making it difficult to support real-time detection of a large number of web pages.

[0006] In view of the above problems, it is necessary to propose a phishing web page detection system and method thereof. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems existing in the background art and propose a phishing web page detection system and method thereof.

[0008] The purpose of the present invention can be achieved through the following technical solutions: A phishing web page detection system and method thereof.

[0009] In a first aspect, the present invention provides a phishing web page detection system, including a web page data scraping module, a risk analysis module, a malicious web page determination module, and a response decision module.

[0010] The web page data scraping module is responsible for extracting web page related data from all web pages accessed by users, providing raw data for subsequent analysis, and obtaining URL information, web page domain name information, and web page content information. The data sources include web page source code, request parameters, domain name information, and network traffic.

[0011] Extract URL information, extract the protocol type, path, and parameter data of each target web page URL to be detected, and obtain the length len(U) of the web page URL and the number of parameters count(P) contained in the web page URL; decompose the path of the web page URL into multiple levels, and match a path complexity eigenvalue U(D) for the web page URL according to the regular expressions of each level.

[0012] Extract web page domain name information, obtain the target web page domain name L to be detected, and match a legal domain name L0 that is most similar to it, and obtain the length len(L) of the target web page domain name L and the length len(L0) of the legal domain name L0 that is most similar to it; obtain the edit distance Levenshtein(L, L0) between the target web page domain name L and the legal domain name L0 that is most similar to it. For the target web page domain name L, obtain the types of characters that appear in it, calculate the occurrence probability of each character; and calculate the entropy entropy(L) of the target web page domain name L.

[0013] Extract web page content information, obtain the page text content of the target web page to be detected, including keywords, links, advertising pictures and text. Obtain the sensitive words in it through natural language processing technology, and determine their malicious sensitivity levels, and match a corresponding feature weight according to the malicious sensitivity levels; obtain the illegal connections in it through content anomaly detection, and determine their malicious levels, and match a corresponding feature weight according to the malicious sensitivity levels; obtain the unknown scripts contained in it through script detection, and match a preset feature weight; record the sensitive words, illegal connections and unknown scripts as malicious features, number all malicious features, and the number symbol is i, i = 1, 2,..., n; n is the total number of malicious features.

[0014] Send the extracted URL information, web page domain name information and web page content information of each target web page to be detected to the risk analysis module.

[0015] The risk analysis module extracts features from the collected URL information, web page domain name information and web page content information and conducts behavior analysis to determine whether there are potential features of phishing attacks in each target web page to be detected.

[0016] Conduct URL information analysis: Obtain the length len(U) of the web page URL, the number of parameters count(P) contained in the web page URL, and the path complexity eigenvalue U(D) from the web page data scraping module. Through a preset formula Calculate the URL risk score ; where μ1 is the URL information eigenvalue, and α1, α2, and α3 are preset weight factors; Conduct web page domain name information analysis: Obtain the length len(L) and entropy entropy(L) of the target web page domain name L, the edit distance Levenshtein(L, L0) between the target web page domain name L and the most similar legitimate domain name L0, and the length len(L0) of the most similar legitimate domain name L0 from the web page data scraping module. Through a preset formula Calculate the domain name risk score ; where μ2 is the domain name information eigenvalue; where β1 and β2 are preset weight coefficients; where is the similarity between the target web page domain name L and the most similar legitimate domain name L0.

[0017] Perform web page content information analysis: Obtain the malicious feature number i and the feature weight fi matched by each malicious feature from the web page data scraping module. Through a preset formula Calculate the web page content risk score ; where is the preset benchmark web page content risk score.

[0018] Perform batch information analysis, and traverse the calculation processes of the risk score, domain name risk score, and web page content risk score for all target web pages to be detected. Send the calculated URL risk scores of all target web pages, the domain name risk scores and the web page content risk scores to the malicious web page determination module.

[0019] The malicious web page determination module determines malicious web pages based on the analysis results of the risk analysis module, and performs risk assessment according to the determination results to match the corresponding threat levels.

[0020] Through a preset formula Calculate the comprehensive risk score of the target web page ; where , and are respectively the preset standard values of the URL risk score, the domain name risk score, and the web page content risk score.

[0021] As a preferred embodiment of the present invention, match the preset risk level according to the comprehensive risk score of the target web page.

[0022] If the comprehensive risk score of the target web page is greater than the maximum preset threshold , then determine that the target web page is a high-risk web page; If the comprehensive risk score of the target web page is less than the minimum preset threshold , it is determined that the target web page is a low-risk web page; In other cases, it is determined that the target web page is a medium-risk web page.

[0023] Perform batch detection, traverse all target web pages to be detected in the comprehensive risk score calculation and risk level matching process of the target web pages, and obtain the risk levels of all target web pages.

[0024] The response decision module takes corresponding response measures according to the risk levels of each target web page to ensure timely blocking or warning of high-risk phishing web pages and reduce the possibility of user victimization.

[0025] Match corresponding response measures for each target web page according to the preset response rules; For high-risk web pages: Immediately intercept and prompt the user with a warning, and prohibit access to the target web page; and generate an emergency response log for the high-risk web page. The content of the emergency response log includes: the URL, domain name, access time, detected malicious features, response measures executed by the system, and the content and time of the user warning message for subsequent analysis and optimization of the response decision; For medium-risk web pages: Provide a warning prompt to remind the user to access the target web page with caution; For low-risk web pages: Record the network traffic data of the target web page and continuously monitor its network traffic changes.

[0026] In a second aspect, the present invention provides an anti-phishing web page detection method, including the following steps: Step 1: Network data scraping; Extract raw data from all web pages accessed by the user through the web page data scraping module, including: Extract URL information: Obtain URL features including the protocol type, path, parameter data, URL length, number of parameters, and path complexity eigenvalue of the web page.

[0027] Extract domain name information: Obtain the domain name of the target web page and calculate domain name features including the domain name length, edit distance from a legitimate domain name, and entropy value.

[0028] Extract web page content information: Obtain web page text, sensitive words, illegal links, advertisements, and unknown scripts, and calculate and match the corresponding feature weights.

[0029] Step 2: Comprehensive risk analysis; Perform behavior analysis on the scraped web page data through the risk analysis module: URL information analysis: Calculate the risk score of the URL based on the URL length, number of parameters, and path complexity eigenvalue; Domain name information analysis: Calculate the risk score of a domain name based on its length, entropy value, and similarity to legitimate domain names.

[0030] Web page content analysis: Calculate the risk score of web page content based on the weight of malicious features and the baseline risk score.

[0031] Step 3: Batch information analysis; Perform batch analysis on all target web pages to be detected. Traverse all target web pages and calculate the URL risk score, domain name risk score, and web page content risk score for each target web page to be detected.

[0032] Step 4: Malicious web page determination; Perform web page determination based on the calculation results of the URL risk score, domain name risk score, and web page content risk score for each target web page to be detected: Obtain the comprehensive risk score of each target web page through the sum operation of the URL risk score, domain name risk score, and web page content risk score.

[0033] Match the preset risk level according to the comprehensive risk score and determine whether the web page is a high-risk, low-risk, or medium-risk web page.

[0034] Step 5: Anti-phishing risk response; Take corresponding defense measures according to the risk level of each target web page; For high-risk web pages: Immediately intercept and prompt the user with a warning, prohibiting access to the target web page; and generate an emergency response log for the high-risk web page. The content of the emergency response log includes: the URL, domain name, access time, detected malicious features, response measures executed by the system, and the content and time of the user warning message for subsequent analysis and optimization of the response decision; For medium-risk web pages: Provide a warning prompt to remind the user to access the target web page with caution; For low-risk web pages: Record the network traffic data of the target web page and continuously monitor its network traffic changes.

[0035] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention extracts multi-dimensional data of web pages through a web page data scraping module, including URL information, domain name information, and web page content information, providing a comprehensive raw data basis and accurate support for subsequent risk analysis; performs batch analysis on all target web pages to be detected, quickly calculates various risk scores and classifies them, supporting large-scale web page phishing risk detection; 2. The present invention conducts comprehensive risk analysis using various web page features including URL features, domain name features, and content features. By calculating the risk scores of URLs, domain names, and web page content, potential phishing web pages can be effectively identified. 3. The present invention classifies web pages according to the comprehensive risk score, intercepts high-risk web pages and generates emergency response logs, and continuously monitors low-risk web pages to ensure safe browsing for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings: Figure 1 is the system block diagram of the present invention; Figure 2 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to Figure 1 as shown, an anti-phishing web page detection system includes a web page data scraping module, a risk analysis module, a malicious web page determination module, and a response decision module.

[0039] The web page data scraping module is responsible for extracting web page-related data from all web pages accessed by users, providing raw data for subsequent analysis, and obtaining URL information, web page domain name information, and web page content information. The data sources include web page source code, request parameters, domain name information, and network traffic.

[0040] Extract URL information, extract the protocol type, path, and parameter data of each target web page URL to be detected, obtain the length len(U) of the web page URL and the number of parameters count(P) included in the web page URL; decompose the path of the web page URL into multiple levels, and match a path complexity feature value U(D) for the web page URL according to the regular expressions of each level.

[0041] Extract the web page domain name information, obtain the target web page domain name L to be detected, match a legal domain name L0 that is most similar to it, obtain the length len(L) of the target web page domain name L and the length len(L0) of the legal domain name L0 that is most similar to it; obtain the edit distance Levenshtein(L, L0) between the target web page domain name L and the legal domain name L0 that is most similar to it. For the target web page domain name L, obtain the types of characters that appear in it, calculate the occurrence probability of each character; and calculate the entropy entropy(L) of the target web page domain name L.

[0042] Extract the web page content information, obtain the page text content of the target web page to be detected, including keywords, links, advertising pictures and texts. Obtain the sensitive words in it through natural language processing technology, and determine their malicious sensitivity levels, and match a corresponding feature weight according to the malicious sensitivity level; obtain the illegal connections in it through content anomaly detection, and determine their malicious levels, and match a corresponding feature weight according to the malicious sensitivity level; obtain the unknown scripts contained in it through script detection, and match a preset feature weight; record the sensitive words, illegal connections and unknown scripts as malicious features, number all malicious features, and the number symbol is i, i = 1, 2,..., n; n is the total number of malicious features.

[0043] Send the URL information, web page domain name information and web page content information of each target web page to be detected that are extracted to the risk analysis module.

[0044] The risk analysis module extracts features from the collected URL information, web page domain name information and web page content information and conducts behavior analysis to determine whether there are potential features of phishing attacks in each target web page to be detected.

[0045] Conduct URL information analysis: Obtain the length len(U) of the web page URL, the number of parameters count(P) contained in the web page URL and the path complexity eigenvalue U(D) from the web page data scraping module. Through a preset formula Calculate the URL risk score ; where μ1 is the URL information eigenvalue, and α1, α2 and α3 are preset weight factors; Conduct web page domain name information analysis: Obtain the length len(L) and entropy entropy(L) of the target web page domain name L, the edit distance Levenshtein(L, L0) between the target web page domain name L and the legal domain name L0 that is most similar to it, and the length len(L0) of the legal domain name L0 that is most similar to it from the web page data scraping module. Through a preset formula Calculate the domain name risk score ; where μ2 is the eigenvalue of domain name information; where β1 and β2 are preset weight coefficients; where is the similarity between the target web page domain name L and the most similar legitimate domain name L0.

[0046] Perform web page content information analysis: Obtain the malicious feature number i and the feature weight fi matched by each malicious feature from the web page data scraping module, and calculate the web page content risk score through a preset formula ; where is the preset benchmark web page content risk score.

[0047] Perform batch information analysis, and traverse the calculation processes of the risk score, domain name risk score, and web page content risk score for all target web pages to be detected. Send the calculated URL risk scores , domain name risk scores and web page content risk scores of all target web pages to the malicious web page determination module.

[0048] The malicious web page determination module determines malicious web pages based on the analysis results of the risk analysis module, and performs risk assessment according to the determination results to match the corresponding threat levels.

[0049] Calculate the comprehensive risk score of the target web page through a preset formula ; where , and are respectively the preset standard values of the URL risk score, domain name risk score, and web page content risk score.

[0050] Furthermore, match the comprehensive risk score of the target web page to the preset risk level.

[0051] If the comprehensive risk score of the target web page is greater than the maximum preset threshold , then determine that the target web page is a high-risk web page; If the comprehensive risk score of the target web page is less than the minimum preset threshold , then determine that the target web page is a low-risk web page; In other cases, determine that the target web page is a medium-risk web page.

[0052] Perform batch detection, and traverse the calculation of the comprehensive risk score of the target web page and the matching process of the risk level for all target web pages to be detected to obtain the risk levels of all target web pages.

[0053] ​​The response decision-making module takes corresponding response measures according to the risk levels of each target web page to ensure timely blocking or warning of high-risk phishing web pages and reduce the possibility of user victimization.

[0054] Match corresponding response measures for each target web page according to the preset response rules; For high-risk web pages: Immediately intercept and prompt the user with a warning, prohibiting access to the target web page; and generate an emergency response log for the high-risk web page. The content of the emergency response log includes: the URL, domain name, access time, detected malicious features, response measures executed by the system, and the content and time of the user warning message, for subsequent analysis and optimization of the response decision; For medium-risk web pages: Provide a warning prompt to remind the user to access the target web page with caution; For low-risk web pages: Record the network traffic data of the target web page and continuously monitor its network traffic changes.

[0055] Please refer to Figure 2 As shown, a method for detecting phishing web pages includes the following steps: Step 1, network data scraping; Extract raw data from all web pages accessed by the user through the web page data scraping module, including: Extract URL information: Obtain URL features including the protocol type, path, parameter data, URL length, number of parameters, and path complexity eigenvalue of the web page.

[0056] Extract domain name information: Obtain the domain name of the target web page and calculate domain name features including the domain name length, edit distance from a legitimate domain name, and entropy value.

[0057] Extract web page content information: Obtain web page text, sensitive words, illegal links, advertisements, and unknown scripts, and calculate and match the corresponding feature weights.

[0058] Step 2, risk comprehensive analysis; Conduct behavioral analysis on the scraped web page data through the risk analysis module: URL information analysis: Calculate the risk score of the URL based on the length, number of parameters, and path complexity eigenvalue of the URL; Domain name information analysis: Calculate the risk score of the domain name based on the length, entropy value, and similarity to a legitimate domain name of the domain name.

[0059] Web page content analysis: Calculate the risk score of the web page content based on the weight of malicious features and the benchmark risk score.

[0060] Step 3, batch information analysis; Perform batch analysis on all target web pages to be detected, traverse all target web pages, and calculate the URL risk score, domain name risk score, and web page content risk score for each target web page to be detected.

[0061] Step Four: Malicious Web Page Judgment; Perform web page judgment based on the calculation results of the URL risk score, domain name risk score, and web page content risk score for each target web page to be detected: Obtain the comprehensive risk score for each target web page through the sum operation of the URL risk score, domain name risk score, and web page content risk score.

[0062] Based on the comprehensive risk score, match the preset risk level to determine whether the web page is a high-risk, low-risk, or medium-risk web page.

[0063] Step Five: Anti-Phishing Risk Response; Take corresponding defense measures according to the risk levels of each target web page; For high-risk web pages: Immediately intercept and prompt the user with a warning, prohibiting access to the target web page; and generate an emergency response log for the high-risk web page. The content of the emergency response log includes: the URL, domain name, access time, detected malicious features, response measures executed by the system, and the content and time of the user warning message, for subsequent analysis and optimization of the response decision; For medium-risk web pages: Provide a warning prompt to remind the user to access the target web page with caution; For low-risk web pages: Record the network traffic data of the target web page and continuously monitor its network traffic changes.

[0064] It should be understood that the terms "including" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0065] It should also be understood that the terms used in this disclosure specification are only for the purpose of describing specific embodiments and are not intended to limit this disclosure. As used in this disclosure specification and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in this disclosure specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations; The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments only. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. An anti-phishing webpage detection system, comprising a webpage data capture module, a risk analysis module and a malicious webpage determination module, characterized in that ; The web page data crawling module is responsible for extracting web page related data from all web pages visited by users; Obtaining URL features including the protocol type, path, parameter data, URL length, number of parameters and path complexity characteristic value of the web page to obtain URL information; Obtain the domain name of the target web page, calculate the domain name features including the domain name length, the edit distance with the legitimate domain name and the entropy value, and obtain the web page domain name information; Obtain web page text, sensitive words, illegal links, advertisements, and unknown scripts, calculate and match the corresponding feature weights, and obtain web page content information; The risk analysis module extracts features from the collected URL information, webpage domain information, and webpage content information and performs behavioral analysis to determine whether there are potential features of phishing attacks in each target webpage to be detected; The malicious web page determination module determines the malicious web page based on the analysis results of the risk analysis module, and performs risk assessment according to the determination results to match the corresponding threat level; The comprehensive risk score of each target web page is obtained by summing up the URL risk score, domain name risk score and web page content risk score; Based on the comprehensive risk score and matching the preset risk level, the web page is determined to be high-risk, low-risk or general-risk.

2. The anti-phishing webpage detection system according to claim 1, characterized in that: Also included is a response decision module; The response decision module takes corresponding response measures according to the risk level of each target web page to ensure timely blocking or warning of high-risk phishing web pages and reduce the possibility of user harm.

3. The anti-phishing webpage detection system according to claim 1, characterized in that: The specific process of extracting URL information is as follows: Extract the protocol type, path, and parameter data of each target web page URL to be detected, obtain the length len (U) of the web page URL, and the number of parameters count (P) contained in the web page URL; decompose the path of the web page URL into multiple levels, and match a path complexity feature value U (D) for the web page URL according to the regular expressions of each level.

4. The anti-phishing webpage detection system according to claim 1, characterized in that: The specific process of extracting web page domain name information is as follows: Obtain the target webpage domain name L to be detected, and match it with a legal domain name L0 that is most similar to it, and obtain the length len(L) of the target webpage domain name L and the length len(L0) of the legal domain name L0 that is most similar to it; Obtain the edit distance Levenshtein (L, L0) between the target webpage domain name L and the most similar legitimate domain name L0; for the target webpage domain name L, obtain the types of characters appearing therein and calculate the probability of occurrence of each character; and calculate the entropy (L) of the target webpage domain name L.

5. The anti-phishing webpage detection system according to claim 1, characterized in that: The specific process of extracting web page content information is as follows: The page text content of the target web page to be detected is obtained, including keywords, links, advertising images and texts; the sensitive words therein are obtained through natural language processing technology, and their malicious sensitivity is determined, and a corresponding feature weight is matched according to the malicious sensitivity; the illegal links therein are obtained through content anomaly detection, and their malicious degree is determined, and a corresponding feature weight is matched according to the malicious sensitivity; the unknown scripts contained therein are obtained through script detection, and a preset feature weight is matched; sensitive words, illegal links and unknown scripts are recorded as malicious features, and all malicious features are numbered, and the numbering symbol is i, i=1, 2,..., n; n is the total number of malicious features.

6. The anti-phishing webpage detection system according to claim 1, characterized in that: The specific process of extracting features from the collected URL information, webpage domain information, and webpage content information and conducting behavior analysis is as follows: Analyze URL information: The length len (U) of the web page URL, the number of parameters count (P) contained in the web page URL, and the path complexity characteristic value U (D) are obtained from the web page data crawling module; the preset formula is used Calculating URL Risk Score ; Where μ1 is the URL information feature value, and α1, α2 and α3 are preset weight factors; Analyze web page domain name information: The length len(L) and entropy entropy(L) of the target webpage domain name L, the edit distance Levenshtein(L, L0) between the target webpage domain name L and its most similar legal domain name L0, and the length len(L0) of its most similar legal domain name L0 are obtained from the webpage data crawling module; through the preset formula Calculating Domain Risk Score ; Where μ2 is the characteristic value of the domain name information; Where β1 and β2 are the preset weight coefficients; Where is the similarity between the target webpage domain name L and its most similar legitimate domain name L0; Conduct web page content information analysis: Obtain the malicious feature number i and the feature weight fi matched by each malicious feature from the webpage data crawling module, and use the preset formula Calculating Web Content Risk Score ;in Provides risk scores for pre-set baseline web content; Perform batch information analysis, and traverse the risk score, domain name risk score and web page content risk score calculation process for all target web pages to be detected; calculate the URL risk scores of all target web pages , Domain Risk Score and web content risk score The data is sent to the malicious web page determination module for further determination of the malicious web page.

7. The anti-phishing webpage detection system according to claim 6, characterized in that: The specific process of determining a malicious web page is as follows: By preset formula Calculate the overall risk score of the target page ;in , and They are respectively the preset URL risk score standard value, domain name risk score standard value and web page content risk score standard value; Match the target webpage to the preset risk level based on its comprehensive risk score; If the target webpage's comprehensive risk score Greater than the maximum preset threshold , then the target web page is determined to be a high-risk web page; If the target webpage's comprehensive risk score Less than the minimum preset threshold , then the target webpage is determined to be a low-risk webpage; In other cases, the target webpage is determined to be a general risk webpage; Batch detection is performed, and the target webpage comprehensive risk score calculation and risk level matching process is traversed over all target webpages to be detected to obtain the risk levels of all target webpages.

8. The anti-phishing webpage detection system according to claim 2, characterized in that: The specific process of taking corresponding response measures according to the risk level of each target web page is as follows: For high-risk web pages: block them immediately and warn the user, prohibiting access to the target web page; Also, generate emergency response logs about high-risk web pages; The contents of the emergency response log include: the URL, domain name, access time of the intercepted web page, the malicious features detected, the response measures executed by the system, and the content and time of the user warning information, for subsequent analysis and optimization of response decisions; For general risk web pages: provide warning prompts to remind users to be cautious when visiting the target web page; For low-risk web pages: record the network traffic data of the target web page and continuously monitor its network traffic changes.

9. A method for detecting phishing pages, characterized in that: The following steps are involved: Step 1: Network data crawling; The web data crawler module extracts raw data from all web pages visited by users, including: Extract URL information: obtain URL features including the protocol type, path, parameter data, URL length, number of parameters and path complexity feature value of the web page; Extract domain name information: obtain the domain name of the target web page, calculate the domain name features including domain name length, edit distance with the legitimate domain name and entropy value; Extract web page content information: obtain web page text, sensitive words, illegal links, advertisements and unknown scripts, calculate and match the corresponding feature weights; Step 2: Comprehensive risk analysis; Conduct behavioral analysis on captured web page data through the risk analysis module: URL information analysis: Calculate the risk score of the URL based on the URL length, number of parameters and path complexity feature values; Domain name information analysis: Calculate the risk score of the domain name based on the length, entropy value and similarity with the legitimate domain name; Web page content analysis: Calculate the risk score of the web page content based on the weight of the malicious features and the baseline risk score; Step 3: Batch information analysis; Perform batch analysis on all target web pages to be tested, traverse all target web pages, and calculate the URL risk score, domain name risk score, and web page content risk score of each target web page to be tested; Step 4: Determine malicious web pages; The webpage judgment is made based on the calculation results of the URL risk score, domain name risk score and webpage content risk score of each target webpage to be detected: The comprehensive risk score of each target web page is obtained by summing up the URL risk score, domain name risk score and web page content risk score; According to the comprehensive risk score, the preset risk level is matched to determine whether the web page is a high-risk, low-risk or general-risk web page; Step 5: Anti-phishing risk response; Take appropriate response measures based on the risk level of each target web page to ensure that high-risk phishing pages are blocked or warned in a timely manner to reduce the possibility of users being harmed.

Citation Information

Patent Citations

  • Method and device for detecting phishing website

    CN104077396A

  • Method and device for conducting security detection on network page

    CN104580092A

  • Method and apparatus for identifying phishing website

    CN105357221A

  • Method clustering phishing page to locate target page

    CN105824822A

  • Malicious external-connection flow detection method and device

    CN108718298A

Cited By

  • Website AI intelligent risk assessment real-time avoiding system and avoiding method thereof

    CN120512308A

  • Fraud website detection method and system

    CN120658504A