A method for identifying malicious URLs based on the PageRank algorithm
Through multi-dimensional feature analysis and optimization based on PageRank algorithm and combined with intelligent analysis and judgment models, the problem of insufficient accuracy and efficiency of malicious website recognition in the existing technology is solved, and more efficient, more comprehensive and more reliable malicious website recognition is achieved.
Patent Information
- Application Number
- CN202510243002.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-03
AI Technical Summary
When facing new and hidden malicious websites, the existing technology lacks recognition accuracy and efficiency, and lacks multi-dimensional analysis, resulting in insufficient comprehensive and accurate analysis.
The malicious URL recognition method based on the PageRank algorithm is adopted to realize the identification of malicious URLs through multi-dimensional feature analysis, PageRank algorithm optimization and intelligent analysis model. The specific steps include data collection and organization, basic analysis and judgment ability analysis, feature extraction and PageRank algorithm model construction, content quality evaluation, user behavior data analysis and time attenuation factor calculation, and ultimately constructing an intelligent analysis and judgment model.
It improves the accuracy and timeliness of malicious URL recognition, realizes comprehensive analysis of multi-dimensional features, improves the comprehensiveness and reliability of analysis and judgment, and can better identify disguised and hidden means.
Smart Images

Figure CN119788409B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security protection, and particularly relates to a method for identifying malicious URLs based on the PageRank algorithm. Background Art
[0002] With the rapid development of Internet technology, website security issues have become increasingly prominent and have become important issues that need to be solved urgently. Especially in recent years, the behavior of splitting malicious website traffic has gradually become a new type of attack method widely used by malicious elements. This attack method cleverly mixes normal traffic with malicious traffic, making it difficult to effectively identify and distinguish normal traffic, which may lead to serious consequences such as a significant decline in website performance and leakage of sensitive data. Traditional security protection measures are unable to provide effective protection when facing this new type of attack.
[0003] Conventional malicious behavior traffic detection mainly relies on content restoration or behavior feature detection of traffic response packets (such as HTTP POST), and content restoration or action detection of traffic request packets (such as HTTP GET). However, in the TCP protocol design, the maximum limit of a single TCP payload packet is 1480 bytes, and the content of an HTTP GET packet is usually small, often not exceeding 500 bytes. Therefore, traffic splitting rarely occurs. This makes the previous detection methods inadequate when facing splitting behavior and unable to conduct effective detection and disposal.
[0004] At the same time, the current traffic splitting mechanism is widely used by lawbreakers in illegal fields to evade detection by existing supervision means. These black and gray production traffic splitting behaviors are mixed in normal traffic splitting, bringing great challenges to detection and disposal. Traditional research and judgment methods mostly rely on means such as domain name feature analysis, record information query, and blacklist filtering. Although these methods can identify and intercept malicious URLs to a certain extent, they are often unable to cope when facing advanced threats and new attack methods.
[0005] Specifically, the problems existing in the prior art mainly include:
[0006] Insufficient accuracy: Traditional methods are difficult to comprehensively and accurately identify all types of malicious URLs, especially those malicious URLs that evade detection through camouflage and concealment means.
[0007] Obvious lag: The research and judgment methods relying on blacklists and preset rules have obvious lag and are difficult to cope with the rapid emergence of new malicious URLs.
[0008] Lack of multi-dimensional analysis: Traditional methods often only focus on single or limited features, lacking comprehensive analysis of multi-dimensional features of malicious URLs, resulting in incomplete and inaccurate research and judgment results.
[0009] Poor timeliness: For newly launched web pages, due to the lack of a sufficient number of links and external references, it is difficult for the traditional PageRank algorithm to give them a reasonable ranking, which affects the timeliness of judgment.
[0010] To solve the above problems, the present invention proposes a method for identifying malicious URLs based on the PageRank algorithm. Summary of the Invention
[0011] In view of the problems mentioned in the background art, the present invention proposes a method for identifying malicious URLs based on the PageRank algorithm, which solves the problems of insufficient accuracy and efficiency of traditional malicious URL judgment methods when facing new and hidden malicious URLs, improves the accuracy of malicious URL judgment; enhances the timeliness of judgment; realizes the comprehensive analysis of multi-dimensional features, and improves the comprehensiveness and reliability of judgment.
[0012] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0013] A method for identifying malicious URLs based on the PageRank algorithm realizes the identification of malicious URLs through multi-dimensional feature analysis, PageRank algorithm optimization, and an intelligent judgment model, and specifically includes the following steps:
[0014] S1: Collection and collation of data to be judged;
[0015] S2: Conduct a basic judgment ability analysis on the sorted data to obtain a judgment result;
[0016] S3: Based on the judgment result, conduct feature extraction and construction of the PageRank algorithm model, and calculate the PageRank value;
[0017] S4: Comprehensively collect and analyze the multi-dimensional features of malicious URLs, construct a comprehensive judgment system, and optimize the PageRank algorithm, specifically including:
[0018] S41: Content quality assessment: Use natural language processing technology to deeply analyze the web page content, construct a multi-dimensional content scoring system, and calculate the content quality score;
[0019] S42: User behavior data analysis: Collect and analyze the multi-dimensional behavior data of users on the web page, construct a user behavior score model, and calculate the user behavior score;
[0020] S43: Time decay factor calculation: Calculate the time decay score according to the time information of the web page;
[0021] S5: Construct an intelligent judgment model and visually display the summary result of the judgment.
[0022] Preferably, in S2, the basic judgment ability analysis includes domain name feature analysis, record information query, domain name inclusion search, blacklist-based filtering method, feature matching and machine learning-based method, data mining and deep learning-based method, and heuristic analysis.
[0023] Preferably, in S3, the specific process of calculating the PageRank value is as follows:
[0024] The calculation formula of the PageRank value PR(A) is as follows:
[0025] ,
[0026] where, is the PageRank value of page A, is the PageRank value of the page pointing to page A, is the page 's number of links pointing to other pages, is the PageRank value of the page pointing to page A, is the page 's number of links pointing to other pages, and d is the damping coefficient.
[0027] Preferably, in S41, content quality evaluation: Use natural language processing technology to deeply analyze the web page content, construct a multi-dimensional content scoring system, and the specific process of calculating the content quality score is as follows:
[0028] The multi-dimensional content scoring system includes: text content analysis, originality scoring, information richness scoring, language fluency scoring, and credibility analysis;
[0029] Introduce a weighted comprehensive ranking index to realize the evaluation of the comprehensive quality of the web page, and the specific calculation formula is:
[0030] Final ranking index = α × PR + (1 - α) × content quality score;
[0031] Specifically:
[0032] ,
[0033] where, α is the weight coefficient, PR is the PageRank value of the web page, is the PageRank value of page A, the content quality score is the score calculated based on the NLP analysis result, is the probability of randomly jumping to a page, N is the total number of web pages; d is the damping coefficient; is the sum of the PageRank values passed from all pages pointing to page A, weighted according to the quality and quantity of the links; (1 - d) is the probability that the user continues to click on a link rather than randomly jump; (1 + α) × CQ(A) is the direct contribution of the content quality score to the PageRank value, adjusted by the weight coefficient α; InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; is a page pointing to page A of the PageRank value, is the page The number of links pointing to other pages.
[0034] Preferably, in S42, user behavior data analysis: collecting and analyzing multi-dimensional behavior data of users on the web page, constructing a user behavior score model, and the specific process of calculating the user behavior score is as follows:
[0035] The multi-dimensional behavior data of users on the web page includes: traffic volume, night behavior, association analysis, number of visits, historical tags, and user evaluation;
[0036] By setting a weight coefficient, combining the user behavior score with the PageRank value to generate a ranking metric, and the calculation formula of the ranking metric is as follows:
[0037] CRS(A)=β×PR(A)+(1 - β)×UB(A),
[0038] where CRS(A) is the final ranking metric of page A; β is the weight coefficient; PR(A) is the PageRank value of page A, and UB(A) is the user behavior score of page A, calculated based on user behavior data, specifically as follows:
[0039] ,
[0040] where n is the number of user behavior metrics; is the weight of the i-th user behavior metric, is the score of page A on the i-th user behavior metric.
[0041] Preferably, in S43, time decay factor calculation: According to the time information of the web page, the specific process of calculating the time decay score is as follows:
[0042] ,
[0043] where PR(A) is the PageRank value of page A; is the time decay factor of page A, calculated based on the online time and content update status of page A; is page T i is the reciprocal of the time decay factor; d is the damping coefficient; InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; is the page pointing to page A PageRank value of, is the page The number of links pointing to other pages;
[0044] The calculation formula of the time decay factor τ is:
[0045] τ = e -λ×Δt,
[0046] where λ is the decay rate, e is the base of the natural logarithm, and Δt is the difference between the page online time and the current time.
[0047] Preferably, in S5, the specific process of constructing the intelligent judgment model is as follows:
[0048] Combined with natural language processing technology, machine learning and time decay mechanism, comprehensively consider the characteristics of multi-dimensional feature data, select a classifier, and construct an intelligent judgment model;
[0049] S51: Data preprocessing: Standardize the extracted feature data; clean the collected website address samples to ensure the accuracy of the data set;
[0050] S52: Model training and verification: Divide the experimental data set into a training set and a test set; use the K-fold cross-validation method to evaluate the model;
[0051] S53: Comprehensive judgment: According to the optimized PageRank algorithm, calculate the comprehensive PageRank value of each website address.
[0052] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0053] (1) Improve the accuracy of malicious website identification: The present invention comprehensively constructs a judgment system for malicious websites through multi-dimensional feature analysis, comprehensively using multi-dimensional features such as domain name features, filing information, search engine inclusion status, content quality, user behavior data, and time information.
[0054] The present invention realizes the optimization of the PageRank algorithm: On the basis of the traditional PageRank algorithm, three key elements of content quality evaluation, user behavior data, and time decay factor are introduced to improve the accuracy and timeliness of judgment; it can better identify malicious websites that avoid detection through camouflage and concealment means.
[0055] (2) Enhance timeliness: By introducing a time decay factor, the present invention can dynamically adjust the PageRank value of web pages to reflect their timeliness and importance; by real-time monitoring and analyzing user behavior data, emerging malicious URLs can be promptly discovered and processed; the system can real-time monitor network traffic, issue early warnings and intercept malicious URLs, achieving real-time monitoring and early warning.
[0056] (3) Achieve multi-dimensional comprehensive analysis: The present invention comprehensively utilizes multi-dimensional features such as domain name characteristics, filing information, search engine inclusion status, content quality, user behavior data, and time information to construct a comprehensive judgment system.
[0057] Data visualization and reports: The system provides data visualization tools to intuitively display the judgment results in the form of charts for users to analyze and make decisions; at the same time, the generated judgment reports also provide users with detailed judgment bases, further enhancing the transparency and credibility of the judgment.
[0058] (4) Improve the level of intelligence: The present invention combines natural language processing technology, machine learning, and a time decay mechanism to construct an intelligent judgment model; by continuously optimizing and iterating model parameters, the intelligent level and real-time response ability of the judgment are improved; it can accurately and efficiently identify malicious URLs and achieve a comprehensive and in-depth analysis of malicious URLs.
[0059] Adaptive adjustment: The system can automatically adjust the judgment strategy according to the dynamic changes of the network environment, improving the adaptability and flexibility of the system.
[0060] (5) The present invention can improve the accuracy of malicious URL judgment, ensure the effective identification of new and hidden malicious URLs; enhance the timeliness of judgment to promptly respond to the threats of emerging malicious URLs; achieve comprehensive analysis of multi-dimensional features, improving the comprehensiveness and reliability of judgment; through an intelligent judgment model, improve the efficiency and accuracy of malicious URL processing, providing strong support for network security protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is the data processing flow chart of the present invention;
[0062] Figure 2 is the schematic diagram of the quantity hypothesis relationship structure of the present invention;
[0063] Figure 3 is the schematic diagram of the quality hypothesis relationship structure of the present invention;
[0064] Figure 4 is the logical architecture diagram of malicious URL judgment of the present invention;
[0065] Figure 5 is the schematic diagram of the data experimental processing flow of the present invention;
[0066] Figure 6 is a schematic diagram of the basic judgment ability analysis in the present invention; Figure 5 in the present invention;
[0067] Figure 7 is the business time sequence diagram of the present invention;
[0068] Figure 8 is the present invention Figure 4 schematic diagram of the content quality assessment process in the present invention;
[0069] Figure 9 is the present invention Figure 4 schematic diagram of the user behavior data process in the present invention;
[0070] Figure 10 is the present invention Figure 4 schematic diagram of the time decay factor process in the present invention;
[0071] Figure 11 is the present invention Figure 4 schematic diagram of the filtering method based on the blacklist in the present invention;
[0072] Figure 12 is the present invention Figure 4 schematic diagram of the method based on feature matching and machine learning in the present invention;
[0073] Figure 13 is the present invention Figure 4 schematic diagram of the method based on data mining and deep learning in the present invention;
[0074] Figure 14 is the present invention Figure 4 schematic diagram of the heuristic analysis process in the present invention. Detailed implementation manners
[0075] The following further clarifies the present invention in combination with specific embodiments. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0076] Embodiment 1
[0077] Aiming at the problems of insufficient accuracy, obvious lag, lack of multi-dimensional analysis, and poor timeliness existing in the existing malicious website judgment technology, this embodiment proposes a malicious website recognition method based on the PageRank algorithm; mainly for the recognition and interception of malicious websites. This method can be integrated into network security devices, security software or cloud platforms to provide real-time malicious website detection and protection services for network users and enterprises.
[0078] In this embodiment, the abbreviations and key terms are defined, and the specific noun explanations are as follows:
[0079] Data packet: A packet is the data unit in TCP / IP protocol communication transmission, and is generally also called a "data packet".
[0080] Traffic fragmentation: The Internet protocol allows IP fragmentation. When a data packet is larger than the maximum transmission unit of the link, it can be decomposed into many small enough fragments so that it can be transmitted on it.
[0081] Three-Way Handshake: In the TCP protocol, the two communicating parties will understand the above information through three TCP packets, and on this basis, establish a TCP connection. The exchange process of the three TCP packet segments of the two communicating parties is also the so-called Three-Way Handshake process for establishing a TCP connection.
[0082] Filtering conditions in traffic mirroring: Include inbound rules and outbound rules, which are used to filter the network traffic mirrored in the mirror session.
[0083] IP Header Length: The function of this field is to describe the length of the IP header because there is a variable-length optional part in the IP header. This part occupies 4 bit positions, and the unit is 32 bit (4 bytes), that is, the value of this area = IP header length (unit: bit) / (8×4). Therefore, the longest length of an IP header is "1111", that is, 15×4 = 60 bytes. The minimum length of the IP header is 20 bytes.
[0084] Total Length of IP Packet: The length of the IP packet calculated in bytes (including the header and data), so the maximum length of the IP packet is 65535 bytes.
[0085] PageRank: PageRank, also known as page rank, Google left rank or PageRank, is a technology calculated based on the hyperlinks between web pages. As one of the elements of web page ranking, Google uses it to reflect the relevance and importance of web pages, and it is often used as one of the factors to evaluate the effectiveness of web page optimization in search engine optimization operations.
[0086] The method provided in this embodiment aims to achieve high-precision and high-efficiency malicious website recognition by comprehensively applying multi-dimensional feature analysis, PageRank algorithm optimization, and intelligent judgment models. The specific implementation steps are as follows:
[0087] S1: Collection and collation of data to be judged;
[0088] Data collection includes network traffic data, web page content data, user behavior data, and other relevant data. The specific information is as follows:
[0089] Network traffic data: Capture HTTP / HTTPS request and response packets from network traffic. These packets contain detailed information about users' web page accesses, such as URLs, request headers, response bodies, etc.
[0090] Use traffic mirroring or traffic scraping tools to collect network traffic data in real-time or at regular intervals to ensure the timeliness and integrity of the data.
[0091] Web page content data: Obtain the HTML content, JavaScript code, CSS styles, etc. of the web pages to be analyzed through web crawler technology or API interfaces. Parse and store the web page content for subsequent natural language processing and feature extraction.
[0092] User behavior data: Collect users' access behavior data on web pages, such as page views, access times, access paths, click behaviors, etc.
[0093] Use browser plugins, log analysis tools, or third-party data providers to obtain this behavior data.
[0094] Other relevant data: Collect other data related to web pages, such as domain name registration information, filing information, search engine indexing status, etc. This data can be obtained from channels such as domain name registrars, filing management agencies, or search engines.
[0095] Data cleaning includes data cleaning, data standardization, data storage, data annotation, etc. Specifically:
[0096] Data cleaning: Clean the collected data to remove duplicate, invalid, or abnormal data. Handle missing values, such as using mean filling, median filling, or model-based filling methods.
[0097] Data standardization: Standardize data from different sources and in different formats to ensure data consistency and comparability. For example, unify the date and time format and convert numerical data to the same dimension.
[0098] Data storage: Store the cleaned data in a database or data warehouse for subsequent querying and analysis. Select an appropriate storage solution according to the type and characteristics of the data, such as relational databases, NoSQL databases, or distributed storage systems.
[0099] Data annotation: Label known normal and malicious URLs to provide sample data for subsequent machine learning model training. The annotation work can be done manually or using semi-automatic or fully automatic annotation tools.
[0100] S2: Analyze the basic judgment ability of the sorted data to obtain the judgment result;
[0101] The basic judgment ability analysis includes domain name feature analysis, record information query, domain name inclusion search, blacklist-based filtering method, feature matching and machine learning-based method, data mining and deep learning-based method, and heuristic analysis.
[0102] Domain name feature analysis: Extract features such as the length of the domain name, character composition (such as the proportion of letters, numbers, and special characters), and whether it contains official names or abbreviations. By analyzing features such as the structure, character combination, and length of the domain name, initially screen out potential malicious URLs.
[0103] Record information query: Query the record information of relevant national competent authorities to obtain the basic information of the website, such as the record-holding unit, registration time, record status, and consistency of the record subject. Comprehensively consider the consistency between the record information and the actual content of the website to improve the accuracy of judgment.
[0104] Domain name inclusion search: That is, the search engine inclusion situation. Query whether the website is included by mainstream search engines, as well as the number of included web pages and the inclusion time. By analyzing the inclusion strategy and update cycle of the search engine, judge the legitimacy of the website.
[0105] Blacklist-based filtering method: The blacklist-based filtering method is a simple and effective malicious URL detection means, which relies on a pre-constructed and maintained malicious URL blacklist database. In the process of malicious URL judgment, the blacklist-based filtering method first matches the URL to be judged with the URLs in the blacklist database. The matching algorithm usually uses string matching, such as exact matching or fuzzy matching, to ensure that the URLs in the blacklist can be accurately identified. If the URL to be judged matches a URL in the blacklist, the system will immediately determine that the URL is a malicious URL and take corresponding disposal measures. The detailed process is as Figure 11 shown.
[0106] Method Based on Feature Matching and Machine Learning: The method based on feature matching and machine learning is a means of malicious URL judgment that combines traditional feature matching techniques and modern machine learning algorithms. This method first extracts multi-dimensional features from the web pages to be judged through feature extraction techniques, such as domain name features, content features, behavior features, etc. These features can comprehensively reflect the attributes and behaviors of web pages, providing rich data support for subsequent judgment. On the basis of feature extraction, this method uses machine learning algorithms to learn and analyze the extracted features. Machine learning algorithms automatically discover the correlations and patterns between features from a large amount of data, thereby achieving accurate identification of malicious URLs. Commonly used machine learning algorithms include Support Vector Machine (SVM), Random Forest, Deep Neural Network (DNN), etc. These algorithms can train efficient classification models based on known normal URL and malicious URL samples for predicting the maliciousness of unknown URLs. The detailed process is as Figure 12 shown.
[0107] Method Based on Data Mining and Deep Learning: The method based on data mining and deep learning is a cutting-edge and powerful malicious URL judgment technology that combines the data exploration and pattern discovery capabilities of data mining and the automatic feature learning and complex pattern recognition capabilities of deep learning. This method first extracts multi-dimensional features related to malicious URLs from a large amount of network data through data mining techniques. On the basis of feature extraction, deep learning algorithms are introduced to further explore the deep correlations and complex patterns between features. Deep learning models, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), etc., can automatically learn the internal representation of data without manual feature design. These models transform the original features into high-level abstract features through multi-layer non-linear transformations, thereby achieving accurate identification of malicious URLs. The detailed process is as Figure 13 shown.
[0108] Heuristic Analysis: Heuristic analysis is a reasoning method based on experience and intuition. In the field of malicious URL judgment, it combines expert knowledge and domain rules, and comprehensively analyzes and evaluates web pages by simulating the decision-making process of human analysts. Heuristic analysis does not rely on a large amount of training data and complex mathematical models, but on analysts' understanding of malicious URL features and keen insight into the network environment.
[0109] In heuristic analysis, a series of heuristic rules related to malicious URLs need to be summarized based on experience first. These rules may be based on multiple dimensions such as domain name characteristics, web page content, and user behavior. Then, analysts will apply these rules to the web pages to be judged. By comparing and analyzing, they will determine whether the web pages conform to the characteristics of malicious URLs. Heuristic analysis is characterized by flexibility and strong interpretability. It can quickly adapt to new network environments and malicious URL characteristics. By continuously adjusting and optimizing heuristic rules, it can improve the accuracy and timeliness of judgment. At the same time, since heuristic rules are constructed based on human - understandable logic and common sense, the judgment results are easy to explain and verify. The detailed process is as Figure 14 shown.
[0110] S3: Based on the judgment results, conduct feature extraction and construction of the PageRank algorithm model, and calculate the PageRank value;
[0111] The PageRank algorithm is one of the core algorithms for Google search engines to evaluate the importance of web pages, and its core concept lies in link analysis. This algorithm believes that the more times a web page is cited by other web pages (i.e., the richer the number of incoming links), the higher its importance in the network. In addition, the PageRank algorithm also deeply considers the quality factor of links, that is, not all links contribute equally to the importance of the target web page. Links from high - quality and authoritative web pages play a more significant role in enhancing the importance of the target web page. This detailed discrimination of link quality enables the PageRank algorithm to more accurately evaluate the true value of web pages and avoid misjudgments that may be caused by simply relying on the number of links.
[0112] S31: Calculate the PageRank value;
[0113] Suppose a group of 4 web pages: A, B, C, and D. If all pages only link to A, then the PR (PageRank) value of A will be the sum of the PageRanks of B, C, and D, that is, PR(A)=PR(B)+PR(C)+PR(D).
[0114] Re - assume that B links to A and C, C only links to A, and D links to all the other 3 pages. A page has only one vote in total, so B gives half a vote to each of A and C. By the same logic, only one - third of the votes cast by D are counted towards the PageRank of A, that is .
[0115] In summary, for a page A, the calculation formula for its PageRank value PR(A) can be expressed as follows:
[0116] ,
[0117] Among them, is the PageRank value of page A, is the page pointing to page A 's PageRank value, is the page 's out-degree, that is, the number of links from the page pointing to other pages, is the page pointing to page A 's PageRank value, is the page 's number of links pointing to other pages, d is the damping factor, usually set to be greater than or equal to 0.7, which represents the probability that a user will continue to randomly click on links after arriving at a page. n represents the nth element in a certain sequence corresponding to time t, and here it is an index of a time or sequence.
[0118] As Figure 4 shown, the process schematic of the PageRank logical algorithm (initial score of the page / number of outgoing links = score corresponding to each outgoing link), such as 0.9 / 3 = 0.3; 1 / 2 = 0.5; the corresponding calculation nodes will accumulate the scores as parameters for subsequent calculations, such as 0.5 + 0.3 = 0.8 pointed to in the previous step.
[0119] S4: Comprehensively collect and analyze multi-dimensional features of malicious URLs, build a comprehensive judgment system, and optimize the PageRank algorithm, specifically:
[0120] S41: Content quality assessment: Use natural language processing technology to deeply analyze the web page content, build a multi-dimensional content scoring system, and calculate the content quality score;
[0121] The multi-dimensional content scoring system includes key aspects such as text content analysis, originality scoring, information richness scoring, language fluency scoring, and credibility analysis.
[0122] To enhance the comprehensiveness and accuracy of web page quality assessment in malicious website judgment, the present invention introduces a content quality assessment mechanism on the basis of the classical PageRank algorithm. The traditional PageRank algorithm mainly relies on the link relationship between web pages to evaluate the importance of web pages. However, this method is difficult to fully reflect the true quality of web page content. Especially for malicious websites filled with spam, malicious advertisements or false content, its limitations are particularly obvious. Therefore, the present invention combines Natural Language Processing (NLP) to deeply analyze web page content and constructs a multi-dimensional content scoring system, aiming to more comprehensively evaluate web page quality, thereby improving the accuracy of web page ranking and the effectiveness of malicious website identification. The multi-dimensional content scoring system specifically covers the following key aspects:
[0123] 1. Bad information detection: That is, text content analysis. Using NLP technology to deeply analyze web page text, accurately identify and eliminate potential malicious content or spam, such as false advertisements, common tricks of phishing websites, etc. The existence of these bad information will directly lower the quality score of the web page.
[0124] 2. Originality scoring: Through advanced text duplication checking technology, strictly examine the originality of web page content. For situations with a large amount of plagiarism or content duplication, a lower originality score will be given, thus affecting its overall quality evaluation.
[0125] 3. Information richness scoring: Comprehensively consider the hierarchical structure of the web page, content type and the richness of information volume. For example, whether the content structure is complete, whether detailed and effective background information is provided, whether diverse media resources (such as pictures, videos, charts, etc.) are incorporated. These factors all have an important impact on web page quality.
[0126] 4. Language fluency scoring: The language quality of the web page is directly related to the user experience. The present invention further reflects the content quality of the web page by carefully analyzing the language fluency of web page text, such as grammar accuracy, spelling correctness, smooth expression, etc.
[0127] 5. Credibility analysis: Deeply check whether the web page contains elements that can enhance credibility, such as company background introduction, industrial and commercial registration information, etc. The lack of credibility elements may imply the existence of fraud risks on the web page, thereby reducing its quality score.
[0128] By introducing this multi-dimensional content scoring system, not only the deficiencies of the traditional PageRank algorithm in content quality assessment are made up for, but also the accuracy and effectiveness of malicious website judgment are significantly improved. When combining the content quality analysis results with the PageRank algorithm, a weighted comprehensive ranking index is also introduced. This index combines the PageRank value of the web page and the content quality score, thus realizing the evaluation of the comprehensive quality of the web page.
[0129] The specific calculation formula is: Final ranking index = α × PR + (1 - α) × Content quality score, where α is the weight coefficient and can be flexibly adjusted according to actual application requirements; PR is the PageRank value of the web page, and the content quality score is the score calculated based on the NLP analysis results.
[0130] The detailed formula is as follows:
[0131]
[0132] Among them, is the PageRank value of page A, is the probability of randomly jumping to a page, N is the total number of web pages, and this part is the basic term in the PageRank algorithm, used to simulate the behavior of users randomly clicking; d is the damping coefficient; is the sum of the PageRank values passed from all pages pointing to page A, weighted according to the quality and quantity of the links; (1 - d) is the probability that the user continues to click on a link instead of randomly jumping; (1 + α) × CQ(A) is the direct contribution of the content quality score to the PageRank value, adjusted by the weight coefficient α. The influence of the two factors can be flexibly balanced according to actual application requirements. InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; is the page pointing to page A 's PageRank value, is the page 's number of links pointing to other pages.
[0133] The following is a simplified example showing the PageRank values, content quality scores and their final ranking indicators of five web pages; in this embodiment, α = 0.7 is set.
[0134] Table 1 Schematic table for calculating content quality scores
[0135]
[0136] In Table 1, the final ranking metric reflects the comprehensive quality of a web page; Web page A has a relatively high PageRank value and a good content quality score, and its final ranking metric is 0.72; Web page C has a relatively low content quality score, and although its PageRank value is high, its final ranking metric is low.
[0137] To verify the optimization effect of introducing content quality assessment, the present invention conducted a series of experiments. The experimental results show that, compared with the traditional PageRank algorithm, the comprehensive ranking method combining content quality scoring can significantly improve the accuracy of web page quality assessment, especially when identifying low-quality content and spam websites. The specific effects are as follows: By comprehensively considering the content quality of web pages, the rankings of spam websites (such as false advertising, malware dissemination pages, etc.) are significantly reduced, effectively avoiding the problem that these websites improve their rankings by means of a large number of external links. Websites with strong originality, rich information, and good user experience are also appropriately promoted in the final rankings, avoiding the problem that high-quality websites have lower rankings due to fewer external links or other factors.
[0138] S42: User behavior data analysis: Collect and analyze multi-dimensional behavior data of users on web pages, construct a refined user behavior score model, dynamically adjust the ranking weights of web pages, and calculate user behavior scores;
[0139] The multi-dimensional behavior data of users on web pages include: page views, night-time behavior, association analysis, number of visits, historical tags, and user evaluations, etc. When exploring the optimization path of the search engine ranking algorithm, user behavior data has become a crucial factor. By collecting and analyzing the multi-dimensional behavior data of users on web pages, the popularity and quality of web pages can be evaluated more accurately, further improving the accuracy of the search engine. The present invention proposes an optimization strategy combining user behavior data. By constructing a refined user behavior score model and dynamically adjusting the ranking weights of web pages, more accurate web page quality assessment can be achieved. The introduction of user behavior data not only reflects the popularity of web pages but also can deeply reveal the actual satisfaction and engagement of users with the content of web pages. Specifically, the core indicators of user behavior data include:
[0140] 1. Page view analysis: Record the number of users who visit a certain web page within a unit of time. Web pages with a relatively large number of page views usually have content that is more favored by users and a higher click-through rate.
[0141] 2. Night-time behavior analysis: By analyzing the access behavior of users at night, judge whether the web page has all-weather attractiveness. If a certain web page can still receive a large number of visits at night, it indicates that its content quality is relatively high or highly meets the needs of users.
[0142] 3. Association analysis: Analyze the access behaviors between homologous web pages. For example, after a user visits a certain page, whether they immediately jump to other pages associated with that page.
[0143] 4. Number of accesses analysis: Examine the multiple access behaviors of users to the same web page within a unit of time. This data helps to determine whether users have a high stickiness to the web page content.
[0144] 5. Historical label (analysis): By analyzing the changes in the access volume of users to a web page within different time periods (such as the last 1 day, the last 3 days, the last 7 days), the long-term popularity of the web page can be identified.
[0145] 6. User evaluation (analysis): Collect and analyze the evaluations of users on this web page on public platforms, including comment content, ratings, etc.
[0146] To effectively integrate user behavior data with the traditional PageRank algorithm, this application proposes a new comprehensive ranking mechanism. This mechanism combines the user behavior score and the PageRank value by setting reasonable weight coefficients to generate a new ranking metric. Specifically, the ranking metric can be calculated by the following formula:
[0147] Final ranking metric = β × PageRank value + (1 - β) × user behavior score,
[0148] That is, CRS(A) = β × PR(A) + (1 - β) × UB(A),
[0149] where CRS(A) is the final ranking metric of page A; β is the weight coefficient used to balance the influence of the PageRank value and the user behavior score on the final ranking; usually, the value of β can be adjusted according to the actual application scenario to find the best balance point; PR(A) is the PageRank value of page A, calculated according to the traditional PageRank algorithm; UB(A) is the user behavior score of page A, which can be calculated based on user behavior data, and the specific calculation formula is:
[0150] ,
[0151] where n is the number of user behavior metrics; is the weight of the i-th user behavior metric, used to reflect the importance of this metric in the comprehensive evaluation; is the score or value of page A on the i-th user behavior metric.
[0152] In addition, for those pages that are frequently marked or reported by users as bad URLs, a dynamic analysis method combining the marking times of bad URLs will be adopted to adjust the PageRank value of the web page in a timely manner. For example, if a certain web page frequently appears negative feedback in user behavior data (such as a high number of bad URL markings), the PageRank value of this web page will be reduced.
[0153] To verify the effectiveness of the proposed optimization method, this application has conducted a large number of experiments and verified them through web page data in the actual network. The experimental results show that the PageRank algorithm combined with user behavior data has achieved a significant improvement in ranking accuracy. In the experiment, after combining the user behavior scores, the rankings of web pages with high negative user evaluations and frequently marked as bad URLs have significantly decreased. Through the refined analysis of user behavior data, the quality of web page content can be accurately quantified, avoiding the blind spots of the traditional ranking algorithm that purely relies on the link structure. The experimental result data is as follows in the table:
[0154] Table 2 Schematic table for calculating user behavior scores
[0155]
[0156] The final ranking indicators in Table 2 reflect the web page rankings after comprehensively considering the PageRank value and user behavior scores. The rankings of Web page A and Web page C are relatively high, and their user behavior scores are also high, indicating that these web pages are more secure and reliable and have higher content quality.
[0157] Combining user behavior data: Combine user behavior data with the PageRank value, and generate a new ranking indicator by setting a reasonable weight coefficient β.
[0158] S43: Calculation of time decay factor: Calculate the time decay score according to the time information of the web page;
[0159] The traditional PageRank algorithm relies too much on the link structure between web pages, which makes it difficult for newly launched web pages to obtain a high score in the ranking due to the limited number of links. To overcome this limitation, this embodiment constructs a more dynamic, fair, and time-sensitive scoring mechanism by introducing a time decay factor. Specifically, the calculation content of the time decay factor includes the following aspects:
[0160] 1. Launch time: That is, the time since the web page was launched. Newer web pages have more growth potential.
[0161] 2. Most recent update time: Reflects the update frequency of the web page content. Pages with frequent updates may be more time-sensitive and credible.
[0162] 3. Response time: That is, the web page response efficiency. Good response speed means higher user experience.
[0163] 4. User access time distribution: Analyze the access behavior of users at different time periods to judge the timeliness and relevance of web page content.
[0164] 5. Registration time: The registration time of the domain name. Domains registered earlier may have higher trust.
[0165] 6. TTL (Time To Live): The survival time of the domain name. A longer TTL usually means the website is more stable and healthier.
[0166] The specific process of calculating the time decay factor is as follows:
[0167]
[0168] Among them, PR(A) is the PageRank value of page A; is the time decay factor of page A, calculated based on the online time and content update status of page A; is the reciprocal of the time decay factor of page T i and is used to weight new pages when passing the PageRank value; d is the damping coefficient; InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; is the page pointing to page A 's PageRank value, is the page 's number of links pointing to other pages.
[0169] To introduce the timeliness factor into the PageRank algorithm, the calculation formula of the time decay factor τ is set as follows:
[0170] τ = e^(-λ×Δt),
[0171] Among them, λ is the decay rate, which is set to 0.1 in this embodiment (this value can be adjusted according to the actual situation), Δt is the difference between the online time of the page and the current time (unit: month); e is the base of the natural logarithm, approximately equal to 2.71828. Calculate the corresponding time decay factor τ according to the online time difference of each page.
[0172] By introducing a time decay factor, the PageRank value of a page will be dynamically adjusted according to its online time and content update status; for new pages, by weighting their PageRank values and multiplying by 1 / τ (i.e., the reciprocal of the time decay factor), their scores can be improved. In this way, newer pages can make up for the score disadvantage caused by fewer links. For old pages, their PageRank values are directly multiplied by τ to reflect the possible decline in value over time. As the existence time of the page increases, its importance will gradually decrease under the influence of the time decay factor. The test results are shown in the following table:
[0173] Table 3 Schematic Table for Calculating Time Decay Factor
[0174]
[0175] It can be seen from Table 3 that as the online time of the page increases, the time decay factor τ gradually decreases. For new pages, the τ value is larger, which can effectively improve their initial scores. For old pages that have existed for a long time, over time, the quality and relevance of the content may change. The time decay factor acts as a "filter", gradually reducing the PageRank values of these pages to avoid their scores being too high and affecting the accuracy of the results.
[0176] By introducing a time decay factor, the PageRank value of the web page is dynamically adjusted to reflect its timeliness and importance. PageRank algorithm optimization: Introduce content quality assessment, user behavior data, and time decay factor to optimize the web page ranking mechanism and improve the accuracy of judgment. Function description: Based on the traditional PageRank algorithm, the system introduces content quality assessment, user behavior data, and time decay factor to optimize the web page ranking mechanism and improve the accuracy of judgment. User interaction: Users can view the comprehensive PageRank value of each website through the interface to understand its importance and credibility in the network. UI interface example: The interface shows the comprehensive PageRank value of the website and the contribution degrees of various influencing factors (such as content quality score, user behavior score, etc.), presented in the form of a bar chart or pie chart.
[0177] The numerical values in this embodiment are all normalized example values, used to illustrate the algorithm optimization effect: Content quality score (0 - 1): Reflects the credibility and quality of the web page content. User behavior score (0 - 1): Quantifies the health of user interaction behavior. Time decay factor (dynamic value): Balances the timeliness differences between new and old pages. Weight coefficients (α, β): Adjust the influence of different dimensions on the final ranking.
[0178] S5: Build an intelligent judgment model;
[0179] The system combines natural language processing technology, machine learning, and a time decay mechanism to build an intelligent judgment model for rapid and accurate identification of malicious URLs.
[0180] S51: Data preprocessing: Standardize the extracted feature data to eliminate biases; strictly clean the collected URL samples to ensure the accuracy of the dataset.
[0181] Extract multi-dimensional features according to the specific requirements of the experiment settings to support subsequent model construction.
[0182] S52: Model training and validation: Divide the experimental dataset into a training set and a test set to ensure that the training set contains a sufficient number of normal and malicious URL samples. Use the training set data to train the constructed malicious URL identification model, and continuously adjust the model parameters to optimize its performance.
[0183] Adopt the K-fold cross-validation method to evaluate the stability and generalization ability of the model.
[0184] S53: Comprehensive judgment: Calculate the comprehensive PageRank value of each URL according to the optimized PageRank algorithm.
[0185] Comprehensively consider the characteristics of multi-dimensional feature data, select excellent classifiers such as Support Vector Machine (SVM), Random Forest, or Deep Neural Network (DNN) to build a malicious URL identification model. Take timely and effective disposal measures for confirmed malicious URLs, such as blocking, warning, or deleting.
[0186] User interaction: Users can view the judgment results through the interface, including performance indicators such as the identification rate, false alarm rate, and missed alarm rate of malicious URLs. UI interface example: The interface displays summary information of the judgment results, such as the list of identified malicious URLs and the judgment accuracy rate, and provides a detailed judgment report for users to view.
[0187] Real-time monitoring and early warning function: The system can monitor network traffic in real-time, give early warnings and block malicious URLs to ensure the safety of network users. User interaction: Users can set the early warning threshold through the interface, receive real-time early warning information, and view the interception records. UI interface example: The interface provides early warning setting options, and users can set the early warning threshold according to their needs (such as the number of malicious URLs, traffic volume, etc.). At the same time, display real-time early warning information and interception records in the form of a timeline or a list.
[0188] Data Visualization and Report Function: The system provides data visualization tools to intuitively display the research and judgment results in the form of charts, facilitating users' analysis and decision-making. At the same time, detailed research and judgment reports are generated for users to consult and archive. User Interaction: Users can select different types of visualization charts (such as line charts, bar charts, pie charts, etc.) through the interface to view the detailed data of the research and judgment results. At the same time, the research and judgment reports can be exported in PDF or Excel format. UI Interface Example: A data visualization area is provided on the interface to display charts of the research and judgment results (such as the trend chart of the number of malicious URLs, the distribution chart of the comprehensive PageRank value, etc.). At the same time, a report export button is provided, and users can download the research and judgment report after clicking it.
[0189] In this embodiment, Startup and Configuration: After the user starts the system, they enter the configuration interface to set system parameters (such as data collection frequency, research and judgment model parameters, etc.). After the configuration is completed, the system starts to automatically collect and analyze the characteristics of malicious URLs in network traffic. Multi-dimensional Feature Analysis: Automatically extract multi-dimensional features of the URLs and output the analysis results. PageRank Algorithm Optimization: Calculate the comprehensive PageRank value of each URL according to the optimized PageRank algorithm and display it. Intelligent Research and Judgment: The system combines multi-dimensional features and the PageRank algorithm to conduct intelligent research and judgment on malicious URLs and output the research and judgment results.
[0190] Software and Hardware Environment:
[0191] 1. Hardware Environment: Server Configuration, In the experiment, 3 high-performance servers are used. Each server is equipped with an Intel(R) Xeon(R) Gold 6238R CPU @ 2.20GHz, a 32-core CPU, 256G of memory, a 20M bandwidth, a gigabit Ethernet network card, and 4 NVIDIA GeForce RTX 3080 Ti graphics cards to ensure sufficient computing resources. Network Architecture: The servers are connected to the Internet and connected to the network traffic monitoring points through switches or routers to collect network traffic data in real time.
[0192] 2. Software Environment: Operating System, The servers run the CentOS Linux release 7.4.1708 operating system to provide a stable operating environment. Development Tools and Libraries: Programming languages such as Python and Java are used, combined with natural language processing libraries (such as NLTK, spaCy), machine learning libraries (such as scikit-learn, TensorFlow / Keras), and big data processing tools (such as Hadoop, Spark) for development and implementation; databases such as MySQL or MongoDB are used to store multi-dimensional feature data, research and judgment results, and user behavior data, etc.
[0193] Embodiment 2
[0194] In this embodiment, improvements to multi-dimensional feature analysis:
[0195] 1. Feature extraction method: In addition to the domain name features, filing information, search engine inclusion status and other features mentioned in Embodiment 1, other features can also be introduced, such as IP address geographical location analysis, DNS resolution records, SSL certificate information, etc., to further enrich the basis for judgment.
[0196] 2. Feature weight assignment: When constructing a multi-dimensional feature scoring system, in addition to using the fixed weight assignment method in Embodiment 1 of this embodiment, a dynamic weight assignment method can also be used to automatically adjust the weights of each feature according to the actual application scenario and data distribution.
[0197] Improvements to the optimization of the PageRank algorithm:
[0198] Optimization method: In addition to introducing content quality evaluation, user behavior data and time decay factor, other optimization methods can also be considered, such as using a topic model (such as LDA) to classify the topics of web page content, further refining the ranking basis of web pages. The topic model can help identify the topics of web pages. For the judgment of malicious URLs, it can distinguish malicious URLs with specific topics such as malware downloads and phishing websites.
[0199] PageRank algorithm: Although the PageRank algorithm performs well in web page ranking, other link analysis algorithms can also be considered, such as the HITS (Hyperlink-Induced Topic Search) algorithm, as an alternative. The HITS algorithm evaluates the importance of web pages by calculating the authority and hub of web pages, and may have more advantages in certain specific scenarios.
[0200] Expansion of application scenarios: In this embodiment, it can be applied to the following scenarios:
[0201] 1) E-commerce security: In the field of e-commerce, the present invention can be used to identify and intercept malicious websites to protect users' personal information and property security. For example, during the payment process, the transaction URL is judged to prevent fraud by malicious URLs such as phishing websites.
[0202] 2) Search engine optimization: In the field of search engine optimization (SEO), the present invention can be used to evaluate the true value and credibility of web pages, helping search engines rank web pages more accurately. For example, by introducing content quality evaluation and user behavior data, the ranking algorithm of search engines is optimized to improve the accuracy and relevance of search results.
[0203] 3) Online advertisement filtering: In the field of online advertisements, the present invention can be used to identify and filter malicious advertisements, improving the browsing experience of users.
[0204] 4) Enterprise information security: In the field of enterprise information security, the present invention can be used to monitor and prevent the spread of malicious URLs in the internal network, protecting the sensitive information and business security of enterprises. For example, by deploying the system of the present invention in the enterprise intranet, the access behaviors of internal employees can be monitored and judged in real time to prevent the spread and infection of malicious URLs.
[0205] 5) Internet of Things security: In the field of the Internet of Things, the present invention can be used to identify and prevent attacks on the Internet of Things system by malicious devices or malicious URLs. For example, by analyzing the access patterns and communication data of Internet of Things devices, potential malicious devices or URLs can be discovered and isolated in a timely manner.
[0206] 6) Social media security: In the field of social media, the present invention can be used to identify and filter malicious links or malicious content, protecting the privacy and security of users. For example, by judging the shared links on social media, the spread and harm of malicious links can be prevented.
[0207] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for identifying malicious websites based on PageRank algorithm, characterized by: Malicious URLs can be identified through multi-dimensional feature analysis, PageRank algorithm optimization, and intelligent judgment models. The specific steps include: S1: Collection and organization of data to be studied; S2: Conduct basic research and judgment capability analysis on the sorted data to obtain the research and judgment results; S3: Based on the research results, feature extraction and PageRank algorithm model construction are performed to calculate the PageRank value; S4: Comprehensively collect and analyze the multi-dimensional characteristics of malicious URLs, build a comprehensive research and judgment system, and optimize the PageRank algorithm, including: S41: Content quality assessment: Use natural language processing technology to conduct in-depth analysis of web page content, build a multi-dimensional content scoring system, and calculate content quality scores, specifically: The multi-dimensional content scoring system includes: text content analysis, originality scoring, information richness scoring, language fluency scoring and credibility analysis; The weighted comprehensive ranking index is introduced to evaluate the comprehensive quality of web pages. The specific calculation formula is: Final ranking index = α × PR + (1-α) × content quality score; Specifically: Among them, α is the weight coefficient, PR is the PageRank value of the web page, PR(A) is the PageRank value of page A, and the content quality score is the score calculated based on the NLP analysis results. is the probability of randomly jumping to a page, N is the total number of web pages; d is the damping coefficient; is the sum of the PageRank values passed from all pages pointing to page A, weighted by the quality and quantity of the links; (1-d) is the probability that the user will continue to click on the link instead of randomly jumping; (1+α)×CQ(A) is the direct contribution of the content quality score to the PageRank value, adjusted by the weight coefficient α; InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; PR(T i ) is page T pointing to page A i The PageRank value of C(T i ) is page T i The number of links pointing to other pages; S42: User behavior data analysis: Collect and analyze the multi-dimensional behavior data of users on the web page, build a user behavior score model, and calculate the user behavior score, specifically: The multi-dimensional behavior data of users on web pages includes: visits, nighttime behavior, association analysis, number of visits, historical tags and user reviews; By setting the weight coefficient, the user behavior score is combined with the PageRank value to generate a ranking index. The calculation formula of the ranking index is as follows: CRS(A)=β×PR(A)+(1-β)×UB(A); Among them, CRS(A) is the final ranking indicator of page A; β is the weight coefficient; PR(A) is the PageRank value of page A, and UB(A) is the user behavior score of page A, which is calculated based on user behavior data, as follows: Where n is the number of user behavior indicators; w i is the weight of the i-th user behavior indicator, Mi(A) is the score of page A on the i-th user behavior indicator; S43: Calculation of time decay factor: Calculate the time decay score based on the time information of the web page, specifically: Among them, PR(A) is the PageRank value of page A; τA is the time decay factor of page A, which is calculated based on the online time and content update status of page A; It is page T i is the inverse of the time decay factor; d is the damping coefficient; InLinks(A) is the sum of the PageRank values passed from all pages pointing to page A; PR(T i ) is page T pointing to page A i The PageRank value of C(T i ) is page T i The number of links pointing to other pages; The calculation formula of the time decay factor τ is: τ=e-λ×Δt; Where λ is the decay rate, e is the base of the natural logarithm, and Δt is the difference between the page online time and the current time; S5: Build an intelligent research and judgment model to visualize the research and judgment summary results, specifically: Combining natural language processing technology, machine learning and time decay mechanism, comprehensively considering the characteristics of multi-dimensional feature data, selecting classifiers, and building intelligent judgment models; S51: Data preprocessing: Standardize the extracted feature data; clean the collected URL samples to ensure the accuracy of the data set; S52: Model training and validation: Divide the experimental data set into training set and test set; Use K-fold cross validation method to evaluate the model; S53: Comprehensive analysis: Based on the optimized PageRank algorithm, calculate the comprehensive PageRank value of each URL.
2. The method for identifying malicious websites based on PageRank algorithm according to claim 1, characterized in that: In S2, basic research and judgment capability analysis includes domain name feature analysis, registration information query, domain name collection search, blacklist-based filtering methods, feature matching and machine learning-based methods, data mining and deep learning-based methods, and heuristic analysis.
3. The method for identifying malicious websites based on PageRank algorithm according to claim 1, characterized in that: In S3, the specific process of calculating the PageRank value is: The calculation formula of PageRank value PR(A) is as follows: Among them, PR(A) is the PageRank value of page A, PR(T i ) is page T pointing to page A i The PageRank value of C(T i ) is page T i The number of links pointing to other pages, PR(Tn) is the PageRank value of page Tn pointing to page A, C(Tn) is the number of links from page Tn to other pages, and d is the damping coefficient.
Citation Information
Patent Citations
Malicious webpage discovery method and system based on feature detection
CN108768921A
Malicious website comprehensive evaluation method and system and storage medium
CN115130104A