Deletable weighted learning type bloom filter and malicious URL detection method

By introducing Perfect Hash tables and weighted learning models into the learning Bloom filter, a weighted learning Bloom filter can be designed using the data popularity characteristics of the URL, which solves the problem of significantly increasing false positive rate and query time in malicious URL detection in the existing technology, and achieves more efficient performance improvements after URL detection and data deletion.

CN120238336AActive Publication Date: 2025-07-01HUBEI UNIV OF ARTS & SCI
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510280318.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-01
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing learning Bloom filters fail to fully utilize the URL's data heat differential characteristics in malicious URL detection, resulting in a decline in the overall performance of the model, especially the problem of significantly increasing false positive rate and query time after data deletion.

Method used

A deletable weighted learning Bloom filter (DWLBF) is designed, including a Perfect Hash table (PH table) and a weighted learning model, which reduces false positive rates and query time by storing and processing URLs based on data heat. The PH table is used to store highly popular benign URLs deleted by the database as a preliminary filter; the weighted learning model provides preliminary detection of URL weights based on data popularity; the partition backup Bloom filter is used to further detect URLs that have not been initially judged as malicious.

Benefits of technology

It effectively reduces the false positive rate after URL detection and data deletion, saves data detection time after data deletion, improves the accuracy of the judgment of the learning model, and improves the detection performance after data deletion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238336A_ABST
    Figure CN120238336A_ABST
Patent Text Reader

Abstract

The invention provides a deletable weighted learning type bloom filter and a malicious URL detection method. The deletable weighted learning type bloom filter comprises a PH table, a weighted learning model and a partition backup bloom filter. The PH table is used for storing part of the benign URL deleted by the database according to the data popularity. If the detected URL is in the PH table, indicating that the URL is deleted by the database, and reminding a user that the URL is a malicious URL without further detection; and the weighted learning model endows the URL with a weight for training according to the data popularity. Acquiring a score through a learning model during detection, if the score is greater than a set threshold, indicating that the access is benign and can be normally accessed, otherwise, preliminarily judging that the access is malicious; the partition backup Bloom filter uses a plurality of sub-Bloom filters to store the URL which is judged to be malicious by the learning model for further detection, if the URL is malicious, a user is reminded, and otherwise, normal access can be achieved. According to the method, the FPR of URL detection is effectively reduced, and meanwhile, the detection performance after data deletion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer communication technologies, and particularly to a deletable weighted learning Bloom filter and a malicious URL detection method. Background Art

[0002] A Bloom filter is a technology that uses a bit array and hash functions for data mapping to detect whether data exists in a set. As a data structure with high space efficiency and query time efficiency, it is widely used in malicious URL detection, name lookup, IP address query, database query optimization, etc. In a big data environment, traditional Bloom filters need to use a large bit array for hash mapping, occupying a large amount of memory space. To solve this problem, Kraska et al. proposed a learning Bloom filter at the 2018 SIGMOD conference, combining machine learning with traditional Bloom filters, effectively reducing memory occupancy while ensuring a low false positive rate and improving the overall performance of the filter, which has received extensive attention and research from scholars at home and abroad.

[0003] In malicious URL detection, the learning Bloom filter trains a learning model through features extracted from URLs, such as a random forest classifier, and then uses a traditional Bloom filter as a backup Bloom filter to prevent the generation of false positive data. The learning Bloom filter can more accurately identify malicious URLs by using a machine learning model to classify URLs, effectively reducing memory overhead and providing a new idea for malicious URL detection.

[0004] The existing several Bloom filters are as follows: (1) Counting Bloom Filter Counting Bloom Filter A variant of the traditional Bloom filter, which realizes the data deletion function by replacing each bit in the bit array with a small counter, solving the problem that the traditional Standared Bloomfilter (SBF) cannot delete data. When mapping to k bits using k hash functions, the corresponding counters in the bit array are incremented by 1. If an element is deleted, the counters of the corresponding k bits are decremented by 1.

[0005] Figure 1 Shows the working principle of the CBF. The three data in the figure 、 、 are inserted into the bit array respectively, and it is necessary to accumulate the hash function mapping to the corresponding positions in the bit array, such as Figure 1 the non-zero positions in the bit array shown. When deleting the data , the counters at the mapped hash positions are decremented by 1, as shown in Figure 1As shown by the new bit array. When querying data separately , , During the query, since and All the positions mapped to the bit array are not 0, so it is judged as positive. However, in reality does not exist in the database and is false positive data. At the same time, The mapped position has 0, so it is judged as negative. It can be seen from this that CBF has the general characteristics of a Bloom filter, that is, there is no false negative data, but there is false positive data. In addition, CBF can effectively support data deletion.

[0006] Due to the use of counters, CBF requires more memory space than SBF. When reserving space for the CBF bit array, it is necessary to select an appropriate number of counters to ensure a low false positive rate (False Positive Rate, FPR) while occupying less memory space. In addition, it is also important to select the optimal number k of hash functions. The optimal number of hash functions is: Equation (1) where m represents the number of counters and n represents the number of stored data. The number of times the i-th counter is incremented is: Equation (2) where represents choosing j times from nk hash mappings, represents the probability that the j-th hash mapping selects the i-th counter, represents the probability that the other nk - j hash mappings do not select the i-th counter. Therefore, the probability that the i-th counter increases by more than j times is: Equation (3) Combining Equation (1) and (3) can obtain: Equation (4) If each counter is allocated 4 bits, it will overflow when the counter reaches 16 bits. At this time, according to Equation (4), it can be obtained that , and this probability is small enough. Therefore, a small FPR can be obtained according to the above allocation.

[0007] (2) Learning Bloom Filter Kraska et al. proposed the Learned Bloom Filter (LBF) at the 2018 SIGMOD conference. It combines machine learning with traditional Bloom filters, aiming to improve the accuracy of set membership queries and reduce the FPR. In LBF, when the prediction score of the classifier is higher than a certain heat threshold, the element is directly recognized as a member of the set without further Bloom filter testing. This method not only improves efficiency but also allows LBF to maintain performance in the face of dynamic data sets, especially in data stream applications such as duplicate detection, malicious URL checking, and web caching. As Figure 2 shown, LBF mainly consists of two parts: the first part uses a machine learning model to simulate the traditional BF as the main filter; the second part combines the first part with a traditional Bloom filter (called the backup Bloom filter) to prevent false negatives. When training the learning model, usually a set of data sets in the database is given as positive data, and a set of data sets that do not exist in the database is given as negative data, and are input into the learning model f(x) for training and follow the following rules: Equation (5) The learning model in LBF can adopt Gradient Boosting (GB), Recurrent Neural Network (RNN), or Convolutional Neural Networks (CNN), etc. Then, by minimizing the optimization loss function the best trained model is obtained for judging whether the queried data is in the set. What is obtained is a probability evaluation value, and then it is compared with the size relationship of the set heat threshold and to output the result. If , it is considered that the queried data x is in the database; if , then x is further input into the backup Bloom filter for judgment to obtain the final judgment result. Most data has been filtered by the learning model, so the data stored in the backup Bloom filter is less, thus achieving the purpose of reducing memory consumption. The FPR of LBF can be expressed as: Equation (6) where represents the learning model FPR, Represents the FPR of the backup Bloom filter. The LBF has opened up a new direction for the development of BFs and has great advantages compared to traditional BFs. First, in the big data environment, traditional BFs need to occupy more memory to reduce the FPR, while the LBF uses a machine learning model to judge the data, so that the FPR will not change with the growth of the data volume. Second, the LBF has higher flexibility and can dynamically adjust the model parameters according to the distribution patterns of different data. Third, compared with the traditional BFs that require a fixed bit array, the LBF only needs to store the parameters of the learning model, reducing the occupied memory space.

[0008] (3) Optimized Learned Bloom Filter In recent years, there have been many optimized learned Bloom filters based on the LBF. Mitzenmacher proposed a Sandwiched Learned Bloomfilter (SandwichedLBF) based on the LBF, which reduces the false positive rate by adding an initial Bloom filter at the beginning of the model. Rae et al. proposed a neural Bloom filter, which applies meta-learning to learn an approximate membership set of a set of data and achieves a higher data compression rate than traditional Bloom filters. In addition, some learning methods are committed to partitioning the Bloom filter area and adaptively setting the hash function to improve the model performance. Dai et al. proposed an Adaptive Learned Bloomfilter (Ada-BF), which reduces the false positive rate by using the complete spectrum of the fractional region. Vaidya et al. proposed a Partitioned Learned Bloomfilter (PLBF), which clusters the fractional space into different regions by dynamic programming and assigns different Bloom filters to each region, thus reducing the memory usage while maintaining a low false positive rate. Bhattacharya et al. proposed a Projection HashBloomfilter (PHBF), which effectively reduces the false positive rate and query time consumption by replacing the standard hash function and cooperating with the SBF.

[0009] None of the above methods consider the impact of data heat characteristics on the model performance. On the one hand, the designs of the above methods do not consider the impact of data heat on the learning effect of the learning model. On the other hand, the above models do not fully consider the problem that the hot data will significantly increase the false positive rate and query time after data deletion, thus affecting the overall performance of the model after data deletion.

[0010] (4) Application of Bloom Filter in Malicious URL Detection In malicious URL detection, by storing the known malicious URL features in a Bloom filter, the system can quickly determine whether a given URL is likely to be a malicious URL, thus giving an early warning and blocking before the user accesses it. In addition, the dynamic nature of the Bloom filter allows the system to continuously update the set of malicious URLs over time to cope with the ever-evolving cyber threats. Despite a certain false positive rate, this probabilistic property of the Bloom filter makes it an effective tool for the initial screening of malicious URLs, especially in scenarios with limited resources or the need for quick response.

[0011] In recent years, scholars at home and abroad have proposed various optimized Bloom filters for malicious URL detection. Feng et al. proposed a new type of multi-level counting Bloom filter for network-based URL filtering to improve cache efficiency and reduce bandwidth consumption. Patgiri et al. combined deep learning with the Bloom filter to propose deepBF, designing a two-dimensional Bloom filter to store malicious and benign URLs respectively for malicious URL detection. Dai et al. proposed an adaptive learning Bloom filter to reduce the false positive rate of malicious URL detection by using the complete spectrum of the fractional region. Shuai et al. proposed a fast matching method based on the Bloom filter for the matching of the trusted whitelist, improving the matching efficiency and reducing the consumption of computing resources. Gebretsadik et al. proposed a new type of enhanced Bloom filter to improve the memory efficiency and query performance of Internet of Things intrusion detection. Dinakar et al. proposed a hybrid model combining traditional machine learning methods and deep learning methods, and evaluated the comprehensive performance of shallow (traditional machine learning) and deep (deep learning) learning methods in malicious URL detection, explaining the advantages and limitations of different methods.

[0012] The above methods use learning models for the design of the Bloom filter, effectively improving the performance of the Bloom filter in malicious URL detection. However, these methods do not consider the important feature of the data popularity of URLs and use it to improve the accuracy and efficiency of malicious URL detection.

[0013] The existing learning Bloom filters have the following problems: The initially designed learning model does not support data update, and the update cost after data insertion and deletion is very high; There are too few optimizations in terms of improving the accuracy of the learning model. The learning model is an important part of the learning Bloom filter, but most optimization methods focus on the design of the overall model or the backup Bloom filter, ignoring the impact of the learning model on the overall performance; The impact of data characteristics on model performance is not considered. In addition, the learning Bloom filter does not consider the impact of data popularity, that is, data access frequency, on the performance of the learning model, nor does it consider the problem of the decline in model performance after data deletion.

[0014] In URL access, the data popularity of different URLs varies greatly. The existing learning Bloom filter fails to make full use of this characteristic to improve the overall performance of the filter for the malicious URL detection task. Summary of the Invention

[0015] This application provides a deletable weighted learning Bloom filter, which can solve the technical problem that the existing learning Bloom filter fails to make full use of the data popularity difference characteristic of URLs to improve the overall performance of the filter for the malicious URL detection task.

[0016] In a first aspect, this application provides a deletable weighted learning Bloom filter, including: A PH table for storing some benign URLs deleted by the database according to data popularity, and identifying them as malicious URLs in the PH table after deletion; A weighted learning model for assigning weights to the URL to be detected according to data popularity and performing preliminary detection to obtain a detection score, and obtaining a preliminary detection result according to the detection score. If the detection score is greater than the set threshold, it is identified as a benign URL and can be accessed normally; otherwise, it is input into the partition backup Bloom filter for further detection; The partition backup Bloom filter includes multiple sub-counting Bloom filters, and the multiple sub-counting Bloom filters are used to store URLs that are not initially detected as malicious for re-detection to obtain a detection result.

[0017] In combination with the first aspect, in an implementation, the minimized loss function L of the weighted learning model is:

[0018] In the formula, represents the input item, representing the URL to be detected, is the weight of the input item X, f(x) represents the judgment result output by the weighted learning model, and y is the actual judgment result.

[0019] In a second aspect, this application provides a malicious URL detection method applied to the deletable weighted learning Bloom filter as described above, which is characterized by including the following steps: Obtain the URL to be detected; Use the PH table to filter the URL to be detected, and preliminarily screen out the URLs that have been deleted by the database in the PH table, and identify them as malicious URLs to remind the user; Use a weighted learning model to perform a preliminary detection on the filtered URLs, obtain detection scores, and based on the detection scores, obtain the malicious URLs preliminarily detected and determined and / or the URLs not preliminarily detected and determined as malicious; Obtain the groups of the URLs determined as malicious by the learning model in the partitioned backup Bloom filter according to the detection scores of the learning model, and perform detections in the corresponding groups. If a malicious URL is detected, the user is reminded. If a benign URL is detected, it can be accessed normally.

[0020] Combined with the second aspect, in an implementation manner, the step of using a PH table to filter the URLs to be detected, obtain the filtered URLs and the malicious URLs filtered out specifically includes the following steps: Filter the URLs to be detected through the PH table; If the URL to be detected exists in the PH table, determine that the URL to be detected is a malicious URL; If the URL to be detected does not exist in the PH table, use the URL to be detected as the filtered URL for the next step of detection.

[0021] Combined with the second aspect, in an implementation manner, the step of using a weighted learning model to perform a preliminary detection on the filtered URLs, obtain detection scores, and based on the detection scores, obtain the malicious URLs preliminarily detected and determined and / or the URLs not preliminarily detected and determined as malicious specifically includes the following steps: Train to obtain a weighted learning model based on minimizing the loss function; Detect and obtain the detection scores of each filtered URL based on the trained weighted learning model; Based on the detection scores and the data demarcation boundaries, perform a preliminary detection on the filtered URLs to obtain a preliminary detection result including malicious URLs and / or URLs not directly determined as malicious.

[0022] Combined with the second aspect, in an implementation manner, the step of grouping the URLs not preliminarily detected and determined as malicious according to the detection scores and storing them in multiple sub-count Bloom filters of the partitioned backup Bloom filter for re-detection to obtain a detection result specifically includes the following steps: Group the URLs not preliminarily detected and determined as benign according to the detection scores; Store the groups of the URLs not preliminarily detected and determined as benign in the corresponding sub-Bloom filters; Test each group of URLs with different numbers of independent hash functions to obtain a detection result including malicious URLs and / or non-malicious URLs.

[0023] In combination with the second aspect, in one embodiment, storing the benign URLs that are misidentified as malicious and deleted by the database according to data heat in the PH table specifically includes the following steps: Sort the benign URLs that have been deleted from the database according to data heat; Set the minimum heat threshold of the URLs stored in the PH table; The access frequency is higher than The deleted URLs are stored in the PH table.

[0024] In combination with the second aspect, in one embodiment, the calculation of the data heat is shown in the following formula:

[0025] In the formula, is the historical access times of the input item X, and N represents the total access times of all URLs.

[0026] In combination with the second aspect, in one embodiment, after grouping and correspondingly storing the URLs that are not initially detected as malicious according to the detection scores in multiple sub-counting Bloom filters of the partition backup Bloom filter and performing malicious detection again to obtain the re-detection results, the following steps are further included: Store some of the benign URLs that have been deleted by the database in the PH according to data heat.

[0027] In combination with the second aspect, in one embodiment, storing some of the benign URLs that have been deleted by the database in the PH according to data heat specifically includes the following steps: Sort the malicious URLs according to data heat; Set the minimum heat threshold; Store the benign URLs that have been deleted by the database and whose access frequency is higher than the minimum heat threshold according to the sorting, and are misidentified as malicious URLs after deletion.

[0028] In a third aspect, the present application provides a computer-readable storage medium, on which a malicious URL detection program is stored. When the malicious URL detection program is executed by a processor, the steps of the malicious URL detection method described in the claims are implemented.

[0029] The beneficial effects brought by the technical solutions provided by the embodiments of the present application at least include: The present application proposes a deletable weighted learning Bloom filter for large-scale malicious URL detection problems, effectively reducing the FPR after URL detection and data deletion, and saving the data detection time after data deletion at the same time; A weighted learning model based on data heat is designed to improve the determination accuracy of the learning model; A PH table based on data heat is designed to improve the detection performance after data deletion. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is the working principle diagram of CBF in the prior art; Figure 2 It is the framework diagram of LBF in the prior art; Figure 3 It is the framework diagram of DWLBF provided by the embodiment of the present application; Figure 4 It is the method flow chart of the malicious URL detection method provided by the embodiment of the present application; Figure 5 It is the FPR comparison diagram of DWLBF and five Bloom filters under different data deletion ratios on three malicious URL data sets provided by the embodiment of the present application; Figure 6 It is the detection time comparison diagram of DWLBF and five Bloom filters under different data deletion ratios on three malicious URL data sets provided by the embodiment of the present application; Figure 7 It is the FPR comparison diagram of using and not using the PH table during the detection of three malicious URL data sets provided by the embodiment of the present application; Figure 8 It is the detection time comparison diagram of using and not using the PH table during the detection of three malicious URL data sets provided by the embodiment of the present application; Figure 9 It is the FPR comparison diagram of setting different heat thresholds during the detection of three malicious URL data sets provided by the embodiment of the present application; Figure 10 It is the detection time comparison diagram of setting different heat thresholds during the detection of three malicious URL data sets provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0032] In the description of the specification, claims and the above - mentioned drawings of this application, the terms "comprising", "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. Descriptions such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequential order, nor do they limit that "first", "second" and "third" are different types.

[0033] In the description of the embodiments of this application, words such as "exemplary", "for example" or "for illustration" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for illustration" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0034] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0035] In some processes described in the embodiments of this application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in order or in parallel, and these operations or steps may be combined.

[0036] First, some technical terms in this application are explained to facilitate the understanding of this application by those skilled in the art.

[0037] DWLBF: Deletable Weighted Learned Bloom Filter, a deletable weighted learned Bloom filter; PH: Perfect Hash, perfect hash; FPR: False - Positive Rate, false positive rate; Ada-BF: Adaptive Learned Bloom Filter, that is, an adaptive learning Bloom filter; URL: Uniform Resource Locator, a uniform resource locator.

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0039] In order to support deletion operations and ensure a relatively low overall FPR for the Bloom filter, this application uses a partition backup filter as the backup filter in DWLBF. This filter was initially designed in Ada-BF.

[0040] In a first aspect, as Figure 3 shown, this application provides a deletable weighted learning Bloom filter, which mainly consists of the following three parts: A PH table, which is used to store some benign URLs deleted by the database according to data popularity, reducing the FPR of the model and the detection time consumption of URLs after data deletion; specifically, it stores the benign URLs deleted by the database with a data popularity greater than the popularity threshold. After deletion, they are recognized as malicious URLs in the PH table; the PH table is a special hash table that ensures that for a given set of keys, the hash function maps each key to a unique bucket or slot, and no two keys share the same hash value. In other words, the PH table achieves zero conflicts during the insertion process. Ideally, the query time complexity of the PH table is O(1); A weighted learning model, which is used to assign weights to the URLs to be detected according to data popularity and perform preliminary detection, obtain a detection score, obtain a preliminary detection result based on the detection score, and delete the malicious URLs detected preliminarily; A partition backup Bloom filter, including multiple sub-counting Bloom filters. The multiple sub-counting Bloom filters are used to store the URLs that are not initially detected as benign for re-detection to obtain a detection result. If a malicious URL is detected, the user is reminded. If a benign URL is detected, it can be accessed normally. It has a deletion function, effectively reducing the FPR of model detection.

[0041] The deletable weighted learning Bloom filter provided by this application performs malicious detection on the URLs to be detected by assigning weights to the data according to the characteristics of data popularity, effectively improving the detection accuracy of the learning model.

[0042] The total memory size M occupied by DWLBF in this application is expressed as: Equation (7) Where, , and represent the memory size occupied by the PH table, the memory size occupied by the learning model, and the memory size occupied by the partition backup filter, respectively. Therefore, only the data with the top 25% data heat ( ) in all data is stored to control the total memory usage. In addition, the deleted data stored is all of relatively high heat, which can improve the query performance while maximizing the resource utilization rate.

[0043] In one embodiment, the learning model uses the Gradient Boosting (GB) method, and the optimization objective is to minimize the loss function L: Equation (8) defines as the weight of the input item X, that is, the URL to be detected, and the calculation formula is as follows: Equation (9) where is the historical access count of the th URL, and N represents the total access count of all URLs.

[0044] This application considers the influence of the data heat characteristic, that is, the data access frequency, on the performance of the learning model. Combining Equation (8) and Equation (9), the minimized loss function L of the weighted learning model of this application is: Equation (10) In the formula, represents the input item, representing the URL to be detected, f(x) represents the judgment result output by the weighted learning model, and y is the actual judgment result.

[0045] In a second aspect, based on the same inventive concept, this application provides a URL-related operation method applied to the deletable weighted learning Bloom filter as described above, including operations such as URL detection, deletion from the database, and insertion of URLs. The following is a detailed description of different specific operations: A malicious URL detection method applied to the deletable weighted learning Bloom filter as described above, as Figure 4 shown, includes the following steps: Step S1: Obtain the URL to be detected, that is, the original URL; Step S2: Use the PH table to filter the URL to be detected, and obtain the filtered URL. If the URL exists in the PH table, it is considered that the URL is a malicious URL and can be accessed normally without performing subsequent operations. Otherwise, proceed to the next step; Step S3: Use a weighted learning model to preliminarily detect the filtered URLs, obtain detection scores, and based on the detection scores, obtain the malicious URLs preliminarily detected and determined and / or the URLs not preliminarily detected and determined as malicious; Step S4: According to the detection scores of the learning model, obtain the groups of the URLs determined as malicious by the learning model in the partitioned backup Bloom filter, and perform detection in the corresponding groups. If a malicious URL is detected, the user is reminded. If a benign URL is detected, it can be accessed normally.

[0046] This application gradually detects malicious URLs from the database through PH table filtering, preliminary detection by a weighted learning model, and re-detection by each sub-counting Bloom filter of the backup filter, and ensures the filtering performance when the URL is deleted from the database.

[0047] In one embodiment, the step S1: Obtain the URL to be detected, i.e., the original URL, specifically includes the following steps: Obtain the original URL by querying the perfect hash function.

[0048] In one embodiment, the step S2: Use the PH table to filter the URL to be detected, obtain the filtered URL and the filtered malicious URL, specifically includes the following steps: Step S21: Map the URL to be detected to the PH table through the perfect hash function for filtering; Step S22A: If the URL to be detected exists in the PH table, it indicates that the URL to be detected is a URL that has been deleted from the database and is a malicious URL, and subsequent detection of whether it is a malicious URL is not required; Step S22B: If the URL to be detected does not exist in the PH table, use the URL to be detected as the URL filtered by the PH table for the next step of detection to determine whether it is a malicious URL.

[0049] In one embodiment, the step S3: Use a weighted learning model to preliminarily detect the filtered URLs, obtain detection scores, and based on the detection scores, obtain the malicious URLs preliminarily detected and determined and / or the URLs not preliminarily detected and determined as malicious, specifically includes the following steps: Step S31: Train to obtain a weighted learning model based on minimizing the loss function; Step S32: Based on the trained weighted learning model, detect and obtain the detection scores of each filtered URL ; Step S33: Perform a preliminary detection on the filtered URLs according to the detection scores and the data - defined boundaries, and obtain a preliminary detection result including malicious URLs and / or URLs not directly determined to be malicious; more specifically, let the data - defined boundary be , be the boundaries for the partitioned backup Bloom filter to identify data as negative and positive; if , then directly determine x as a malicious URL, otherwise insert x into the specified sub - counting Bloom filter (Counting Bloom Filter, CBF) to further determine whether it is a malicious URL.

[0050] In one embodiment, the step S4: Group the URLs not preliminarily detected as malicious according to the detection scores and store them in multiple sub - counting Bloom filters of the partitioned backup Bloom filter for re - detection, and obtain the detection result. Specifically, it includes the following steps: Step S41: Group the URLs not preliminarily detected as malicious to obtain grouped URLs; specifically, compare the detection scores of the URLs not preliminarily detected as malicious with the corresponding group score - popularity threshold , and divide the URLs not preliminarily detected as malicious into d groups. Step S42: Store the URLs in each group not preliminarily detected as malicious in the corresponding sub - Bloom filter. Step S43: Test each group of URLs with different numbers of independent hash functions to obtain a detection result including malicious URLs and / or non - malicious URLs; specifically, for the j - th group of URLs, use independent hash functions for testing. Here, the specific value of can be adjusted according to the characteristics of the URLs in each group to control the expected FPR of each group; adjust the parameters, specifically including adjusting the number of hash functions for testing each group of URLs, the score - popularity threshold used for grouping , etc. hyper - parameters to make the expected FPR of different groups more balanced.

[0051] In one embodiment, after the step S4: Group the URLs preliminarily detected as malicious according to the detection scores and store them in multiple sub - counting Bloom filters of the partitioned backup Bloom filter, and perform malicious detection again to obtain the re - detection result, the following steps are further included: Step S5: Store some benign URLs deleted by the database according to the data popularity, and determine them as malicious URLs after deletion.

[0052] In one embodiment, to reduce the FPR and detection time consumption after data deletion, step S5: Store some of the legitimate URLs deleted by the database in PH according to data popularity, which specifically includes the following steps: Step S51: Sort the malicious URLs according to data popularity; Step S52: Set the minimum popularity threshold ; Step S53: Store the legitimate URLs with access frequencies higher than the minimum popularity threshold in the PH table, which are recognized as malicious URLs after being deleted.

[0053] In one embodiment, the calculation of the data popularity in step S51 is shown by the following formula:

[0054] In the formula, is the historical access count of input item X, and N represents the total access count of all URLs.

[0055] The following gives an algorithm pseudocode example of operations such as URL detection, deletion from the database, and insertion of URLs applied to the deletable weighted learning Bloom filter as described above: The algorithm for the URL detection process is as follows:

[0056] The URL detection process described in Algorithm 1 is mainly divided into the following steps: First, use the PH table to test the query data (lines 2 and 3 of the algorithm). If 's hash value is in , it indicates that is a URL that has been deleted from the database, that is, this URL is not in the whitelist, and set the status of to False (lines 4 and 5). Otherwise, input into the weighted learning model (lines 6 and 7); then, use to obtain the score of (line 8). If indicates that the learning model determines that is in the database, set the status of to True. Otherwise, according to , input into the corresponding sub-backup filter (lines 9 - 12). Third, use to test (Line 13), and set according to the result of the status (Line 14). Fourth, return the result of

[0057] The algorithm for the URL deletion process is as follows:

[0058] The process of URL deletion is described in Algorithm 2, which is mainly divided into two steps: First, input into the weighted learning model to obtain the score of (Lines 2 and 3). If , then obtain the popularity ranking value of (Lines 4 - 6). If , it means that the popularity of is above , then re - hash and store in , and set the status of to True (Lines 7 - 10). Otherwise, according to input into the corresponding sub - backup filter ; Second, delete from , and set the status of

[0059] Here, abnormal situations where the deleted URL is not stored in the database or has been deleted before are not considered. In addition, if it is a batch deletion of URLs, the sorting operation in Line 5 of Algorithm 2 is placed at the beginning of the algorithm to save the time overhead of deletion.

[0060] The algorithm for the URL insertion process is as follows:

[0061] Algorithm 3 describes the URL insertion process, which is mainly divided into the following steps: First, use the PH table to test the inserted URL (Lines 2 - 3). If the hash value of is in , it proves that is hot data and has been deleted before. At this time, remove from Input into the weighted learning model (lines 7 - 8); then, use to obtain the score of (line 9). If , it means that the learning model determines that is a benign URL, so directly set to True (lines 10 - 11). Otherwise, according to insert into the corresponding sub-backup filter (lines 12 - 13); third, insert into , and set to True to indicate successful insertion (lines 14 - 15).

[0062] To verify the above effects, 68,909 URLs obtained from Kaggle are used for all experiments here, including benign URLs as positive data and malicious URLs (including malware URLs, phishing URLs, and spam URLs) as negative data. To simulate different retrieval frequencies of users for different URLs, different retrieval times are assigned to all URLs, and the detailed information of all URLs is described in Table 1. In the training phase, all positive samples and 85% of the negative samples are used to train the model, and 10% of the negative samples are used for validation. In the testing phase, all positive data and 5% of the negative samples are used for testing.

[0063] Table 1 Detailed information of URLs used in the experiment

[0064] This application compares the proposed DWLBF method with five Bloom filter methods, and the comparison methods are described as follows: Counting Bloom Filter (CBF): CBF is a traditional Bloom filter structure, which is optimized on the basis of SBF. Each bit of the bit array is extended to a counter so that the data structure can support the deletion of dynamic sets.

[0065] Learning Bloom Filter (LBF): LBF is a new type of Bloom filter that combines a machine learning model with a traditional BF. It uses the predicted probability score to reduce the number of keys that need to be mapped into the bit array, thereby reducing the FPR and memory usage.

[0066] Sandwiched Bloom Filter (SandwichedLBF): SandwichedLBF is a model further optimized based on LBF. Its performance is optimized by adding an initial filter before the learning model.

[0067] Adaptive Bloom Filter (Ada-BF): Ada-BF is an improved LBF. It optimizes performance by adaptively adjusting the parameters of the backup filter in different score regions and using the full score range region. Compared with the LBF method, it has a lower FPR and memory usage.

[0068] Partitioned Learning Bloom Filter (PLBF): PLBF is also an improved LBF. It clusters the score space into different regions through dynamic programming and assigns different BFs to each region, thus reducing memory usage while maintaining a low FPR.

[0069] The experimental comparison was carried out on the Ubuntu 20.04 LTS system with an Intel(R) Xeon(R) CPU E3-1245 v3@3.40GHz processor and 32GB of memory. In addition, Python3 was used as the programming language for all methods.

[0070] Specific comparisons were made in terms of the detection FPR and detection time of the filter when there was no deleted data, the detection performance of the filter after deleting some data, the usage performance of the PH table, the setting of the popularity threshold, etc. in the FPR test and detection time test as follows: I. Comparison of the detection performance of the filter when there is no deleted data: 1.1. FPR comparison: In this part, the FPRs of all methods during the query process were compared, and only the case of non-deleted URLs was considered. The bit array size of the backup filter was set between with an interval of . In addition, for the SandwichedLBF method being compared, there is an initial filter and a backup filter. The bit array size used in the initial filter was set to one-tenth of the total bit array size, and the rest was for the backup filter.

[0071] The FPR comparison results of the detection process on three malicious URL datasets are shown in Tables 2 to 4 respectively. It can be seen from these three tables that at a fixed bit array size, the FPR of the DWLBF method is the lowest, while the FPR of the traditional Bloom filter CBF is very high. On the three URLs, when the bit array size used by the Bloom filter was set to When the value is and , the FPR of CBF can reach 81.0739%. When the size of the bit array is small, the FPR of other comparison methods is relatively high, while the FPR of DWLBF is the lowest. In addition, the FPR of all methods decreases as the size of the bit array increases and then tends to be stable. When the size of the bit array is large, the FPR of DWLBF stabilizes at a relatively small value. It can be clearly seen from the comparison results that the DWLBF method has the smallest FPR when using the same size of the bit array. Especially in Tables 2 and 3, when the size of the bit array is greater than

[0072] Table 2 Comparison results of FPR when no data is deleted on Malware (FPR: %)

[0073] Table 3 Comparison results of FPR when no data is deleted on Phishing (FPR: %)

[0074] Table 4 Comparison results of FPR when no data is deleted on Spam (FPR: %)

[0075] There are several reasons for the above results. First, the traditional Bloom filter CBF uses a bit array to determine whether a URL is positive, so the given size of the bit array has a great impact on the FPR. Second, for the compared learning Bloom filters, SandwichedLBF adds an initial filter, which can reduce the FPR to a certain extent. However, the initial filter must occupy some bits, so it has a relatively high FPR when the memory space is small. Ada-BF and PLBF use partitioning to optimize the partitioned backup Bloom filter, but they ignore the impact of the learning performance of the learning model on the FPR. Third, the proposed DWLBF method uses data popularity for weighting during the model training process, which effectively improves the accuracy of the learning model, so it obtains the lowest FPR compared with other Bloom filters.

[0076] 1.2. Comparison of detection time This application further compares the query time consumption in the case of no data deletion. The size of the bit array is set to 。The comparison results are shown in Table 5. On the one hand, since CBF only needs to confirm that the hash value at the corresponding position in the bit array is not 0, while the learning-based Bloom filter needs to be verified through the combination of the learning model and the backup Bloom filter, which consumes more detection time, so the time consumption of CBF is the lowest. On the other hand, for all learning-based Bloom filters, their time consumptions are very close, and the time consumption of the proposed DWLBF method is not more than that of other learning-based Bloom filters. Therefore, considering the FPR results comprehensively, the DWLBF method is superior to other methods in improving the query performance.

[0077] Table 5 Comparison results of detection time without data deletion (time: s)

[0078] II. Comparison of the detection performance of filters after deleting some data 2.1. FPR comparison This part compares the FPR of filters after deleting different percentages of data on three malicious URL datasets. The percentage of deleted data in the corresponding URL dataset in the database ranges from 5% to 30% with an interval of 5%, and the size of the bit array is set to . Set the data heat threshold to 25%, and all deleted URL data with heat values greater than are stored in the PH table. The experimental results are shown in Figures 5(a) - Figure 7 (c). Obviously, the DWLBF method shows the best performance among all learning-based methods. When the FPR of other learning-based methods gradually increases with the increase of the proportion of deleted URLs, the FPR of DWLBF still remains at a low value.

[0079] Figure 5 In (a), it is the FPR comparison diagram of DWLBF and five Bloom filters under different data deletion ratios on the Malware malicious URL dataset; Figure 5 In (b), it is the FPR comparison diagram of DWLBF and five Bloom filters under different data deletion ratios on the Phishing malicious URL dataset; Figure 5 In (c), it is the FPR comparison diagram of DWLBF and five Bloom filters under different data deletion ratios on the Spam malicious URL dataset.

[0080] There are several reasons for the above results: First of all, DWLBF uses the PH table to store a part of the deleted hot data. When querying this part of the data stored in the PH table, it will be directly filtered out, avoiding being input into the subsequent model and being judged as positive data; Secondly, the PH table stores hot data, and the probability of querying this part of the data is higher than that of other data. Therefore, the increase of FPR can be effectively controlled. Thirdly, although CBF is a traditional Bloom filter, after data deletion, a deletion mark will be set at the corresponding position in the bit array of the hash mapping to keep the false positive rate stable. However, when the data volume is too large and the size of the bit array used is small, its FPR is still large.

[0081] In summary, the DWLBF method proposed in this application can effectively solve the problem of the increase of FPR after data deletion.

[0082] 2.2. Detection time comparison This application further compares the detection time in the case of data deletion. All settings are the same as those in Section D.2.3.2(1). The comparison results are as Figure 6 shown. Since the CBF method is a traditional Bloom filter method, its detection time is still lower than that of the learning-based method after data deletion. For all learning-based methods, before data deletion, the query times of all methods are very close. However, after data deletion, the DWLBF method shows better performance because some of the deleted data is stored in the PH table, and this part of the data does not need to be detected by the learning model and the backup Bloom filter, so the query time is greatly reduced.

[0083] Figure 6 Figure (a) is a comparison chart of the detection time of DWLBF and five Bloom filters under different data deletion ratios on the Malware malicious URL dataset; Figure 6 Figure (b) is a comparison chart of the detection time of DWLBF and five Bloom filters under different data deletion ratios on the Phishing malicious URL dataset; Figure 6 Figure (c) is a comparison chart of the detection time of DWLBF and five Bloom filters under different data deletion ratios on the Spam malicious URL dataset.

[0084] III. Performance test of the PH table 3.1. FPR test This application tests the impact of the PH table on the FPR after data deletion. The experiment is carried out on three malicious URL datasets. The FPR results of URL detection with and without using the PH table after data deletion are compared. The data deletion percentage is between 10% and 30%, with an interval of 10%. The size of the bit array on all URL datasets is set to . The comparison results are as Figure 7As shown, it is obvious that the FPR is effectively reduced after using the PH table. Therefore, the PH table has a good effect on reducing the FPR after data deletion.

[0085] Figure 7 In (a), it is a comparison chart of FPR when using the PH table and not using the PH table after data deletion in the Phishing URL dataset; Figure 7 In (b), it is a comparison chart of FPR when using the PH table and not using the PH table after data deletion in the Malware URL dataset; Figure 7 In (c), it is a comparison chart of FPR when using the PH table and not using the PH table after data deletion in the Spam URL dataset.

[0086] 3.2 Detection Time Test This application tests the influence of the PH table on the detection time after data deletion. All experimental settings are the same as those in Section D.2.3.3(1). The experimental results are as Figure 8 shown. It can be seen from the results that using the PH table can effectively reduce the detection time consumption of the model after data deletion. Combining the comparison results of FPR, using the PH table can effectively improve the detection performance of the model after data deletion.

[0087] Figure 8 In (a), it is a comparison chart of the detection time when using the PH table and not using the PH table after data deletion in the Phishing URL dataset; Figure 8 In (b), it is a comparison chart of the detection time when using the PH table and not using the PH table after data deletion in the Malware URL dataset; Figure 8 In (c), it is a comparison chart of the detection time when using the PH table and not using the PH table after data deletion in the Spam URL dataset.

[0088] IV. Setting of Heat Threshold Setting 4.1 FPR Test This application tests the influence of setting different heat thresholds on the FPR results. The experiment is carried out on three malicious URL datasets. The heat threshold is between 5% and 25%, with an interval of 10%. The data deletion percentage is between 10% and 30%, with an interval of 10%. The bit array size on all URL datasets is set to . The experimental results are as Figure 9 shown. Whether the data deletion percentage is small or large, the FPR decreases as the value increases. This is because When the value increases, the more data with a larger value stored in the PH table, the more the deleted data detected during detection will be detected as negative, so the lower the FPR.

[0089] Figure 9 Among them, (a) is a comparison chart of FPR when setting different thresholds for detecting the Phishing URL dataset; Figure 9 Among them, (b) is a comparison chart of FPR when setting different thresholds for detecting the Malware URL dataset; Figure 9 Among them, (c) is a comparison chart of FPR when setting different thresholds for detecting the Spam URL dataset. 4.2. Detection time test This application tests the impact of setting different heat thresholds on the detection time. The experimental settings are the same as those in part 4.1. The experimental results are as Figure 10 shown. Whether the percentage of deleted data is small or large, the time consumption decreases as the value increases. This is because when the value increases, the more data with a larger value stored in the PH table, the more the deleted data detected during detection will be detected as negative, and the deleted data stored in the PH table does not need to be detected by the learning model anymore, so the detection time is shortened. Considering the results of FPR in 4.1, the setting of the heat threshold has a certain impact on the model performance after data deletion. In addition, the more deleted data stored in the PH table, the greater the memory space occupied. Therefore, only is taken in the experiment.

[0090] Figure 10 Among them, (a) is a comparison chart of detection time when setting different thresholds for detecting the Phishing URL dataset; Figure 10 Among them, (b) is a comparison chart of detection time when setting different thresholds for detecting the Malware URL dataset; Figure 10 Among them, (c) is a comparison chart of detection time when setting different thresholds for detecting the Spam URL dataset.

[0091] In the third aspect, the embodiments of this application provide a deletable weighted learning Bloom filter device for malicious URL detection. The deletable weighted learning Bloom filter device for malicious URL detection can be a device with data processing functions such as a personal computer (PC), a laptop, a server, etc.

[0092] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of components inside the device of the deletable weighted learning Bloom filter for malicious URL detection, and interfaces for implementing the interconnection between the device of the deletable weighted learning Bloom filter for malicious URL detection and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.

[0093] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0094] The processor can be a general-purpose processor, which can call the deletable weighted learning Bloom filter program stored in the memory and execute the deletable weighted learning Bloom filter provided by the embodiments of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the deletable weighted learning Bloom filter program for malicious URL detection is called can refer to the various embodiments of the deletable weighted learning Bloom filter of the present application, which will not be elaborated here.

[0095] Fourthly, the embodiments of the present application also provide a readable storage medium.

[0096] The deletable weighted learning Bloom filter program for malicious URL detection is stored on the readable storage medium of the present application. When the deletable weighted learning Bloom filter program for malicious URL detection is executed by a processor, the steps of the deletable weighted learning Bloom filter as described above are implemented.

[0097] Among them, the method implemented when the deletable weighted learning Bloom filter program for malicious URL detection is executed can refer to the various embodiments of the deletable weighted learning Bloom filter of the present application, which will not be elaborated here.

[0098] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0100] The above are only the preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A deletable weighted learning Bloom filter, characterized in that: include: The PH table is used to store some benign URLs that have been deleted from the database according to data popularity, and are identified as malicious URLs after being deleted; The weighted learning model is used to assign weights to the URLs to be detected according to the data popularity and perform preliminary detection to obtain the detection score. If the detection score is greater than the set threshold, the URL is judged to be a benign URL; if the detection score is less than the set threshold, the URL is preliminarily judged to be a malicious URL; The partition backup Bloom filter includes multiple sub-counting Bloom filters, which are used to store URLs that are initially detected as malicious for re-detection, obtain re-detection results, and send a reminder user instruction if the URL is again determined to be malicious, otherwise normal access rights are granted to the URL.

2. The deletable weighted learning Bloom filter according to claim 1, characterized in that: The minimization loss function L of the weighted learning model is: In the formula, Indicates the input item, representing the URL to be detected. is the weight of the input item X, f(x) represents the judgment result output by the weighted learning model, and y is the actual judgment result.

3. A method for detecting malicious URLs in a removable weighted learning Bloom filter as claimed in any one of claims 1 to 2, characterized in that: The following steps are involved: Get the URL to be tested; Use the PH table to filter the URL to be detected. If the URL exists in the PH table, it is considered a malicious URL and no subsequent operations are required. Otherwise, proceed to the next step. Use the weighted learning model to perform preliminary detection on the filtered URLs and obtain the detection score. If the detection score is greater than the set threshold, the URL is considered to be a benign URL and normal access rights are granted without subsequent operations. Otherwise, proceed to the next step. According to the detection score, the grouping of the URLs judged as malicious by the weighted learning model in the partition backup Bloom filter is obtained, and re-detection is performed in the corresponding grouping. If it is detected as a malicious URL, a reminder user instruction is sent; if it is detected as a benign URL, the URL is granted normal access rights.

4. The malicious URL detection method according to claim 3, characterized in that: The step of obtaining the URL to be detected specifically includes the following steps: Get the URL to be detected by querying the hash function.

5. The malicious URL detection method according to claim 3, characterized in that: The method of using a weighted learning model to perform preliminary detection on the filtered URLs, obtaining detection scores, and obtaining malicious URLs determined by preliminary detection and / or URLs that are not determined to be malicious by preliminary detection based on the detection scores specifically includes the following steps: A weighted learning model is obtained by training based on minimizing the loss function; The detection score of each URL after filtering is obtained based on the weighted learning model detection obtained through training; Perform preliminary detection on the filtered URLs based on the detection scores and data demarcation boundaries to obtain preliminary detection results including malicious URLs and / or URLs that are not directly determined to be malicious.

6. The malicious URL detection method according to claim 3, characterized in that: The method of grouping the URLs that are not initially detected as malicious according to the detection scores and performing re-detection in a plurality of sub-counting Bloom filters correspondingly stored in the partition backup Bloom filter to obtain the detection results specifically includes the following steps: Store the URL group URLs that are not initially detected as malicious in the corresponding sub-Bloom filter; Each group of URLs is tested using a different number of independent hash functions to obtain detection results including malicious URLs and / or non-malicious URLs.

7. The malicious URL detection method according to claim 3, characterized in that: The detection scores of the URLs that are not determined to be malicious by the preliminary detection are grouped and stored in a plurality of sub-counting Bloom filters of the partition backup Bloom filter, and malicious detection is performed again. After obtaining the re-detection results, the following steps are also included: Benign URLs that are deleted from the database are stored based on data popularity and are identified as malicious URLs after being deleted.

8. The malicious URL detection method according to claim 7, characterized in that: The malicious URLs that are deleted from the database according to the data heat storage part are identified as malicious URLs after being deleted, specifically including the following steps: Sort malicious URLs based on data popularity; Set the minimum heat threshold; Benign URLs deleted by the database with access frequencies higher than the minimum heat threshold are stored in a sorted manner. At this time, the URL is identified as a malicious URL.

9. The malicious URL detection method according to claim 7, characterized in that: The data heat The calculation of is as follows: In the formula, is the historical visit count of input item X, and N is the total visit count of all URLs.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a malicious URL detection program, wherein when the malicious URL detection program is executed by a processor, the steps of the malicious URL detection method according to any one of claims 3 to 9 are implemented.

Citation Information

Patent Citations

  • Threshold determination model determining method, device, medical detection equipment and storage medium

    CN109599178A

  • Hybrid phishing website detection method and device, electronic equipment and storage medium

    CN114070653A

  • Multi-layer malicious URL identification method based on learning type bloom filter

    CN115941327A

  • Website processing method and device, computing equipment and storage medium

    CN117493623A

  • Data reading method and device, electronic equipment, medium and product

    CN118567561A