A method and system for discovering malicious URL communities based on URL fission
By extracting and analyzing the multi-dimensional features of malicious URLs, dynamically constructing the URL relationship diagram and using clustering algorithms, the problem of difficult to identify malicious URL communities in the existing technology is solved, and more efficient and accurate malicious URL recognition is achieved.
Patent Information
- Application Number
- CN202510272467.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
It is difficult for the existing technology to accurately mine the malicious URL community from a single malicious URL, and it is difficult to accurately identify the emerging malicious URL patterns when facing the rapid changes and hidden propagation of malicious URLs.
By obtaining the list of known URLs, extracting domain name features, IP address features and web page content features, calculating the feature similarity between unmarked URLs and tagged URLs, dynamically constructing a URL relationship diagram, and using the Gilvan-Newman algorithm for clustering to identify malicious URL communities.
It improves the accuracy of malicious URL recognition, can effectively reveal the relationship between URLs, identify closely related malicious URL communities, and improves the efficiency of network security protection.
Smart Images

Figure CN119766581B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method and system for discovering malicious website groups based on website fission, belonging to the technical field of network security. Background Art
[0002] With the rapid development of the Internet, the security situation in cyberspace has become increasingly complex and severe. Malicious URLs, as important tools for cyber attacks and criminal activities, are showing a trend of diversification, concealment and scale. These malicious URLs may be used to spread malware, steal user privacy information, commit cyber fraud and other illegal and criminal acts, seriously threatening the information security and property security of individuals, enterprises and even the country. Malicious URL communities are usually composed of multiple interrelated malicious URLs. They evade security detection and protection by sharing resources, coordinated attacks and rapid replacement, which brings great challenges to network security protection. Therefore, how to accurately mine the associated malicious URL community from a single malicious URL has become a key issue that needs to be solved in the current field of network security.
[0003] At present, there are many types of detection and discovery methods for malicious URLs, including filtering methods based on blacklist libraries, association analysis methods based on graph mining, and classification methods based on machine learning. Among them:
[0004] The filtering method based on the blacklist library maintains a blacklist library containing known malicious URLs to match and filter the URLs visited by users. However, most of these methods have a lag because fraudsters can quickly generate a large number of new malicious URLs through means such as domain name generation algorithms, thereby bypassing the filtering of the blacklist library.
[0005] The association analysis method based on graph mining abstracts network entities such as URLs or IP addresses into nodes in the graph, and the relationships between entities into edges, and uses graph mining algorithms to discover associations in the network. However, the results of graph mining algorithms often lack interpretability, making it difficult to intuitively understand the specific formation mechanism of malicious URL communities.
[0006] The classification method based on machine learning extracts the features of the URL and classifies the URL using a machine learning algorithm to identify potential malicious URLs. For example, the Chinese invention patent "A method and system for detecting fraud in a large number of URLs" with publication number CN117596003A discloses a method and system for detecting fraud in a large number of URLs, which relates to the field of URL detection technology and is applied to a URL detection server. The URL detection server is deployed on an elastic container instance and includes the following steps: obtaining a large number of URLs that have not been detected from a message middleware device; determining the relevant information generated after the URL is accessed through headless browser technology; inputting the relevant information into a trained machine learning model, generating a detection result through the trained machine learning model; and returning the detection result to the message middleware device. The above invention deploys the URL detection service on an elastic container instance. When a large number of URLs that have not been detected need to be detected, multiple elastic containers can be opened to detect these URLs at the same time, which greatly improves the detection efficiency. It can not only detect a large number of URLs in a short time, but also greatly improve the accuracy of judgment. However, the above invention mainly focuses on detecting whether a single URL is fraudulent, without paying attention to the relationship between URLs, and cannot dig out the associated malicious URL community from a single malicious URL. Moreover, the above invention relies on machine learning models for detection. Like traditional machine learning methods, it is difficult to accurately identify the emerging malicious URL patterns when facing the rapid changes and hidden propagation of malicious URLs. Moreover, the above invention does not take into account the dynamic characteristics of malicious URL communities, such as coordinated attacks and resource sharing between URLs, and therefore cannot effectively deal with the security threats brought by malicious URL communities. Summary of the invention
[0007] In order to solve the above problems existing in the prior art, the present invention proposes a method and system for discovering malicious URL communities based on URL fission.
[0008] The technical solution of the present invention is as follows:
[0009] On the one hand, the present invention provides a method for discovering a malicious URL community based on URL fission, the method comprising:
[0010] Obtain a list of known URLs and mark a list of malicious URLs in the list of known URLs;
[0011] Extract features from the known URL list to obtain domain name features, IP address features, and web page content features, calculate the feature similarity of unmarked URLs and marked URLs in the known URL list, and calculate the comprehensive similarity of unmarked URLs based on the feature similarity. If the comprehensive similarity of unmarked URLs exceeds a preset similarity threshold, the unmarked URLs are regarded as candidate malicious URLs and added to the candidate malicious URL set.
[0012] Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights, and the Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities;
[0013] Calculate the nearest community distance of each node in the URL relationship graph, calculate the URL comprehensive score based on the comprehensive similarity and the nearest community distance, remove the candidate malicious URLs whose URL comprehensive scores are lower than the preset comprehensive threshold from the candidate malicious URL set, and obtain the final candidate malicious URL set and the corresponding URL relationship graph.
[0014] As a preferred embodiment of the present invention, the method further includes performing data preprocessing on the known website list before extracting features from the known website list, including data cleaning, URL standardization and data denoising.
[0015] As a preferred implementation of the present invention, the feature extraction of the known website list to obtain domain name features, IP address features and web page content features is specifically as follows:
[0016] Divide domain names at all levels based on known URL information to obtain domain name features;
[0017] The domain name feature is parsed to obtain the corresponding IP address, and the longitude and latitude corresponding to the IP address are queried using the GeoLite2 database to obtain the IP address feature;
[0018] The web page content features include text features and image features. The text and image data in the web page corresponding to the known URL are captured, and the text data is vectorized using the large language model GPT-3 to obtain text features. The image data is feature extracted using an autoencoder to obtain image features.
[0019] As a preferred implementation of the present invention, the calculation of the feature similarity between the unmarked URLs and the marked URLs in the known URL list is specifically as follows:
[0020] The edit distance between the domain name features of the unlabeled URL and the domain name features of the labeled URL is calculated by the Damerau-Levenshtein distance, and normalized to be expressed as follows:
[0021] ;
[0022] In the formula, is the normalized edit distance; To mark the domain name corresponding to the URL; The domain name corresponding to the unmarked URL; For domain name and domain name The Damerau-Levenshtein distance is calculated for the string within. and Domain names and domain name The length of the corresponding string; To perform the maximum value operation;
[0023] in, Calculated by dynamic programming algorithm, specifically:
[0024] ;
[0025] ;
[0026] In the formula, To perform the minimum value operation; Domain name Before characters; Domain name Before characters; is the indicator function; Domain name Middle Bit character; Domain name Middle Bit character;
[0027] from , Start to increase and , and calculate the values of , until , ,get ;
[0028] The semantic similarity between the domain name features of the unlabeled URL and the domain name features of the labeled URL is calculated and expressed as:
[0029] ;
[0030] In the formula, and From domain name and domain name The extracted vocabulary set; Domain name and domain name The semantic similarity of and Domain names and domain name Index of vocabulary; Domain name The number of words in the vocabulary set;
[0031] The domain name feature similarity is calculated based on the normalized edit distance and semantic similarity, and is expressed as follows:
[0032] ;
[0033] is the domain name feature similarity; is the proportionality coefficient;
[0034] The distance between the IP address features of the unmarked URL and the IP address features of the marked URL is calculated by the Haversine formula, and normalized to obtain the IP address feature similarity, which is expressed as:
[0035] ;
[0036] ;
[0037] In the formula, is the IP address feature similarity; is the intermediate parameter; and Respectively represent the longitude and latitude of the IP address of the marked URL; and Respectively represent the longitude and latitude of the unmarked URL IP address;
[0038] Calculate the cosine similarity between the webpage content features of the unmarked URL and the webpage content features of the marked URL, including the cosine similarity corresponding to the text features and the cosine similarity corresponding to the image features. Normalize the cosine similarity of the webpage content features to obtain the webpage content feature similarity, which is expressed as:
[0039] ;
[0040] In the formula, is the similarity of web page content features; and The text features representing the labeled URLs and the unlabeled URLs respectively; and Represent the image features of labeled URLs and unlabeled URLs respectively.
[0041] As a preferred implementation of the present invention, the comprehensive similarity of unlabeled URLs is calculated based on the feature similarity and is expressed as:
[0042] ;
[0043] In the formula, is the comprehensive similarity; , and They are the weights of domain name feature similarity, IP address feature similarity and web page content feature similarity respectively.
[0044] As a preferred implementation of the present invention, the Girvan-Newman algorithm is used to cluster the URL relationship graph as follows:
[0045] Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as:
[0046] ;
[0047] In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths;
[0048] The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stopping condition is reached, so as to obtain a URL relationship graph divided into multiple subgraphs, wherein the stopping condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community.
[0049] As a preferred embodiment of the present invention, the nearest community distance of each node in the URL relationship graph is calculated. , expressed as:
[0050] ;
[0051] In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated;
[0052] The comprehensive score of the URL is calculated based on the comprehensive similarity and the nearest community distance, expressed as:
[0053] ;
[0054] In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively.
[0055] On the other hand, the present invention also provides a malicious URL community discovery system based on URL fission, the system includes a data collection module, a feature extraction module, a URL fission module, a result filtering module and an output module, wherein:
[0056] The data collection module is used to obtain a list of known web addresses and mark a list of malicious web addresses in the list of known web addresses;
[0057] The feature extraction module is used to extract features from a known website list to obtain domain name features, IP address features and web page content features;
[0058] The URL fission module is used to calculate the feature similarity between the unmarked URLs and the marked URLs in the known URL list, and calculate the comprehensive similarity of the unmarked URLs based on the feature similarity, which is expressed as:
[0059] ;
[0060] In the formula, is the comprehensive similarity; , and Domain name feature similarity , IP address feature similarity Similarity with web page content features The weight of the domain name feature similarity is calculated based on the normalized edit distance and semantic similarity; the distance between the IP address feature of the unmarked URL and the IP address feature of the marked URL is calculated by the Haversine formula to obtain the IP address feature similarity; the cosine similarity between the web page content feature of the unmarked URL and the web page content feature of the marked URL is calculated to obtain the web page content feature similarity;
[0061] If the comprehensive similarity of the unmarked URL exceeds the preset similarity threshold, the unmarked URL is regarded as a candidate malicious URL and added to the candidate malicious URL set;
[0062] Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights. The Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities, specifically:
[0063] Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as:
[0064] ;
[0065] In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths;
[0066] The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stop condition is reached, so as to obtain a URL relationship graph divided into a plurality of subgraphs, wherein the stop condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community;
[0067] The result filtering module is used to calculate the nearest community distance of each node in the URL relationship graph , expressed as:
[0068] ;
[0069] In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated;
[0070] The comprehensive score of the URL is calculated based on the comprehensive similarity and the nearest community distance, expressed as:
[0071] ;
[0072] In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively;
[0073] Remove candidate malicious URLs whose comprehensive scores are lower than a preset comprehensive threshold from the candidate malicious URL set;
[0074] The output module is used to output the final set of candidate malicious URLs and the corresponding URL relationship graph.
[0075] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for discovering malicious URL communities based on URL fission as described in any embodiment of the present invention is implemented.
[0076] On the other hand, the present embodiment further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for discovering malicious URL communities based on URL fission as described in any embodiment of the present invention.
[0077] The present invention has the following beneficial effects:
[0078] 1. The present invention is a method and system for discovering malicious URL communities based on URL fission. The method extracts domain name features, IP address features and web page content features from a list of known URLs, characterizes URLs from multiple dimensions, and comprehensively captures malicious URL features. Because different malicious URLs have abnormal performances in different features, the features are comprehensively considered to improve recognition accuracy and effectively identify disguised malicious URLs. The present invention also calculates the similarity of the corresponding features of unmarked URLs and marked URLs, and then obtains a comprehensive similarity. If the similarity exceeds a preset threshold, it is considered as a candidate malicious URL. Through similarity analysis, potential malicious URLs with similar features to known malicious URLs are discovered, thereby avoiding missed judgments and broadening the scope of malicious URL discovery.
[0079] 2. The present invention is a method and system for discovering malicious URL communities based on URL fission. Based on a set of candidate malicious URLs, a URL relationship graph with URLs as nodes and comprehensive similarity as edge weights is dynamically constructed, and the Girvan-Newman algorithm is used to cluster the graphs. The community structure is revealed, and the degree of association between URLs is intuitively displayed through the URL relationship graph in the above steps. The closely related URL communities are automatically identified through the clustering algorithm, thereby improving the speed and accuracy of mining malicious URL communities.
[0080] 3. The present invention is a method and system for discovering malicious URL communities based on URL fission. By calculating the nearest community distance of each node in the URL relationship graph and combining the comprehensive similarity to calculate the comprehensive score of the URL, the candidate malicious URLs below the preset comprehensive threshold are removed from the set. By comprehensively considering the degree of association and similarity between the URL and the community, the real malicious URL community is accurately determined, and the misjudged candidate malicious URLs are further eliminated, thereby improving the accuracy of the final result. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 is a flow chart of the method of the present invention;
[0082] Figure 2 It is an overall framework diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0083] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0084] It should be understood that the step numbers used in this document are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0085] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0086] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0087] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.
[0088] Embodiment 1:
[0089] See also Figure 1 This embodiment provides a method for discovering malicious URL communities based on URL fission, identifying candidate fraud-related URLs based on feature similarity calculation and graph clustering analysis, and finally outputting a set of candidate fraud-related URLs and a corresponding URL relationship graph, including the following steps:
[0090] S1. Obtain a list of known URLs and mark a list of malicious URLs in the list of known URLs;
[0091] S11, the obtaining of the known website list specifically includes actively obtaining a large amount of website information through search engines, social networks and other channels. In this embodiment, a crawler is written by using the Scrapy framework of Python to obtain the website addresses on the search result page from search engines such as Baidu and Google, and a known website list is constructed based on the website addresses obtained by the search;
[0092] Preferably, when obtaining the list of known URLs, malicious URLs are determined based on feedback information of URLs in the search engine. For example, some browsers provide a function of reporting malicious websites. When a large number of users report a URL, the URL will be marked as malicious. Malicious URLs in the list of known URLs are marked based on the above feedback information.
[0093] S12. Perform data preprocessing on the known URL list, including data cleaning, URL standardization and data denoising. Further:
[0094] S121, the data cleaning includes removing duplicate URLs and processing invalid URLs. There may be duplicate URLs in the known URL list. These duplicate URLs will increase the burden of subsequent calculations and have no substantial help to the results. Therefore, it is necessary to remove duplicate URLs. In this embodiment, the removal of duplicate URLs is specifically to perform hash processing on the known URL list and establish a hash table, traverse each URL in the known URL list, and if the hash value of the URL already exists in the hash table, remove it from the known URL list; the processing of invalid URLs is specifically to check whether the format of the URL conforms to the standard URL format, such as whether it contains a legal protocol (such as http, https) and domain name, etc., and mark and remove URLs that do not conform to the format; at the same time, try to access each URL, and remove the URLs that cannot be accessed (such as returning an error status code such as 404, 500, etc.) from the known URL list;
[0095] S122, the URL standardization includes unifying the protocol, removing unnecessary parameters and decoding the URL encoding, wherein: the unifying protocol specifically unifies the protocol of all URLs to https or http to avoid feature differences caused by different protocols, for example, all URLs starting with "http: / / " are uniformly converted to "https: / / " if supported; the removing unnecessary parameters specifically removes some parameters in known URLs that do not affect their essential features, such as session ID, timestamp, etc.;
[0096] S123, the data denoising specifically includes removing short links, because short links usually hide the real website information, which is not conducive to subsequent feature extraction and analysis. In this embodiment, the short link is restored to the original long URL by calling the API of the short link service provider;
[0097] Through the data preprocessing in step S12, invalid, redundant or abnormal data is removed to ensure the quality of the collected website data and the accuracy and effectiveness of subsequent processing;
[0098] S2, extracting features from the known URL list to obtain domain name features, IP address features, and web page content features, calculating feature similarities between unmarked URLs and marked URLs in the known URL list, and calculating comprehensive similarities of unmarked URLs based on the feature similarities. If the comprehensive similarities of unmarked URLs exceed a preset similarity threshold, the unmarked URLs are considered candidate malicious URLs and added to the candidate malicious URL set.
[0099] S21. Feature extraction:
[0100] S211, dividing domain names at different levels based on known website information to obtain domain name features;
[0101] S212, resolving the domain name feature to obtain the corresponding IP address, querying the longitude and latitude corresponding to the IP address using the GeoLite2 database to obtain the IP address feature;
[0102] S213, the web page content features include text features and image features, crawling the text and image data in the web page corresponding to the known URL, using the large language model GPT-3 to perform text vectorization processing on the text data to obtain text features, and using the autoencoder to perform feature extraction on the image data to obtain image features. Since the extraction of text features and the extraction of image features are both existing technologies, they will not be repeated here;
[0103] S22, calculating the feature similarity between the unmarked URL and the marked URL in the known URL list is specifically:
[0104] S221. Calculate the edit distance between the domain name feature of the unlabeled URL and the domain name feature of the labeled URL by using the Damerau-Levenshtein distance, and perform normalization processing to express it as follows:
[0105] ;
[0106] In the formula, is the normalized edit distance; To mark the domain name corresponding to the URL; The domain name corresponding to the unmarked URL; For domain name and domain name The Damerau-Levenshtein distance is calculated for the string within. and Domain names and domain name The length of the corresponding string; To perform the maximum value operation;
[0107] in, Calculated by dynamic programming algorithm, specifically:
[0108] ;
[0109] ;
[0110] In the formula, To perform the minimum value operation; Domain name Before characters; Domain name Before characters; is the indicator function; Domain name Middle Bit character; Domain name Middle Bit character;
[0111] from , Start to increase and , and calculate the values of , until , ,get ;
[0112] The semantic similarity between the domain name features of the unlabeled URL and the domain name features of the labeled URL is calculated and expressed as:
[0113] ;
[0114] In the formula, and From domain name and domain name The extracted vocabulary set; Domain name and domain name The semantic similarity of and Domain names and domain name Index of vocabulary; Domain name The number of words in the vocabulary set;
[0115] The domain name feature similarity is calculated based on the normalized edit distance and semantic similarity, and is expressed as follows:
[0116] ;
[0117] is the domain name feature similarity; is the proportionality coefficient, which is dynamically adjusted according to the ratio of spelling errors to semantic deception in the known URL information;
[0118] S222. Calculate the distance between the IP address feature of the unmarked URL and the IP address feature of the marked URL by using the Haversine formula, and perform normalization processing to obtain the IP address feature similarity, which is expressed as:
[0119] ;
[0120] ;
[0121] In the formula, is the IP address feature similarity; is the intermediate parameter; and Respectively represent the longitude and latitude of the IP address of the marked URL; and Respectively represent the longitude and latitude of the unmarked URL IP address;
[0122] S223, calculating the cosine similarity between the webpage content features of the unmarked URL and the webpage content features of the marked URL, including the cosine similarity corresponding to the text features and the cosine similarity corresponding to the image features, normalizing the cosine similarity of the webpage content features, and obtaining the webpage content feature similarity, which is expressed as:
[0123] ;
[0124] In the formula, is the similarity of web page content features; and The text features representing the labeled URLs and the unlabeled URLs respectively; and Image features representing labeled URLs and unlabeled URLs respectively;
[0125] S23, calculating the comprehensive similarity of the unmarked URLs based on the feature similarity, and expressing it in the formula:
[0126] ;
[0127] In the formula, is the comprehensive similarity; , and They are the weights of domain name feature similarity, IP address feature similarity, and web page content feature similarity;
[0128] Preferably, in this embodiment, the weights of the domain name feature similarity, the IP address feature similarity and the web page content feature similarity are obtained by KL divergence calculation. In this embodiment, the domain name feature and the web page content feature are discrete data, and the IP address feature is a continuous feature. Therefore:
[0129] , ;
[0130] ;
[0131] In the formula, is the characteristic KL divergence; For the corresponding features, specifically, when hour, is a domain name feature; when hour, is the IP address feature; when hour, It is the IP address feature; is the probability distribution of the corresponding feature in the reference sample set; is the probability distribution of the corresponding feature in the data set to be calculated. In this embodiment, the data set to be calculated is a set of candidate malicious URLs;
[0132] S3. Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights, and the Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities;
[0133] S31, taking each URL in the candidate malicious URL set as a node, adding it to the initial URL relationship graph, calculating the comprehensive similarity between any two nodes, and for node pairs whose similarity is greater than a preset node threshold, adding an edge to the URL relationship graph, and using the comprehensive similarity as the weight of the edge;
[0134] This embodiment adopts a dynamic update mechanism, specifically:
[0135] When a new candidate malicious URL is added to the candidate malicious URL set, it is added as a new node in the URL relationship graph, and the comprehensive similarity between the new node and the existing nodes in the URL relationship graph is calculated. For node pairs with similarity greater than the node threshold, a new edge is added and the edge weight is set;
[0136] When a candidate malicious URL is removed from the candidate malicious URL set, the corresponding node and all edges connected to it are deleted from the URL relationship graph;
[0137] This embodiment recalculates the comprehensive similarity between all nodes periodically or in real time to reflect changes in URL features, and updates the weights of the edges in the URL relationship graph based on the new comprehensive similarity results. If the comprehensive similarity of some edges is lower than the node threshold, consider deleting these edges; if a new node pair with a comprehensive similarity greater than the node threshold appears, add a new edge;
[0138] Furthermore, in order to reduce the computational overhead, in this embodiment, when a new candidate malicious URL is added to the candidate malicious URL set, the comprehensive similarity between all nodes in the entire URL relationship graph is not recalculated immediately, but the local calculation of the new node and the existing node is focused on. Specifically:
[0139] Extract the features of the new node in detail, including domain name features, IP address features, and web page content features; hash the features of each dimension separately. For example, for domain name features, split them into multiple sub-parts and hash each sub-part. This way, complex features can be converted into fixed-length hash values for easy storage and quick comparison.
[0140] Constructing a local feature vector composed of hash values corresponding to the features to represent the new node, the local feature vector not only takes up less space, but also facilitates improving the calculation speed in the subsequent similarity calculation;
[0141] For each existing node in the URL relationship graph, its features are hashed in advance and stored in a hash index table. Based on the hash feature values of the nodes in the hash index table, the corresponding node information can be quickly located. In this embodiment, a Bloom filter is used as the structure of the hash index table, and the hash feature values of the existing nodes are stored in the Bloom filter. When a new node is added, the Bloom filter is used to quickly determine whether some features of the new node may be similar to those of the existing nodes, thereby reducing unnecessary detailed similarity calculations.
[0142] When calculating the comprehensive similarity between the new node and the existing nodes, a hash-based similarity calculation method is used: a set of existing nodes that may be similar to the new node is quickly screened out through a Bloom filter, and for the screened nodes, their hash feature vectors are further compared, including using methods such as Hamming distance and cosine similarity to calculate the similarity between hash vectors. For example, the Hamming distance of the hash value of the domain name feature is calculated, and the Hamming distance is used as the hash similarity indicator of the domain name feature. The smaller the Hamming distance, the higher the similarity.
[0143] The hash similarity indicators of different dimensions are weighted and combined to obtain the comprehensive similarity between the new node and each existing node. Preferably, in the weighting process, a suitable weight is assigned to each dimension according to the actual application scenario and the importance of different features;
[0144] After the local similarity calculation is completed, only the node pairs whose comprehensive similarity is greater than the preset similarity node threshold are processed, that is, new edges are added to the URL relationship graph and the edge weights are set, which avoids large-scale recalculation of the entire graph and significantly reduces the computational overhead;
[0145] Preferably, this embodiment also adopts a recursive strategy to repeat the process of dynamically constructing the URL relationship graph in step S3 until any of the following conditions is met: ① the set of candidate malicious URLs is empty; ② the number of recursions reaches a set recursion threshold; ③ the proportion of unmarked URLs for which the comprehensive similarity is to be calculated is lower than a set calculation threshold; the recursion terminates;
[0146] S32. After the recursion is completed, the Girvan-Newman algorithm is used to perform cluster analysis on the URL relationship graph to identify high-density clustered malicious URL communities, specifically:
[0147] Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as:
[0148] ;
[0149] In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths;
[0150] The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stop condition is reached, so as to obtain a URL relationship graph divided into a plurality of subgraphs, wherein the stop condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community;
[0151] S4, calculating the nearest community distance of each node in the URL relationship graph, calculating the URL comprehensive score based on the comprehensive similarity and the nearest community distance, removing the candidate malicious URLs whose URL comprehensive scores are lower than a preset comprehensive threshold from the candidate malicious URL set, and obtaining the final candidate malicious URL set and the corresponding URL relationship graph;
[0152] S41. Calculate the nearest community distance of each node in the URL relationship graph , expressed as:
[0153] ;
[0154] In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated;
[0155] S42. Calculate the comprehensive score of the website based on the comprehensive similarity and the nearest community distance, expressed as:
[0156] ;
[0157] In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively.
[0158] Embodiment 2:
[0159] like Figure 2 As shown, this embodiment provides a malicious URL community discovery system based on URL fission, the system includes a data collection module, a feature extraction module, a URL fission module, a result filtering module and an output module, wherein:
[0160] The data collection module is used to obtain a list of known web addresses and mark a list of malicious web addresses in the list of known web addresses;
[0161] The feature extraction module is used to extract features from a known website list to obtain domain name features, IP address features and web page content features;
[0162] The URL fission module is used to calculate the feature similarity between the unmarked URLs and the marked URLs in the known URL list, and calculate the comprehensive similarity of the unmarked URLs based on the feature similarity, which is expressed as:
[0163] ;
[0164] In the formula, is the comprehensive similarity; , and Domain name feature similarity , IP address feature similarity Similarity with web page content features The weight of the domain name feature similarity is calculated based on the normalized edit distance and semantic similarity; the distance between the IP address feature of the unmarked URL and the IP address feature of the marked URL is calculated by the Haversine formula to obtain the IP address feature similarity; the cosine similarity between the web page content feature of the unmarked URL and the web page content feature of the marked URL is calculated to obtain the web page content feature similarity;
[0165] If the comprehensive similarity of the unmarked URL exceeds the preset similarity threshold, the unmarked URL is regarded as a candidate malicious URL and added to the candidate malicious URL set;
[0166] Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights. The Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities, specifically:
[0167] Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as:
[0168] ;
[0169] In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths;
[0170] The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stop condition is reached, so as to obtain a URL relationship graph divided into a plurality of subgraphs, wherein the stop condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community;
[0171] The result filtering module is used to calculate the nearest community distance of each node in the URL relationship graph , expressed as:
[0172] ;
[0173] In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated;
[0174] The comprehensive score of the URL is calculated based on the comprehensive similarity and the nearest community distance, expressed as:
[0175] ;
[0176] In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively;
[0177] Remove candidate malicious URLs whose comprehensive scores are lower than a preset comprehensive threshold from the candidate malicious URL set;
[0178] The output module is used to output the final set of candidate malicious URLs and the corresponding URL relationship graph.
[0179] It is worth noting that the system described in this embodiment is used to run the method described in Example 1.
[0180] Embodiment 3:
[0181] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a method for discovering malicious URL communities based on URL fission as described in any embodiment of the present invention is implemented.
[0182] Embodiment 4:
[0183] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, a method for discovering a malicious URL community based on URL fission as described in any embodiment of the present invention is implemented.
[0184] It is worth noting that the system, electronic device and computer-readable storage medium described in the present invention are all based on the same inventive concept as the method described in Example 1 of the present invention, and will not be described in detail here.
[0185] In the embodiments of the present invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c may be single or multiple.
[0186] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0188] In several embodiments provided by the present invention, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk and other media that can store program codes.
[0189] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for discovering malicious URL communities based on URL fission, characterized in that: The method comprises: Obtain a list of known URLs and mark a list of malicious URLs in the list of known URLs; Perform feature extraction on the known URL list to obtain domain name features, IP address features and web page content features, calculate the feature similarity of the unmarked URLs and marked URLs in the known URL list, and calculate the comprehensive similarity of the unmarked URLs based on the feature similarity, which is expressed as: ; In the formula, is the comprehensive similarity; , and Domain name feature similarity , IP address feature similarity Similarity with web page content features The weight of the domain name feature similarity is calculated based on the normalized edit distance and semantic similarity; the distance between the IP address feature of the unmarked URL and the IP address feature of the marked URL is calculated by the Haversine formula to obtain the IP address feature similarity; the cosine similarity between the web page content feature of the unmarked URL and the web page content feature of the marked URL is calculated to obtain the web page content feature similarity; If the comprehensive similarity of the unmarked URL exceeds the preset similarity threshold, the unmarked URL is regarded as a candidate malicious URL and added to the candidate malicious URL set; Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights. The Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities, specifically: Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as: ; In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths; The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stop condition is reached, so as to obtain a URL relationship graph divided into a plurality of subgraphs, wherein the stop condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community; Calculate the nearest community distance of each node in the URL relationship graph , expressed as: ; In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated; The comprehensive score of the URL is calculated based on the comprehensive similarity and the nearest community distance, expressed as: ; In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively; The candidate malicious URLs whose comprehensive scores are lower than a preset comprehensive threshold are removed from the candidate malicious URL set, and a final relationship diagram between the candidate malicious URL set and the corresponding URLs is obtained.
2. According to claim 1, a method for discovering malicious URL communities based on URL fission, characterized in that: The method also includes performing data preprocessing on the known website list before extracting features from the known website list, including data cleaning, URL standardization and data denoising.
3. A method for discovering malicious URL communities based on URL fission according to claim 2, characterized in that: The feature extraction of the known website list to obtain domain name features, IP address features and web page content features is specifically as follows: Divide domain names at all levels based on known URL information to obtain domain name features; The domain name feature is parsed to obtain the corresponding IP address, and the longitude and latitude corresponding to the IP address are queried using the GeoLite2 database to obtain the IP address feature; The web page content features include text features and image features. The text and image data in the web page corresponding to the known URL are captured, and the text data is vectorized using the large language model GPT-3 to obtain text features. The image data is feature extracted using an autoencoder to obtain image features.
4. A method for discovering malicious URL communities based on URL fission according to claim 3, characterized in that: The specific method of calculating the feature similarity between the unmarked URLs and the marked URLs in the known URL list is as follows: The edit distance between the domain name features of the unlabeled URL and the domain name features of the labeled URL is calculated by the Damerau-Levenshtein distance, and normalized to be expressed as follows: ; In the formula, is the normalized edit distance; To mark the domain name corresponding to the URL; The domain name corresponding to the unmarked URL; For domain name and domain name The Damerau-Levenshtein distance is calculated for the string within. and Domain names and domain name The length of the corresponding string; To perform the maximum value operation; in, Calculated by dynamic programming algorithm, specifically: ; ; In the formula, To perform the minimum value operation; Domain name Before characters; Domain name Before characters; is the indicator function; Domain name Middle Bit character; Domain name Middle Bit character; from , Start to increase and , and calculate , until , ,get ; The semantic similarity between the domain name features of the unlabeled URL and the domain name features of the labeled URL is calculated and expressed as: ; In the formula, and From domain name and domain name The extracted vocabulary set; Domain name and domain name The semantic similarity of and Domain names and domain name Index of vocabulary; Domain name The number of words in the vocabulary set; The domain name feature similarity is calculated based on the normalized edit distance and semantic similarity, and is expressed as follows: ; is the domain name feature similarity; is the proportionality coefficient; The distance between the IP address features of the unmarked URL and the IP address features of the marked URL is calculated by the Haversine formula, and normalized to obtain the IP address feature similarity, which is expressed as: ; ; In the formula, is the IP address feature similarity; is the intermediate parameter; and Respectively represent the longitude and latitude of the IP address of the marked URL; and Respectively represent the longitude and latitude of the unmarked URL IP address; Calculate the cosine similarity between the webpage content features of the unmarked URL and the webpage content features of the marked URL, including the cosine similarity corresponding to the text features and the cosine similarity corresponding to the image features. Normalize the cosine similarity of the webpage content features to obtain the webpage content feature similarity, which is expressed as: ; In the formula, is the similarity of web page content features; and The text features representing the labeled URLs and the unlabeled URLs respectively; and Represent the image features of labeled URLs and unlabeled URLs respectively.
5. A malicious URL community discovery system based on URL fission, characterized in that: The system includes a data collection module, a feature extraction module, a URL fission module, a result filtering module and an output module, wherein: The data collection module is used to obtain a list of known web addresses and mark a list of malicious web addresses in the list of known web addresses; The feature extraction module is used to extract features from a known website list to obtain domain name features, IP address features and web page content features; The URL fission module is used to calculate the feature similarity between the unmarked URLs and the marked URLs in the known URL list, and calculate the comprehensive similarity of the unmarked URLs based on the feature similarity, which is expressed as: ; In the formula, is the comprehensive similarity; , and Domain name feature similarity , IP address feature similarity Similarity with web page content features The weight of the domain name feature similarity is calculated based on the normalized edit distance and semantic similarity; the distance between the IP address feature of the unmarked URL and the IP address feature of the marked URL is calculated by the Haversine formula to obtain the IP address feature similarity; the cosine similarity between the web page content feature of the unmarked URL and the web page content feature of the marked URL is calculated to obtain the web page content feature similarity; If the comprehensive similarity of the unmarked URL exceeds the preset similarity threshold, the unmarked URL is regarded as a candidate malicious URL and added to the candidate malicious URL set; Based on the candidate malicious URL set, a URL relationship graph is dynamically constructed with URLs as nodes and comprehensive similarity as edge weights. The Girvan-Newman algorithm is used to cluster the URL relationship graph to obtain candidate malicious URL communities, specifically: Calculate the edge betweenness centrality of each edge in the URL relationship graph, expressed as: ; In the formula, For edge The edge betweenness centrality of Represents a slave node To Node The number of shortest paths; Indicates that the edge Slave Node To Node The number of shortest paths; The edge betweenness centrality of each edge in the URL relationship graph is sorted in descending order, and each edge is deleted in turn until a set stop condition is reached, so as to obtain a URL relationship graph divided into a plurality of subgraphs, wherein the stop condition includes that the proportion of the remaining edges is reduced to a preset boundary threshold, and the subgraph represents a malicious URL community; The result filtering module is used to calculate the nearest community distance of each node in the URL relationship graph , expressed as: ; In the formula, The coordinates of the center node of the malicious URL cluster; It is the central node of the malicious URL community; is the node coordinates for which the nearest community distance is to be calculated; The comprehensive score of the URL is calculated based on the comprehensive similarity and the nearest community distance, expressed as: ; In the formula, Comprehensive score for the website; and are the weights of comprehensive similarity and nearest community distance, respectively; Remove candidate malicious URLs whose comprehensive scores are lower than a preset comprehensive threshold from the candidate malicious URL set; The output module is used to output the final set of candidate malicious URLs and the corresponding URL relationship graph.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements a malicious URL community discovery method based on URL fission as described in any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements a malicious URL community discovery method based on URL fission as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method and system for carrying out fraud-related detection on large-batch websites
CN117596003A
Website verification method and device
CN105099996A
Rule discovery method and system
CN108171053A