Website fingerprint calculation method, system, storage medium and terminal
By calculating the structural vector values of the website, analyzing static resources and feature fields, automatically classifying and determining the characteristics of the sample website as fingerprints, the inefficiency and false positive problems caused by manual collection in the prior art are solved, and efficient and automated website fingerprint calculation is achieved.
Patent Information
- Application Number
- CN202111487908.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-07
AI Technical Summary
In the prior art, fingerprint calculation of batch website samples relies on manual collection, resulting in low computing efficiency and prone to false alarms.
By obtaining website samples, calculating the structural vector value of the document objectified model of the target website, analyzing the static resource list and feature fields, classifying the website based on this information, and determining the characteristics of the sample website as the website fingerprint.
Automatic website fingerprint calculation can be realized, and similar websites can be found in massive sample websites, reducing labor investment, improving computing efficiency and reducing labor costs.
Smart Images

Figure CN114154043B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security, and in particular to a website fingerprint calculation method, a calculation system, a storage medium and a terminal. Background Art
[0002] Currently, in application development, it is often necessary to obtain the identity of the website application, that is, to obtain the website fingerprint. However, fingerprinting for batch website samples mainly relies on manual collection, which requires collecting the characteristic fields of each website and performing pairwise comparisons between websites based on the characteristic fields. Once the number of website samples is large, the calculation efficiency of the website fingerprint will be greatly reduced, and false positives are prone to occur.
[0003] Therefore, how to improve the computational efficiency of website fingerprints is a technical problem that those skilled in the art need to solve urgently. Summary of the invention
[0004] The purpose of this application is to provide a website fingerprint calculation method, a calculation system, a storage medium and a terminal, which can improve the calculation efficiency of the website fingerprint.
[0005] In order to solve the above technical problems, this application provides a method for calculating website fingerprints. The specific technical solution is as follows:
[0006] Obtaining a website sample, and determining a target website from the website sample;
[0007] Calculate the structure vector value of the document object model corresponding to the target website;
[0008] Crawling the target website, obtaining a static resource list, parsing the static file resource list of the target website, and outputting a website list corresponding to each static resource in the website sample;
[0009] Parsing the characteristic fields of the target website;
[0010] Classifying the websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determining the sample website;
[0011] The characteristics of the example website are used as the website fingerprint.
[0012] Optionally, the step of calculating the structure vector value of the document object model of the target website includes:
[0013] Obtain the target website html page and construct the document object model;
[0014] In the document object model, a parent node is selected as a head element as a target node, and the element name and attribute of each target node are concatenated into a string;
[0015] Calculate the hash value of the string, and multiply the hash value by the weight of the target node to obtain the weight value corresponding to the target node; wherein, the greater the node depth of the target node, the more nodes are the same as the target node, and the smaller the weight of the target node;
[0016] Accumulate the weight values of all target nodes to get the structure vector value.
[0017] Optionally, the parsing of the static file resource list of the target website includes:
[0018] Preprocessing the static resources in the static file resource list to remove characteristic information of public library resources and static resources;
[0019] Constructing a static resource dictionary, calculating a static hash value for adjacent static file resource names using a preset formula, and establishing a mapping relationship between the static hash value, the static file resource name list, and the web page address corresponding to the static file resource;
[0020] Calculate the hash value of each static file resource name in the static file resource list to obtain a hash value list corresponding to the static file resource list;
[0021] The static hash value is calculated by using a preset formula for adjacent static file resource names;
[0022] Determine whether the static resource dictionary contains the static hash value;
[0023] If so, it is determined that there is an intersection between the static file resource lists of the target website and other websites, and the web page address of the target website is added to the list of web page addresses corresponding to the static file resources;
[0024] If not, save the static hash value and the corresponding static file resource name list, and the web page address corresponding to the static file resource.
[0025] Optionally, the preset formula is:
[0026]
[0027] Where i is the number of adjacent static file resources taken in each calculation and i is greater than 2, j is the index number of the first static file resource in the static file resource list among the static file resources taken in each calculation, k is iterative traversal, which is used to traverse all static resources with index numbers in the interval [j, j+i-1], x ij A static hash value.
[0028] Optionally, preprocessing the static resources in the static file resource list to remove characteristic information of public library resources and static resources includes:
[0029] Configure path blacklist and / or file name blacklist for public library resources;
[0030] Delete at least one of the version number and the random number in the static file resource name, and remove the domain name or IP address in the path corresponding to the static resource.
[0031] Optionally, the websites are classified according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and the example websites are determined to include:
[0032] According to the structure vector value, the website list corresponding to each static resource and the characteristic field, the websites in the website sample are analyzed for association and classified, and any original website in each class has at least one similar website, and the original website and the similar website have at least two of the structure vector value, the website list corresponding to each static resource and the characteristic field pair website in common;
[0033] At least one example website is identified in each category of websites.
[0034] Optionally, when defining the sample website, also include:
[0035] Discard sample websites that do not have corresponding similar websites.
[0036] The present application also provides a website fingerprint computing system, including:
[0037] A website acquisition module, used to acquire website samples and determine a target website from the website samples;
[0038] A structure vector value calculation module, used to calculate the structure vector value of the document object model corresponding to the target website;
[0039] A static resource analysis and calculation module is used to crawl the target website, obtain a static resource list, parse the static file resource list of the target website, and output a website list corresponding to each static resource in the website sample;
[0040] A feature field acquisition module, used to parse the feature fields of the target website;
[0041] An association analysis module, used to classify websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determine the example website;
[0042] The fingerprint calculation module is used to use the characteristics of the example website as the website fingerprint.
[0043] The present application also provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the steps of the above-mentioned method when executed by a processor.
[0044] The present application also provides a terminal, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps of the above method when calling the computer program in the memory.
[0045] The present application provides a method for calculating a website fingerprint, comprising: obtaining a website sample, and determining a target website from the website sample; calculating a structural vector value of a document object model corresponding to the target website; crawling the target website to obtain a static resource list, parsing the static file resource list of the target website, and outputting a website list corresponding to each static resource in the website sample; parsing a feature field of the target website; classifying websites according to the structural vector value, the website list corresponding to each static resource, and the feature field, and determining an example website; and using the features of the example website as the website fingerprint.
[0046] This application uses the calculated structural vector value of the website to analyze the static file resource list and feature fields of the website, and classifies the website according to the structural vector value, the website list and feature fields corresponding to each static resource, so as to determine representative sample websites, and use the features of the sample websites as website fingerprints. It can automatically find similar websites in massive sample websites, and extract the common features of similar websites into fingerprints, which can greatly reduce manpower investment, improve the calculation efficiency of website fingerprints, and reduce labor costs.
[0047] The present application also provides a website fingerprint detection system, a computer-readable storage medium and a terminal, which have the above-mentioned beneficial effects and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0049] Figure 1 A flowchart of a method for calculating a website fingerprint provided in an embodiment of the present application;
[0050] Figure 2 A calculation flow chart of the structure vector value provided in the embodiment of the present application;
[0051] Figure 3A schematic diagram of a website fingerprint computing system provided in an embodiment of the present application:
[0052] Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0054] See also Figure 1 , Figure 1 A flowchart of a method for calculating a website fingerprint provided in an embodiment of the present application, the method comprising:
[0055] S101: Obtain website samples, and determine a target website from the website samples;
[0056] This step aims to obtain a website sample and determine the target website in the website sample. The target website refers to the website whose web page similarity is to be calculated. Usually, each website in the website sample can be used as a target website. Some websites can also be selected from the website sample as a set to determine the website fingerprints corresponding to the websites in the set. In this case, the target website can be any website in the set. It should be noted that in the process of calculating the website fingerprint, it is necessary to refer to each web page of the website to calculate the similarity between the web pages, so as to find similar websites in a large number of websites, and use the common features of similar websites as website fingerprints.
[0057] There is no limitation on how to obtain website samples. You can crawl websites in real time or obtain websites from a website database as website samples in this embodiment. When determining the target website, you need to ensure that each website sample can respond. Specifically, you can use Python's requests library to construct a request header and send a request to each website sample. If the response of the website sample can be correctly obtained, the parsing operation of the target website is performed, otherwise try to obtain the response of the next website sample.
[0058] S102: Calculating the structure vector value of the document object model corresponding to the target website;
[0059] This step aims to calculate the structural vector value of the document object model corresponding to the target website, that is, the structural vector value of the DOM (Document object model tree structure) tree, which can also be called DOM_Value. DOM is a standard of W3C (World Wide Web Consortium), and HTML (Hyper Text Markup Language) DOM is a standard on how to get, modify, add or delete HTML elements. In HTML DOM, everything is a node, and the nodes are presented in the form of a tree. This node tree starts from the root node, showing the collection of nodes and the connections between nodes, and then branches out to the text nodes at the lowest level of the tree.
[0060] There is no specific limitation on how to calculate the structure vector value, which is usually based on the process of calculating the hash value of the string and configuring the weight of the node. Those skilled in the art can customize the structure vector value calculation process, which is not specifically limited here.
[0061] See also Figure 2 , Figure 2 Figure 2 The calculation flow chart of the structure vector value provided in the embodiment of the present application is a specific implementation method of this step, and the process can be as follows:
[0062] S1021: Obtain the HTML page of the target website and construct a document object model;
[0063] S1022: selecting a parent node as a head element as a target node in the document object model, and concatenating the element name and attributes of each target node into a string;
[0064] S1023: Calculate the hash value of the string, and multiply the hash value by the weight of the target node to obtain the weight value corresponding to the target node;
[0065] It should be noted that the greater the node depth of the target node, the more nodes are the same as the target node, and the smaller the weight of the target node;
[0066] S1024: Accumulate the weight values of all target nodes to obtain a structure vector value.
[0067] First, for the website pages of the target website, that is, the HTML pages, a document object model is constructed, that is, the logical relationship between the website pages is converted into corresponding nodes, so that the document object model is obtained according to the logical relationship between the nodes. After that, the parent node is selected as the head element (also known as "element") as the target node, the element name and attributes of each target node are concatenated into a string, the hash value of the string is calculated, and then all the hash values are multiplied by the node weight and accumulated. The final value is the structural vector value of the web page. It should be noted that the greater the node depth, the more nodes are the same as the node, and the smaller the weight of the node.
[0068] The structural vector value calculated in this way has the following characteristics:
[0069] If the contents of the header elements of two web pages are consistent, then the structural vector values of the two websites are consistent; if the two web pages are not much different in the header elements, then the structural vector values of the two websites are similar.
[0070] S103: crawling the target website, obtaining a static resource list, parsing the static file resource list of the target website, and outputting a website list corresponding to each static resource in the website sample;
[0071] This step requires crawling the target website to obtain the static file resource list of the target website. The static file resource list contains static resources that can be used as website fingerprints, so this step requires crawling to obtain the static file resource list in order to select static resource files suitable as website resources.
[0072] However, it should be noted that not all static resources in the static file resource list are suitable as website fingerprints. Therefore, in this step, the static resources in the static file resource list can be pre-processed to remove the characteristic information of public library resources and static resources.
[0073] Specifically, the following processes may be included:
[0074] Step 1: Configure the path blacklist and / or file name blacklist of public library resources;
[0075] The second step is to delete at least one of the version number and the random number in the static file resource name, and remove the domain name or IP address in the path corresponding to the static resource.
[0076] Since public library resources cannot be used as application fingerprints, such as 'index.css', 'jquery.js', etc., it is necessary to set blacklists for paths and file names. For example, you can set four blacklists for path, css, js, and ico. Some common path names or file names are saved in the blacklist. If the path and file name of a static file are both in the blacklist (such as ' / js / jquery.min.js'), the data cannot be used as a fingerprint. This type of data can be avoided when parsing the static file resource list.
[0077] Similarly, some static resources have version numbers in their file names, and some have random numbers in their file names. If these two contents are retained, it will be impossible to classify the same static files into one category when extracting common features later. Therefore, the version number and random number need to be removed. For example, after removing the version number and random number from 'uni.webview.1.5.1.js' and 'chunk-vendors.71719f9c.js', 'uni.webview.js' and 'chunk-vendors.js' are retained.
[0078] However, the paths of some static resources retain the domain name or IP address of the website, which also cannot be used as fingerprints. It is necessary to remove the domain name and IP address of the website and only retain the relative path part of the static resource.
[0079] like:
[0080] 'https: / / truckresource.g7s.huoyunren.com / js / es6-promise / es6-promise.min.js', you can remove the website domain name and keep ' / js / es6-promise / es6-promise.min.js'.
[0081] After the above static resource preprocessing, a valid static resource list of the target website can be obtained. However, if the static resource list is directly used as the feature of the website for counting and comparison, there may be many omissions. For example, the static_list1 of website 1 is ['1.css', '2.css', '3.css', '4.js'], and the static_list2 of website 2 is ['1.css', '2.css', '3.css', '5.js']. They have common static resources ['1.css', '2.css', '3.css'], but directly comparing the static resource lists cannot obtain effective common features.
[0082] To this end, a preferred execution idea of this step can be as follows: for a website's static file resource list, traverse all its subsets with a length greater than 1, use a preset formula to calculate a static hash value for each subset, and determine whether the static hash value exists in the resource dictionary. If it does, it means that the static file resource lists of this website and other websites have an intersection, and these websites can be considered to have similarities and be classified into one category; if it does not exist, first save this static hash value and its corresponding subset static_text_list and the website's web address. Finally, once the length of the web address list is greater than 1, it means that multiple websites contain these static resource files, and these static resource file names and corresponding website lists are output. The specific implementation process can be as follows:
[0083] S1031: construct a static resource dictionary, calculate the static hash value of adjacent static file resource names by a preset formula, and establish a mapping relationship between the static hash value, the static file resource name list and the web page address corresponding to the static file resource;
[0084] S1032: Calculate the hash value of each static file resource name in the static file resource list to obtain a hash value list corresponding to the static file resource list;
[0085] S1033: Calculate the adjacent static file resource names using a preset formula to obtain a static hash value;
[0086] S1034: Determine whether the static resource dictionary contains a static hash value; if so, proceed to S1035; if not, proceed to S1036;
[0087] S1035: Determine that there is an intersection between the static file resource lists of the target website and other websites, and add the web page address of the target website to the web page address list corresponding to the static file resources;
[0088] S1036: Save the static hash value and the corresponding static file resource name list, and the web page address corresponding to the static file resource.
[0089] The preset formula is:
[0090]
[0091] Where i is the number of adjacent static file resources taken in each calculation and i is greater than 2, j is the index number of the first static file resource in the static file resource list among the static file resources taken in each calculation, k is iterative traversal, which is used to traverse all static resources with index numbers in the interval [j, j+i-1], x ij A static hash value.
[0092] In order to better describe the above process, the following example illustrates steps S1031 to S1036:
[0093] In step S1031, a static resource dictionary static_hash_dict{} is constructed, mapped to 'static_hash:(static_text_list,url_list)', where static_hash is the value obtained by calculating multiple adjacent static resource file names through a preset formula, static_text_list represents multiple static resource file names corresponding to the value, and url_list stores all website URLs that have these static resource files.
[0094] The static file resource list obtained by each website crawler is processed as follows:
[0095] If the static file resource list obtained by the website crawler is [s0, s1, ..., sn-1], calculate the hash value of each static resource file name to obtain the hash value list hash_list [h0, h1, ..., hn-1] corresponding to the static file resource list.
[0096] For i=n,n-1,...,2,j=0,1,...,ni, xij is calculated using the preset formula. i represents the number of adjacent static resources taken each time, and j represents the index number of the first static resource in the static file resource list among the multiple static resources taken each time. k is an iteration variable, which will traverse all static resources with index numbers in the interval [j,j+i-1]. xij is the value calculated using the preset formula for given i and j, which is equivalent to static_hash mentioned above.
[0097] If the static resource dictionary static_hash_dict contains the static hash value xij, then add the URL of the website to its corresponding web address list url_list;
[0098] If the static resource word static_hash_dict does not contain the static hash value xij, then add a mapping of 'xij:([sj,sj+1,...,sj+i-1],[url])' to static_hash_dict.
[0099] After all target websites are processed according to the above process, the static resource dictionary static_hash_dict is traversed. Once the length of the web address list is greater than 1, that is, len(url_list)>1, the static resource file name and the corresponding website list are saved and output, that is, (static_text_list, url_list).
[0100] The meaning of the preset formula is: each time a subset of length i (i ≥ 2) is selected, index number j is selected as the iteration starting point, and iteration variable k is used to iterate in the index number interval [j, j+i-1], and each h k Multiply k by the serial number (k-j+1) in the interval [j,j+i-1], and then add up the adjacent numbers with different numbers to get xij. For example, the static resource dictionary static_list1 of website 1 is ['1.css', '2.css', '3.css', '4.js'], and its corresponding hash_list is [h1, h2, h3, h4]. Then all the static hash values calculated by the preset formula are:
[0101] 'h1-2h2+3h3-4h4,h1-2h2+3h3,h2-2h3+3h4,h1-2h2,h2-2h3,h3-2h4'.
[0102] Adopting the method of using adjacent numbers with different signs can prevent overflow caused by the value being too large, and also include the requirements for the deployment order of static resources. Multiplying by a coefficient can effectively prevent the occurrence of the same name and cause the calculation result to be 0.
[0103] To prevent repeated comparisons of the same static resource, an identifier can be set for each static resource. If it is determined that there are static resources on other websites that are consistent with the static resource, no subsequent comparison will be performed. For example, in the above example, static_list1['1.css','2.css','3.css','4.js'] and static_list2['1.css','2.css','3.css','5.js'] have been identified as having the same static resources ['1.css','2.css','3.css'], so ['1.css','2.css'] will no longer be identified.
[0104] S104: parsing the characteristic fields of the target website;
[0105] For other feature information obtained during the crawling process, it is necessary to parse and obtain corresponding feature fields. The specific feature fields are not limited here, and may include one of the following, or any combination of several. Of course, those skilled in the art may also parse and obtain other feature fields, which are not limited here by examples:
[0106] ①Server: Server field in the http response header;
[0107] ②Cookies: The Set-Cookie field in the http response header only takes its key as a feature;
[0108] ③Www-Authenticate:Www-Authenticate field in the http response header;
[0109] ④X-Csrf-Token: X-Csrf-Token field in the http response header;
[0110] ⑤X-Powered-By: X-Powered-By field in the http response header;
[0111] ⑥title: website title information;
[0112] ⑦application-name:name="application-name" in the meta tag of the html page, take its content value;
[0113] ⑧meta_copyright: name="copyright" in the meta tag of the html page, take its content value;
[0114] ⑨description: In the meta tag of the html page, name = "description", take its content value;
[0115] ⑩generator: In the meta tag of the html page, name = "generator", take its content value;
[0116] keywords:name="keywords" in the meta tag of the html page, take its content value;
[0117] author:name="author" in the meta tag of the html page, take its content value;
[0118] Copyright: The description marked with "All rights reserved" or "copyright" in the body of the HTML page, such as "Copyright©; 2001-2020, Tencent Cloud.";
[0119] powered_by: indicates the description of "powered by" in the body of the html page, such as "Powered byDiscuz!X3.4";
[0120] favicon: the page icon of the website, take its md5 value for storage;
[0121] Robots.txt: The file obtained by adding "robots.txt" to the website path, takes its md5 value for storage.
[0122] It should be noted that in other application embodiments of the present application, this step can be executed simultaneously with the previous step, or this step can be executed first and then the previous step, that is, there is no established execution order relationship between the process of parsing the characteristic fields of the target website and the process of parsing the static file resource list of the target website and outputting the website list corresponding to each static resource in the website sample.
[0123] S105: Classifying the websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determining the example website;
[0124] After the above steps, the websites are classified according to the obtained structural vector values, website lists and feature fields, so as to determine the sample websites.
[0125] There is no specific limitation on how to classify here. The essence is to perform correlation analysis between features. Some features, such as 'Server', 'X-Powered-By', etc., can generate fingerprints alone. But there are other features that are not very reliable when used alone to generate fingerprints. For example, the favicon feature is used as the page icon of a website. Some companies may use the same favicon for all their products, so the favicon feature cannot be used as the fingerprint of a single application. In this case, multiple features need to be combined to generate a fingerprint. The purpose of feature correlation analysis is to find websites with multiple identical features in the crawler results and classify them into one category.
[0126] The websites in the website sample can be associated and classified based on the structural vector value, the website list corresponding to each static resource and the characteristic fields. If there is at least one similar website for any original website in each class, and the original website and the similar website have at least two of the same structural vector value, the website list corresponding to each static resource and the characteristic fields for the website, then at least one example website can be determined in each category of websites.
[0127] For example: Website 1 and Website 2 have the same structure vector value and favicon feature, and Website 2 and Website 3 have the same website list and favicon feature. The probability that these three websites use the same system is very high. First, using the structure vector value as the basic feature, it can be found that Website 1 and Website 2 are similar, and Website 1 and Website 2 are classified into the same category. Then using static_list as the basic feature, it can be found that Website 2 and Website 3 are similar, and Website 3 is then classified into the category of Website 1 and Website 2.
[0128] Feature association can find similar websites in the crawler website samples and classify similar websites into one category, so that a large number of crawler website samples are divided into multiple similar website categories. If a website cannot find a similar website in the sample, the website will be discarded.
[0129] During the execution of this step, sample websites that do not have corresponding similar websites can also be discarded. Select a feature, classify websites with the same feature into one category, use a website classification list to save all websites with the feature, and display the number of websites with the feature. If a feature only appears once in the entire crawler result, it cannot be used as fingerprint data and needs to be discarded. Correspondingly, if a website cannot find a similar website in the sample, the website will be discarded.
[0130] S106: Using the features of the example website as the website fingerprint.
[0131] The common features of similar websites may become fingerprints. The sample website is a representative of several similar websites, and its feature information is the common features of multiple similar websites. It is possible to extract the features of the sample website to obtain fingerprints.
[0132] The embodiment of the present application utilizes the calculated structural vector value of the website, analyzes the static file resource list and feature fields of the website, and classifies the website according to the structural vector value, the website list and feature fields corresponding to each static resource, thereby determining representative example websites, and using the features of the example websites as website fingerprints. It can automatically find similar websites in massive sample websites, and extract the common features of similar websites into fingerprints, which can greatly reduce manpower input and reduce labor costs.
[0133] A website fingerprint calculation system provided in an embodiment of the present application is introduced below. The website fingerprint calculation system described below and the website fingerprint calculation method described above can refer to each other.
[0134] See also Figure 3 , Figure 3This is a schematic diagram of a website fingerprint computing system provided in an embodiment of the present application. The present application also provides a website fingerprint computing system, including:
[0135] A website acquisition module, used to acquire website samples and determine a target website from the website samples;
[0136] A structure vector value calculation module, used to calculate the structure vector value of the document object model corresponding to the target website;
[0137] A static resource analysis and calculation module is used to crawl the target website, obtain a static resource list, parse the static file resource list of the target website, and output a website list corresponding to each static resource in the website sample;
[0138] A feature field acquisition module, used to parse the feature fields of the target website;
[0139] An association analysis module, used to classify websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determine the example website;
[0140] The fingerprint calculation module is used to use the characteristics of the example website as the website fingerprint.
[0141] Based on the above embodiment, as a preferred embodiment, the structure vector value calculation module includes:
[0142] A document object model building unit, used to obtain the HTML page of the target website and construct the document object model;
[0143] A string construction unit, used for selecting a parent node as a head element as a target node in the document object model, and concatenating the element name and attributes of each target node into a string;
[0144] A hash value calculation unit, used to calculate the hash value of the string, and multiply the hash value by the weight of the target node to obtain the weight value corresponding to the target node; wherein, the greater the node depth of the target node, the more nodes are the same as the target node, and the smaller the weight of the target node;
[0145] The structural vector value calculation unit is used to accumulate the weight values of all target nodes to obtain the structural vector value.
[0146] Based on the above embodiment, as a preferred embodiment, the static resource analysis and calculation module includes:
[0147] A preprocessing unit, used for preprocessing the static resources in the static file resource list to remove characteristic information of public library resources and static resources;
[0148] A dictionary generation unit is used to construct a static resource dictionary, calculate the static hash value of adjacent static file resource names by a preset formula, and establish a mapping relationship between the static hash value, the static file resource name list and the web page address corresponding to the static file resource;
[0149] A hash value list generating unit, used to calculate the hash value of each static file resource name in the static file resource list, and obtain a hash value list corresponding to the static file resource list;
[0150] A static hash value calculation unit, used to calculate the static hash value of adjacent static file resource names using a preset formula;
[0151] A static resource analysis unit is used to determine whether the static resource dictionary contains the static hash value; if so, determine whether the static file resource lists of the target website and other websites have an intersection, and add the web address of the target website to the list of web addresses corresponding to the static file resources; if not, save the static hash value and the corresponding static file resource name list, and the web address corresponding to the static file resource.
[0152] The preset formula is:
[0153]
[0154] Where i is the number of adjacent static file resources taken in each calculation and i is greater than 2, j is the index number of the first static file resource in the static file resource list among the static file resources taken in each calculation, k is iterative traversal, which is used to traverse all static resources with index numbers in the interval [j, j+i-1], x ij A static hash value.
[0155] Based on the above embodiment, as a preferred embodiment, the preprocessing unit includes:
[0156] A blacklist configuration subunit is used to configure a path blacklist and / or a file name blacklist of public library resources;
[0157] The invalid data processing unit is used to delete at least one of the version number and the random number in the static file resource name, and remove the domain name or IP address in the path corresponding to the static resource.
[0158] Based on the above embodiment, as a preferred embodiment, the association analysis module is a module for performing the following steps:
[0159] According to the structure vector value, the website list corresponding to each static resource and the characteristic field, the websites in the website sample are analyzed for association and classified, and any original website in each class has at least one similar website, and the original website and the similar website have at least two of the structure vector value, the website list corresponding to each static resource and the characteristic field pair website in common;
[0160] At least one example website is identified in each category of websites.
[0161] The present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps of the malicious encrypted traffic detection method provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0162] The present application also provides a terminal, which may include a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the website fingerprint calculation method provided in the above embodiment can be implemented. Of course, the terminal may also include various network interfaces, power supplies and other components. Figure 4 , Figure 4 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal in this embodiment may include: a processor 2101 and a memory 2102.
[0163] Optionally, the terminal may further include a communication interface 2103 , an input unit 2104 , a display 2105 and a communication bus 2106 .
[0164] The processor 2101 , the memory 2102 , the communication interface 2103 , the input unit 2104 , and the display 2105 all communicate with each other via the communication bus 2106 .
[0165] In the embodiment of the present application, the processor 2101 may be a central processing unit (CPU), a specific application integrated circuit, a digital signal processor, a readily available programmable gate array or other programmable logic device, etc.
[0166] The processor may call the program stored in the memory 2102. Specifically, the processor may execute the operations executed by the terminal in the above embodiment.
[0167] The memory 2102 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions. In the embodiment of the present application, the memory at least stores a program for implementing the following functions:
[0168] Obtaining a website sample, and determining a target website from the website sample;
[0169] Calculate the structure vector value of the document object model corresponding to the target website;
[0170] Crawling the target website, obtaining a static resource list, parsing the static file resource list of the target website, and outputting a website list corresponding to each static resource in the website sample;
[0171] Parsing the characteristic fields of the target website;
[0172] Classifying the websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determining the sample website;
[0173] The characteristics of the example website are used as the website fingerprint.
[0174] In one possible implementation, the memory 2102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function (such as a topic detection function, etc.); the data storage area may store data created during the use of the computer.
[0175] In addition, the memory 2102 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.
[0176] The communication interface 2103 may be an interface of a communication module, such as an interface of a GSM module.
[0177] The present application may further include a display 2105 and an input unit 2104 and the like.
[0178] Figure 4 The structure of the terminal shown does not constitute a limitation on the terminal in the embodiment of the present application. In actual applications, the terminal may include Figure 4 More or fewer components than shown, or combinations of certain components.
[0179] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the system provided in the embodiment, since it corresponds to the method provided in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0180] Specific examples are used herein to illustrate the principles and implementation methods of the present application, and the description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
[0181] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. A method for calculating a website fingerprint, characterized in that: include: Obtaining a website sample, and determining a target website from the website sample; Calculate the structure vector value of the document object model corresponding to the target website; Crawling the target website, obtaining a static resource list, parsing the static file resource list of the target website, and outputting a website list corresponding to each static resource in the website sample; Parsing the characteristic fields of the target website; Classifying the websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determining the sample website; Using the characteristics of the example website as the website fingerprint; The websites are classified according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and the example websites are determined to include: According to the structure vector value, the website list corresponding to each static resource and the characteristic field, the websites in the website sample are analyzed for association and classified, and any original website in each class has at least one similar website, and the original website and the similar website have at least two of the structure vector value, the website list corresponding to each static resource and the characteristic field pair website in common; At least one example website is identified in each category of websites.
2. The method for calculating website fingerprint according to claim 1, characterized in that: The calculating of the structure vector value of the document object model corresponding to the target website includes: Obtain the target website html page and construct the document object model; In the document object model, a parent node is selected as a head element as a target node, and the element name and attribute of each target node are concatenated into a string; Calculate the hash value of the string, and multiply the hash value by the weight of the target node to obtain the weight value corresponding to the target node; wherein, the greater the node depth of the target node, the more nodes are the same as the target node, and the smaller the weight of the target node; Accumulate the weight values of all target nodes to get the structure vector value.
3. The method for calculating website fingerprint according to claim 1, characterized in that: The static file resource list of the target website to be parsed includes: Preprocessing the static resources in the static file resource list to remove characteristic information of public library resources and static resources; Constructing a static resource dictionary, calculating a static hash value for adjacent static file resource names using a preset formula, and establishing a mapping relationship between the static hash value, the static file resource name list, and the web page address corresponding to the static file resource; Calculate the hash value of each static file resource name in the static file resource list to obtain a hash value list corresponding to the static file resource list; The static hash value is calculated by using a preset formula for adjacent static file resource names; Determine whether the static resource dictionary contains the static hash value; If so, it is determined that there is an intersection between the static file resource lists of the target website and other websites, and the web page address of the target website is added to the list of web page addresses corresponding to the static file resources; If not, save the static hash value and the corresponding static file resource name list, and the web page address corresponding to the static file resource.
4. The method for calculating website fingerprint according to claim 3, characterized in that: The preset formula is: ; Where i is the number of adjacent static file resources obtained in each calculation and i is greater than 2, j is the index number of the first static file resource in the static file resource list among the static file resources obtained in each calculation, and k is an iterative traversal, which is used to traverse all static resources with index numbers in the interval [j, j + i - 1]. is a static hash value, It is the hash value corresponding to the static file resource.
5. The method for calculating website fingerprint according to claim 3, characterized in that: Preprocessing the static resources in the static file resource list to remove the characteristic information of the public library resources and the static resources includes: Configure path blacklist and / or file name blacklist for public library resources; Delete at least one of the version number and the random number in the static file resource name, and remove the domain name or IP address in the path corresponding to the static resource.
6. The method for calculating website fingerprint according to claim 1, characterized in that: When identifying sample sites, also include: Discard sample websites that do not have corresponding similar websites.
7. A website fingerprint calculation system, characterized in that: include: A website acquisition module, used to acquire website samples and determine a target website from the website samples; A structure vector value calculation module, used to calculate the structure vector value of the document object model corresponding to the target website; A static resource analysis and calculation module is used to crawl the target website, obtain a static resource list, parse the static file resource list of the target website, and output a website list corresponding to each static resource in the website sample; A feature field acquisition module, used to parse the feature fields of the target website; An association analysis module, used to classify websites according to the structure vector value, the website list corresponding to each static resource and the characteristic field, and determine the example website; A fingerprint calculation module, used for taking the characteristics of the example website as the website fingerprint; The association analysis module is a module for performing the following steps: According to the structure vector value, the website list corresponding to each static resource and the characteristic field, the websites in the website sample are analyzed for association and classified, and any original website in each class has at least one similar website, and the original website and the similar website have at least two of the structure vector value, the website list corresponding to each static resource and the characteristic field pair website in common; At least one example website is identified in each category of websites.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for calculating a website fingerprint as described in any one of claims 1 to 6 are implemented.
9. A terminal, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the method for calculating the website fingerprint according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and device for detecting website bugs
CN103632100A
Network crawler based website fingerprint information scanning method and apparatus
CN109376291A