Web Application Type Recognition Method, Device and System Based on Multidimensional Digital Features
By automatically extracting the header and Body fields of the Web response page, building a benchmark database, and using the improved Simhash and TF-IDF algorithms to generate digital feature vectors, solving the problem of manual rule extraction in the existing technology, and achieving automated recognition and accuracy of Web application types.
Patent Information
- Application Number
- CN202211649809.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-12-21
AI Technical Summary
The existing web application type recognition method relies on manual extraction of regular rules, which has problems such as difficulty in extracting rules, difficulty in updating rule databases, and many manual dependencies.
Using a multi-dimensional digital feature method, the header and Body fields of the web response page are automatically extracted, the benchmark database is constructed, and digital feature vectors are generated using improved Simhash and TF-IDF algorithms, clustering and similarity calculations are performed to realize automated web application type recognition.
It realizes automated recognition of web application types, get rid of the pain points of manual rule extraction, improves the accuracy and efficiency of recognition, and reduces manual intervention.
Smart Images

Figure CN116010257B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cyberspace mapping and network entity fingerprint recognition, and more specifically, relates to a method, device and system for identifying Web application types based on multi-dimensional digital features. Background Art
[0002] Web application type identification refers to the process of judging and identifying the application type and version that provides Web services, and is an important aspect of fingerprint identification. Web service application type refers to the product model and version number of the specific application provided by the Web service, such as the specific model and version number of the type of Web application such as firewall, VPN, router, mail system, etc. Web application type identification is an important part of Web vulnerability exploitation, cyberspace mapping, and network security assessment entity identification. Through Web application type identification, the model, version and other information of the Web application can be determined, which helps to select vulnerability verification tools and payloads for vulnerability verification and security assessment, and achieve rapid vulnerability discovery and security response.
[0003] In the application field, the current mainstream Web application type identification method is to manually extract regular rules of text information that can represent the uniqueness of Web response content, build a Web application response text information rule database, and use regular matching methods to identify and judge Web application types. Mainstream Web application type identification tools such as Nmap, WhatWeb, and Wappalyzer all use this method to implement Web application type fingerprint identification.
[0004] However, the Web application type identification method based on regular rule matching requires manual extraction of precise regular rules to identify the content of Web pages in order to achieve the purpose of type identification. There are pain points such as difficulty in extracting regular rules, difficulty in updating the rule library, small number of rule libraries, and high reliance on manual labor. Summary of the invention
[0005] In view of the above defects or improvement needs of the prior art, the present invention provides a method, device and system for identifying the type of Web application based on multi-dimensional digital features, thereby solving the technical problems existing in the existing methods for identifying the type of Web application, such as difficulty in extracting regular rules, difficulty in updating the rule base, small number of rule bases and high reliance on manual work.
[0006] To achieve the above object, according to one aspect of the present invention, a method for identifying Web application types based on multi-dimensional digital features is provided, comprising:
[0007] Benchmark database establishment phase:
[0008] Extract the Header fields and Body fields of N pages in a Web application, and preprocess them respectively to obtain the first digital feature vector and the second digital feature vector;
[0009] Use the first digital feature vector and the second digital feature vector as the digital feature vectors to be processed respectively to perform database building operations, obtaining the first Web benchmark database and the second Web benchmark database;
[0010] The database building operation includes: dividing all digital feature vectors to be processed into a training set and a data set; clustering the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculating the F1 score of each center point vector using the digital feature vectors in the test set; storing the relationship between the F1 score, radius corresponding to each center point vector and the type of the Web application in the Web benchmark database;
[0011] Application stage:
[0012] Initiate an HTTP request to the Web application to be detected, and obtain the third digital feature vector corresponding to the Header field and the fourth digital feature vector corresponding to the Body field in its response;
[0013] Determine the first target center point vector and its F1 score in the first Web benchmark database; the Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; determine the second target center point vector and its F1 score in the second Web benchmark database; the Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius;
[0014] Take the Web application type to which the target center point vector with the larger F1 score belongs as the detection result.
[0015] According to the second aspect of the present invention, there is provided a Web application type recognition device based on multi-dimensional digital features, including:
[0016] A benchmark database generation unit, configured to extract the Header fields and Body fields of N pages in a Web application, and preprocess them respectively to obtain the first digital feature vector and the second digital feature vector;
[0017] Use the first digital feature vector and the second digital feature vector as the digital feature vectors to be processed respectively to perform database building operations, obtaining the first Web benchmark database and the second Web benchmark database;
[0018] The database building operation includes: dividing all digital feature vectors to be processed into a training set and a data set; clustering the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculating the F1 score of each center point vector by using the digital feature vectors in the test set; storing the relationship between the F1 score, radius corresponding to each center point vector and the type of the Web application in the Web benchmark database;
[0019] An identification unit is configured to send an HTTP request to the Web application to be detected, and obtain a third digital feature vector corresponding to the Header field and a fourth digital feature vector corresponding to the Body field in its response;
[0020] Determine the first target center point vector and its F1 score in the first Web benchmark database; the Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; determine the second target center point vector and its F1 score in the second Web benchmark database; the Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius;
[0021] Take the Web application type to which the target center point vector with a larger F1 score belongs as the detection result.
[0022] According to a third aspect of the present invention, there is provided a Web application type identification system based on multi-dimensional digital features, including: a computer-readable storage medium and a processor;
[0023] The computer-readable storage medium is used to store executable instructions;
[0024] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method as described in the first aspect.
[0025] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects can be achieved:
[0026] 1. The Web application type recognition method based on multi-dimensional digital features provided by the present invention first realizes the automatic extraction of digital features of Web pages by automatically extracting keywords and mapping multi-dimensional vectors from the two dimensions of the Header and Body fields of the Web response page. Secondly, in the Web application type recognition method, a digital feature fingerprint benchmark library of Web applications is constructed, and the benchmark center point vectors, detection radii, and result validity evaluation F1 scores of the Header and Body fields of Web pages of known application types are generated. The fingerprint recognition result is obtained by comparing the digital features of different dimensions of the Web page with the digital feature benchmark library. Finally, the fingerprint recognition results of multi-dimensional digital features are merged by comparing the sizes of the result validity F1 scores, and the final fingerprint recognition result is determined. The present invention realizes the automatic generation of digital feature fingerprints of Web applications, getting rid of the pain point of the difficult manual extraction of fingerprint rules in conventional methods.
[0027] 2. The Web application type recognition method based on multi-dimensional digital features provided by the present invention takes into account the problem that the traditional Simhash algorithm cannot discover the internal information of feature words. However, for feature words with large weights, even changing one character will cause a great change in the digital features of the page. Therefore, further, the present invention improves the traditional Simhash algorithm and proposes a multiple Simhash algorithm for preprocessing fields to generate digital feature vectors corresponding to the fields. The multiple Simhash algorithm changes the hash algorithm used in the traditional Simhash algorithm from a conventional random hash algorithm to the Simhash algorithm, further filters and segments the input text feature words, and performs operations such as hashing, weighting, merging, and dimensionality reduction on these segments. The obtained hash value is used as the output of the entire hash, which can make full use of the information given by the feature words and effectively avoid the problem of feature word modification.
[0028] 3. The Web application type recognition method based on multi-dimensional digital features provided by the present invention improves the TF-IDF algorithm and proposes a TF-IDF-M algorithm model for generating a unique digital feature vector of the Web response page, which can realize the generation of feature vectors of text information in the Web response page. In the present invention, the word length information M is introduced into the TF-IDF calculation process, which can distinguish the influence of word length on TF-IDF, increase the information volume of TF-IDF, and help improve the accuracy of TF-IDF. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is one of the flowcharts of the Web application type recognition method based on multi-dimensional digital features provided by the embodiments of the present invention;
[0030] Figure 2 The second flowchart of the Web application type recognition method based on multi-dimensional digital features provided by the embodiments of the present invention;
[0031] Figure 3 The flowchart for generating the benchmark digital feature library of Web pages of known application types provided by the embodiments of the present invention;
[0032] Figure 4 The flowchart for generating the recognition result of the Web application type provided by the embodiments of the present invention;
[0033] Figure 5 The flowchart for generating the digital features of the Body / Header fields provided by the embodiments of the present invention;
[0034] Figure 6 The structural diagram of the Web application type recognition device based on multi-dimensional digital features provided by the embodiments of the present invention. Specific embodiments
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0036] The embodiments of the present invention provide a Web application type recognition method based on multi-dimensional digital features, as Figure 1-2 shown, including:
[0037] Benchmark database establishment stage:
[0038] Extract the Header fields and Body fields of N pages in each type of known Web application, and preprocess them respectively to obtain the first digital feature vector and the second digital feature vector;
[0039] Use the first digital feature vector and the second digital feature vector as the digital feature vectors to be processed respectively to perform the database building operation, and obtain the first Web benchmark database and the second Web benchmark database;
[0040] The database building operation includes: dividing all the digital feature vectors to be processed into a training set and a data set; clustering the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculating the F1 scores of each center point vector using the digital feature vectors in the test set; storing the relationship between the F1 scores, radii corresponding to each center point vector and the type of the Web application in the Web benchmark database;
[0041] Application stage:
[0042] Send an HTTP request to the Web application to be detected, and obtain the third digital feature vector corresponding to the Header field and the fourth digital feature vector corresponding to the Body field in its response (i.e., the response page);
[0043] Determine the first target center point vector and its F1 score in the first Web reference database; the Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; determine the second target center point vector and its F1 score in the second Web reference database; the Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius;
[0044] Select the Web application type to which the target center point vector with the larger F1 score belongs among the first target center point vector and the second target center point vector as the detection result.
[0045] Specifically, as Figure 3 shown, in the reference database establishment stage, construct a digital feature reference fingerprint library for the Header and Body fields of the Web application responses of known application types, including the following steps:
[0046] S101, Basic dataset construction: Collect the Header and Body field text contents of N (N>200) pages for each Web application of known application types respectively to construct a basic dataset;
[0047] S102, Generation of known application type page Body / Header digital feature vectors: Process the text content of the Header field in the dataset to generate the Header digital feature vectors of the known application type pages; process the text content of the Body field in the dataset to generate the Body digital feature vectors of the known application type pages;
[0048] S103, Training dataset construction: Divide the Header / Body digital feature vector sets of each known type of Web page into a digital feature vector training set and a digital feature vector test set according to a preset ratio (for example: 6:4);
[0049] S104, Generation of the center point vector and detection radius of the Header / Body digital feature vectors of known application type pages: Cluster the digital feature vectors in the training set to obtain the center point vectors and radii of K clusters.
[0050] Preferably, the clustering includes:
[0051] The DBSCAN clustering algorithm is used to cluster the digital feature vectors in the training set to obtain multiple clustering clusters, and the number of the multiple clustering clusters is used as the K value of the K-Means clustering algorithm. The K-Means clustering algorithm based on the Hamming distance is used to obtain the center point vectors and radii of K clusters.
[0052] Specifically, the DBSCAN clustering algorithm is used to cluster the training set (other density clustering algorithms can also be used) to obtain the number of clustering clusters (i.e., the number of clusters) K of the digital feature vectors of the known type Web page Header / Body; K is used as the K value parameter of K-Means, and the K-Means algorithm based on the Hamming distance is used to cluster the training set to obtain the center point vectors and radii of K clusters; the center point vectors and cluster radii of each cluster are extracted as the center point vectors and detection radii of the digital feature vectors of the Header / Body of the known application type page.
[0053] S105. Verify the center point vectors and detection radii and generate the F1 score of the digital feature effectiveness: Use the Hamming distance ranging algorithm to calculate the similarity bit by bit for the center point vectors and detection radii of the K different Web application type vector clusters generated for the vectors in the test set to verify the effectiveness of the center point vectors and detection radii, and at the same time generate the F1 score.
[0054] It can be understood that for a Web application of a known application type, there are usually multiple different versions of pages. Therefore, the more versions can be collected in step S101, the more comprehensive the fingerprint database will be. At the same time, after the fingerprint database is established, the center point vectors and detection radii of each known type of Web application in the fingerprint database can be verified through the test set to check whether the pages of different versions of the Web applications collected in step S101 are comprehensive. The specific verification method is: if the Hamming distance between the digital feature vectors in the test set and any center point vector is within the detection radius corresponding to the center point vector and exceeds a preset ratio (for example: 90%), it is considered valid.
[0055] Furthermore, the calculation method of the F1 score of each center point vector is:
[0056] Calculate the Hamming distance between the digital feature vectors in the test set and each center point vector, and obtain the F1 score of each center point vector according to the radius corresponding to each center point vector.
[0057] Since for each type of Web application, the fingerprint recognition process is actually a binary classification process, that is, whether a Web page is or is not of this type of Web application, therefore, precision, recall, and F1-score are selected to measure the effect of fingerprint recognition results during the evaluation process in the test set. Among them, precision is for the recognition results, and its meaning is the probability that the actual positive samples are among all the samples recognized as positive in the test set; recall is for the original samples, and its meaning is the probability that the samples recognized as positive are among the actual positive samples in the test set. The F1-score calculation formula is as follows:
[0058]
[0059] S106, generate a digital feature fingerprint library for known type Web benchmark pages: Store the center point vector, detection radius, effective F1-score, and page type label of the Header / Body digital feature vectors of the known application type pages obtained in step S105 into the Web benchmark page Header / Body digital feature database to obtain the established Header digital feature database (i.e., the first Web benchmark database) and Body digital feature database (i.e., the second Web benchmark database).
[0060] In the application stage, as Figure 4 shown, perform application type recognition on the Web application to be detected based on the established first Web benchmark database and second Web benchmark database, including the following steps:
[0061] S201, obtain the text content of two dimensions of the Header / Body fields of the page to be detected: Initiate an HTTP request to the Web application to be detected and recognized, and obtain the text content of the Header fields and Body fields in its response.
[0062] S203, generate digital features of the Header / Body page text: Preprocess the Header information and the text content of the Body field to generate a digital feature vector of the Header field of the page to be detected (i.e., the third digital feature vector) and a digital feature vector of the Body field (i.e., the fourth digital feature vector).
[0063] S203, fingerprint recognition based on page digital features: Use the digital feature vector to calculate the similarity bit by bit with the center point vector and detection radius of the digital features of the known type Web applications in the Web benchmark database using the Hamming distance ranging algorithm to determine the Web application to which the page digital feature vector belongs, and obtain the fingerprint recognition result based on the similarity.
[0064] Specifically, the third digital feature vector and the center point vector in the first Web benchmark database are used to calculate the similarity bit by bit using the Hamming distance ranging algorithm; among them, if the Hamming distance between the third digital feature vector and any center point vector in the first Web benchmark database is less than the detection radius of the center point vector, it is considered that the third digital feature vector is similar to the center point vector, and their Web application types are the same, so as to obtain the fingerprint result identified using the Header field; similarly, if the Hamming distance between the fourth digital feature vector and any center point vector in the second Web benchmark database is less than the detection radius of the center point vector, it is considered that the fourth digital feature vector is similar to the center point vector, and their Web application types are the same, so as to obtain the fingerprint result identified using the Body field.
[0065] S204, Recognition result merging and fingerprint result generation: Compare and merge the fingerprint results identified using the Header and Body fields based on the fingerprint validity evaluation parameter F1 score, and take the fingerprint result with the larger F1 score as the final fingerprint recognition result.
[0066] S205, Storage of Web application page digital features and application type labels: After generating the fingerprint recognition result based on the page digital features, store the Header digital features, Body field digital features and the determined application type labels of this result into the corresponding Web application entity nodes in the database.
[0067] Furthermore, the improved SimHash algorithm is used to preprocess the fields to obtain the digital feature vectors corresponding to the fields, that is, for the extracted Header and Body fields, the improved multiple Simhash fuzzy hashing algorithm is respectively used to generate the digital feature vectors of the Header field and the Body field, including the following steps:
[0068] The fields are successively subjected to word segmentation processing and stop word filtering to obtain the keywords of the fields;
[0069] After filtering the special characters of the keywords, word segmentation is performed again to obtain the updated keywords; the updated keywords are respectively converted into hash value strings represented in binary and then successively subjected to weighting, merging and dimensionality reduction processing to generate digital feature vectors.
[0070] Furthermore, as Figure 5 shown, after successively performing word segmentation processing and stop word filtering on the fields, it further includes:
[0071] Calculate the TF-IDF value or TF-IDF-M value of each keyword in the keyword list, and retain only the top N keywords with larger TF-IDF values or TF-IDF-M values;
[0072] During the weighting process, the TF-IDF value or TF-IDF-M value of each updated keyword is used as its weight;
[0073] The TF-IDF-M value is the product of the number of characters, word frequency and reverse document frequency of the keyword.
[0074] First, let's explain the traditional Simhash algorithm. The traditional Simhash algorithm can be divided into 5 steps:
[0075] 1) Word segmentation and stop word filtering: Use word segmentation methods such as NLTK and Jieba to segment the text and select a fixed number of feature words (i.e. keywords) to obtain the word segmentation results; filter the word segmentation results with stop words, delete the interfering words in the word segmentation results, and obtain a keyword list; give the weight corresponding to each feature word in the keyword list. Usually, the weight of each word can be set to its TF-ID value.
[0076] Preferably, further screening can be performed based on the TF-ID value of each feature word to extract the optimal keyword list that best represents the Header and Body field contents: select the top N feature words with larger TF-ID values in the keyword list to obtain the optimal keyword list.
[0077] 2) Hash: Each keyword in the optimal keyword list is converted into a hash value represented in binary through a hash algorithm.
[0078] 3) Weighting: Change the 0 in the binary hash value in 2) to -1, and then multiply each digit by the weight of the corresponding word to get the weighted feature string corresponding to each word. For example, if the binary hash value of the first word is "1010" and the weight is 5, first change "1010" to "1-11-1", and then multiply each digit by 5 to get "5-55-5".
[0079] 4) Merge: Add the corresponding bits of the weighted feature string corresponding to each feature word of a document. For example, the first document has 3 feature words, and the corresponding weighted feature strings are "5-55-5", "33-33" and "-6-66-6", then the result after merging is "2-88-8". The first bit "2" is obtained by adding the first bit "5" of the first feature string, "3" of the second feature string and "-6" of the third feature string, and the rest of the bits are similar.
[0080] 5) Dimensionality reduction: For each digit in the result obtained in step 4, if the digit is greater than 0, change it to 1, otherwise change it to 0 to obtain a "01" string. For example, if the result obtained in the fourth step is "2-88-8", the result after dimensionality reduction is "1010".
[0081] Thus, it can be seen that the digital feature vectors of both the Header header field and the Body field are "01" strings.
[0082] The method for determining the similarity between two documents is as follows: Calculate the Hamming distance between the "01" strings obtained by the two documents through the traditional Simhash algorithm, that is, the number of different corresponding bits between the two "01" strings. The smaller the distance, the more similar the two vectors are, and the larger the distance, the more different the two vectors are.
[0083] The method for calculating the similarity between the digital feature vector of the Header header / Body field and the center point vector is as follows: Determine whether the Hamming distance between the digital feature vector of the Header header / Body field and the center point vector is less than the detection range of the center point vector. If so, they are similar; if not, they are not similar.
[0084] In practice, it is found that the traditional Simhash algorithm has the problem of being unable to discover the internal information of feature words. For feature words with a large weight, even changing one character will cause a large change in the digital features of the page. Therefore, further, the present invention improves the traditional Simhash algorithm and proposes a multiple Simhash algorithm.
[0085] The difference between the multiple Simhash algorithm and the traditional Simhash algorithm is reflected in step 2). In order to make full use of the information given by the feature words, the multiple Simhash algorithm changes the hash algorithm used in the hash process in step 2) of the traditional Simhash algorithm from a conventional random hash algorithm to the Simhash algorithm. That is, after filtering the input feature words by special characters (such as. / = _), further word segmentation is performed to achieve as fine a granularity of word segmentation as possible, and then these segmented words are subjected to operations such as hashing, weighting, merging, and dimensionality reduction in the traditional Simhash to obtain the hash value as the output of the entire step 2). In this way, the dual Simhash algorithm is obtained.
[0086] Further, by making the same changes to step 2 of the second-level Simhash, a triple Simhash algorithm can be obtained, and so on, an N-level Simhash algorithm can also be obtained, where N is a positive integer representing the multiplicity of the Simhash algorithm. In practical applications, its value can be adjusted according to the fineness of the word segmentation result and the accuracy requirement of the algorithm.
[0087] In addition, it is found in the keyword generation experiment that the more characters a feature word has, the more representative it is. Based on this, in the word segmentation step of the multiple Simhash algorithm, the optimal keyword list that best represents the content of the Header and Body fields is extracted based on the TF-IDF-M value of each word in the keyword list; in the weighting step, the weight of each word is its corresponding TF-IDF-M value.
[0088] That is, preferably, after the field is sequentially subjected to word segmentation processing and stop word filtering, it further includes: calculating the TF-IDF-M value of each keyword in the keyword list, and only retaining the top N keywords with larger values among them;
[0089] When performing weighting processing, the TF-IDF-M value of each updated keyword is used as its weight;
[0090] The TF-IDF-M value is the product of the number of characters, term frequency, and inverse document frequency of the keyword.
[0091] Specifically, the calculation method of the keyword TF-IDF-M is as follows: Denote the total number of documents as N, and the N documents are d1, d2, …, d N , and the number of words in the N documents are n1, n2, …, n N , d i The jth word in is w ij , w ij The number of times it appears in d i is t ij , d1, d2, …, d N The number of documents containing w ij is c ij . The calculation formula of the term frequency (TF) is:
[0092]
[0093] Where TF ij represents the term frequency of w ij .
[0094] The calculation formula of the inverse document frequency (IDF) is:
[0095]
[0096] Where IDF ij represents the inverse document frequency of w ij .
[0097] The calculation formula of the TF-IDF value is:
[0098] TF-IDFij = TF ij × IDF ij
[0099] where TF-IDF ij represents the TF-IDF value of w ij .
[0100] In the keyword generation experiment, it is found that the more characters a feature word has, the more representative it is. Therefore, the present invention modifies the TF-IDF calculation formula by adding a correction term M ij , M ij = lg(len(w ij ))), where len(w ij ) represents the number of characters contained in w ij . After adding the correction term, the calculation formula of the new TF-IDF-M is:
[0101] TF-IDF-M ij = TF ij × IDF ij × M ij
[0102] The fingerprint recognition method provided by the present invention is compared with common fingerprint recognition tools WhatWeb and Wappalyzer. The present invention has advantages in the precision rate and recall rate of the Web application type recognition result, and has a good effect in the Web application type recognition.
[0103] In summary, the present invention uses the K-Means and DBSCAN clustering algorithms to generate the central point vectors of the digital features of the known Web application type pages, the detection radius, and the effective row evaluation F1 score, and on this basis, realizes the similarity detection of unknown pages; finally, the invention starts from two dimensions of the Header header and the Body field in the Web application response page, respectively performs similarity detection on the digital feature vectors of the Header header and the Body field in the Web application response page, and combines the different results obtained based on the F1 score. The present invention realizes the automatic generation of the digital feature fingerprints of Web applications, proposes a new method for Web application type recognition, gets rid of the pain point of the difficult manual extraction of fingerprint rules in the conventional method, and has strong practicability.
[0104] Next, the Web application type recognition device based on multi-dimensional digital features provided by the present invention will be described. The Web application type recognition device based on multi-dimensional digital features described below can be mutually corresponding and referred to the Web application type recognition method based on multi-dimensional digital features described above.
[0105] An embodiment of the present invention provides a Web application type recognition device based on multi-dimensional digital features, including:
[0106] A reference database generation module, configured to extract the Header fields and Body fields of N pages in a Web application, and preprocess them respectively to obtain a first digital feature vector and a second digital feature vector;
[0107] Perform a database building operation using the first digital feature vector and the second digital feature vector as digital feature vectors to be processed respectively, to obtain a first Web reference database and a second Web reference database;
[0108] The database building operation includes: dividing all digital feature vectors to be processed into a training set and a data set; clustering the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculating the F1 scores of each center point vector using the digital feature vectors in the test set; storing the relationship between the F1 score, radius corresponding to each center point vector and the type of the Web application in the Web reference database;
[0109] An identification module, configured to initiate an HTTP request to a Web application to be detected, and obtain a third digital feature vector corresponding to the Header field in its response and a fourth digital feature vector corresponding to the Body field;
[0110] Determine a first target center point vector and its F1 score in the first Web reference database; the Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; determine a second target center point vector and its F1 score in the second Web reference database; the Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius;
[0111] Take the Web application type to which the target center point vector with a larger F1 score belongs as the detection result.
[0112] Specifically, as Figure 6 shown, the Web application type recognition device based on multi-dimensional digital features includes five main functional units
[0113] ① A Web page response acquisition unit 100, which is mainly used to acquire the page text content of the Web application response;
[0114] The Web page response acquisition unit 100 acquires the text content of the Header and Body fields of the page response of the Web application through an HTTP request, and inputs the obtained result to the Web page digital feature generation unit 200 for processing.
[0115] ②The Web page digital feature generation unit 200 generates digital feature vectors for the Header and Body field information of the Web application response content using the TF-IDF-Hash algorithm;
[0116] The Web page digital feature generation unit 200 includes a Header digital feature generation module 210 and a Body digital feature generation module 220.
[0117] The Header digital feature generation module 210 refers to a functional program that receives the Web response Header text information input by the Web page response acquisition unit 100 and generates digital features therefrom.
[0118] The Body digital feature generation module 220 refers to a functional program that receives the Web response Body field text information input by the Web page response acquisition unit 100 and generates digital features therefrom.
[0119] The Web page digital feature generation unit 200 receives the Web page text content input by the Web page response acquisition unit 100, extracts the Header and Body field text information therefrom, respectively extracts digital features from the text information, and inputs the extracted digital features to the Web application type recognition unit 300 based on page digital features for Web application type recognition.
[0120] ③The known type Web reference page digital feature fingerprint library generation unit 300 mainly constructs training sets and test sets of digital feature vectors for the Headers / Bodies of Web pages of known application types, calculates the reference digital feature center point vectors, detection radii for similarity detection, and the F1 scores for evaluating the effectiveness of the detection results of the reference digital feature center point vectors for the Headers and Bodies of Web reference pages of known application types, and stores the center point vectors, detection radii, and F1 scores of the Headers and Bodies of Web pages of known application types in the reference page digital feature fingerprint library for fingerprint recognition of Web applications based on digital features.
[0121] The known type Web reference page digital feature fingerprint library generation unit 300 includes a Header vector center point and radius generation module 310, a Body vector center point and radius generation module 320, a Header reference digital feature fingerprint library 330, and a Body reference digital feature fingerprint library 340.
[0122] The Header vector center point and radius generation module 310 refers to a functional program that constructs a training set and a test set from the Header vector data of known application type Web pages, and generates the center point vector and detection radius of the known type Header through DBSCAN and K-Means clustering methods.
[0123] The Body vector center point and radius generation module 320 refers to a functional program that constructs a training set and a test set from the Body vector data of known application type Web pages, and generates the center point vector and detection radius of the known type Body fields through DBSCAN and K-Means clustering methods.
[0124] The Header reference digital feature fingerprint library 330 refers to a database used to store the digital feature fingerprint information and application types of the page Headers of known type Web applications. The information stored includes the center point vector, detection radius, F1 score used to evaluate the effectiveness of similarity recognition, and application type label information of the Headers of known type Web applications, and is used for Header digital feature fingerprint recognition.
[0125] The Body reference digital feature fingerprint library 340 refers to a database used to store the digital feature fingerprint information and application types of the Body fields of known type Web applications. The information stored includes the center point vector, detection radius, F1 score used to evaluate the effectiveness of similarity recognition, and application type label information of the Body fields of known type Web applications, and is used for Body field digital feature fingerprint recognition.
[0126] The known type Web reference page digital feature fingerprint library generation unit 300 stores the digital feature fingerprints of the Headers and Body fields of the response pages of known type Web applications and the corresponding Web application type labels, and is used to compare the digital features of the Headers and Body fields of the page to be tested for similarity, so as to realize Web application type recognition.
[0127] ④ The Web application type recognition unit 400 based on page digital features calculates the similarity between the digital feature vectors in two dimensions of the Header and Body fields generated for the page to be tested and the center point vectors and detection radii of the Headers and Body fields of the fingerprints in the reference digital feature fingerprint library respectively using the Hamming distance algorithm based on bit operations, realizes the determination and classification of Web page similarity, and at the same time merges the results of Header and Body fingerprint recognition based on the F1 score.
[0128] The Web application type recognition unit 400 based on page digital features includes a Header fingerprint recognition module 410, a Body fingerprint recognition module 420, and a result merging module 430.
[0129] The Header fingerprint recognition module 410 refers to a program that performs a similarity comparison of Header digital features. Its function is to perform a similarity comparison between the page Header digital features generated by the Web page Header digital feature generation module 210 and the digital features in the Header digital feature fingerprint data set 330, and obtain the application type information of the page Header.
[0130] The Body fingerprint recognition module 420 refers to a program that performs a similarity comparison of Body digital features. Its function is to perform a similarity comparison between the page Body digital features generated by the Web page Body digital feature generation module 220 and the digital features in the Body digital feature fingerprint data set 340, and obtain the application type information of the page Body.
[0131] The result merging module 430 refers to a program that effectively merges the fingerprint recognition results of the page Header digital features and the Body digital features. Its function is to effectively merge the results of the Header and Body digital feature fingerprint recognition according to the F1 score of the fingerprint recognition, and take the fingerprint result corresponding to the larger F1 score as the final application type recognition result.
[0132] The Header fingerprint recognition module 410 receives the Header digital features generated by the Web page Header digital feature generation module 210, performs a similarity comparison between them and the digital feature fingerprints in the Header digital feature fingerprint data set 330, and obtains the application type result of the page Header based on page similarity; the Body fingerprint recognition module 420 receives the Body digital features generated by the Web page Body digital feature generation module 220, performs a similarity comparison between them and the digital feature fingerprints in the Body digital feature fingerprint data set 340, and obtains the application type result of the page Body field based on page similarity; the recognized Header fingerprint result and Body field fingerprint result are input to the result merging module 430, and are effectively merged according to the F1 score of the fingerprint recognition, and the fingerprint result corresponding to the larger F1 score is taken as the final application type recognition result; the result merging module 430 inputs the recognized final result to the Web application type recognition result data storage unit 500 for storage.
[0133] ⑤The Web application type recognition result data storage unit 500 is used to store the recognized Web application type entity data, including Web application page response data, Header and Body field digital feature information, Web application types, etc.
[0134] Its working mechanism is as follows: The Web page response acquisition unit 100 obtains the Web application response text content through an HTTP request and passes it to the Web page digital feature generation unit 200; after receiving the incoming page response content, the Web page digital feature generation unit 200 uses the Header digital feature extraction module 210 and the Body digital feature extraction module 220 to extract and generate the digital features of the Header and Body fields in the Web page; the known type Web reference page digital feature fingerprint library generation unit 300 uses the Header vector center point and radius generation module 310 to generate the Header reference digital feature fingerprint library 330, and uses the Body vector center point and radius generation module 320 to generate the Body reference digital feature fingerprint library 340; the Web application type recognition unit 400 based on page digital features compares the page digital features generated by the Web page digital feature generation unit 200 with the corresponding reference digital feature fingerprint data in the known type Web reference page digital feature fingerprint library 300 to recognize the Web application type, and stores the recognized application digital features and type label results in the Web application type recognition result data storage unit 500.
[0135] The Web application type recognition result data storage unit 500 refers to a database that stores the effective Web application type recognition results. The stored content includes the digital fingerprint information of the Web application type and the corresponding Header and Body fields in two dimensions, as well as the type label information of the Web application. Its function is to store the digital features and application types of the effective Web application type recognition results.
[0136] An embodiment of the present invention provides a Web application type recognition system based on multi-dimensional digital features, including: a computer-readable storage medium and a processor;
[0137] The computer-readable storage medium is used to store executable instructions;
[0138] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the Web application type recognition method based on multi-dimensional digital features as described in any of the above embodiments.
[0139] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying Web application types based on multi-dimensional digital features, characterized in that Including: Benchmark database establishment stage: Extract the Header fields and Body fields of N pages in the Web application, and preprocess them respectively to obtain the first digital feature vector and the second digital feature vector; Use the first digital feature vector and the second digital feature vector as the digital feature vectors to be processed respectively to perform the database building operation, and obtain the first Web benchmark database and the second Web benchmark database; The database building operation includes: dividing all digital feature vectors to be processed into a training set and a data set; Cluster the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculate the F1 score of each center point vector using the digital feature vectors in the test set; Store the relationship between the F1 score, radius corresponding to each center point vector and the type of the Web application in the Web benchmark database; Application stage: Initiate an HTTP request to the Web application to be detected, and obtain the third digital feature vector corresponding to the Header field and the fourth digital feature vector corresponding to the Body field in its response; Determine the first target center point vector and its F1 score in the first Web benchmark database; The Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; Determine the second target center point vector and its F1 score in the second Web benchmark database; The Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius; Take the Web application type to which the target center point vector with the larger F1 score belongs as the detection result; The calculation method of the F1 score of each center point vector is: Calculate the Hamming distance between the digital feature vectors in the test set and each center point vector, and obtain the F1 score of each center point vector according to the radius corresponding to each center point vector.
2. The method according to claim 1, characterized in that, The clustering includes: Use the DBSCAN clustering algorithm to cluster the digital feature vectors to obtain multiple clustering clusters, take the number of the multiple clustering clusters as the K value of the K-Means clustering algorithm, and use the K-Means clustering algorithm based on the Hamming distance to obtain K center point vectors and radii.
3. The method according to claim 1, characterized in that, Use an improved SimHash algorithm to preprocess the fields to obtain the digital feature vectors corresponding to the fields, including: Perform word segmentation and stop word filtering on the field in sequence to obtain a keyword list; Perform special character filtering on the keywords in the keyword list and then perform word segmentation again to obtain updated keywords; Convert the updated keywords into hash value strings represented in binary respectively, and then perform weighted, merged and dimensionality reduction processing in sequence to generate digital feature vectors.
4. The method according to claim 3, wherein After performing word segmentation processing and stop word filtering on the field in sequence, it also includes: Calculate the TF-IDF value or TF-IDF-M value of each keyword in the keyword list, and only retain the top N keywords with larger TF-IDF values or TF-IDF-M values; When performing weighted processing, use the TF-IDF value or TF-IDF-M value of each updated keyword as its weight; Among them, the TF-IDF-M value is the product of the character count, term frequency, and inverse document frequency of the keyword.
5. A Web application type recognition device based on multi-dimensional digital features, characterized in that, It includes: A reference database generation module, which is used to extract the Header field and Body field of N pages in the Web application, and perform preprocessing on them respectively to obtain a first digital feature vector and a second digital feature vector; Perform database building operations on the first digital feature vector and the second digital feature vector respectively as the digital feature vectors to be processed, and obtain a first Web reference database and a second Web reference database; The database building operation includes: dividing all digital feature vectors to be processed into a training set and a data set; Cluster the digital feature vectors in the training set to obtain the center point vectors and radii of multiple clusters, and calculate the F1 score of each center point vector using the digital feature vectors in the test set; store the relationship between the F1 score, radius corresponding to each center point vector, and the type of the Web application in the Web reference database; An identification module, which is used to initiate an HTTP request to the Web application to be detected, and obtain a third digital feature vector corresponding to the Header field in its response and a fourth digital feature vector corresponding to the Body field; Determine the first target center point vector and its F1 score in the first Web reference database; the Hamming distance between the first target center point vector and the third digital feature vector is less than its radius; determine the second target center point vector and its F1 score in the second Web reference database; the Hamming distance between the second target center point vector and the fourth digital feature vector is less than its radius; Use the Web application type to which the target center point vector with the larger F1 score belongs as the detection result; The calculation method of the F1 score of each center point vector is: Calculate the Hamming distance between the digital feature vectors in the test set and each center point vector, and obtain the F1 score of each center point vector according to the radius corresponding to each center point vector.
6. A Web application type recognition system based on multi-dimensional digital features, characterized in that, It includes: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Deep-learning-based intrusion detection method, system and computer program for web applications
US10778705B1
System and Method for Blockchain Automatic Tracing of Money Flow Using Artificial Intelligence
US20220067738A1