Web asset surveying and mapping method based on response structure identification

By building a variety of HTTP request payloads and a deep semantic feature extraction model based on RoBERTa model, combining the fingerprint recognition classification model of random forests and feedforward neural networks, the problems of high false positive rate and maintenance threshold of traditional Web asset fingerprint recognition methods are solved, and high-accurate Web asset mapping is achieved.

CN120017428AActive Publication Date: 2025-05-16GUANGZHOU JINHANG NETWORK TECH CO LTD +2

Patent Information

Application Number
CN202510502401.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The traditional web asset fingerprint recognition method based on rule matching has problems such as high false positive rate and high maintenance threshold, which is difficult to adapt to the rapidly iterative web application architecture, and it is difficult to extract effective semantic features.

Method used

Using a web asset mapping method based on response structure recognition, a deep semantic feature extraction model based on RoBERTa model is designed, and a coarse and fine-grained web server fingerprint recognition classification model is constructed by constructing multiple HTTP request payloads, obtaining response header data and performing standardized processing, and a deep semantic feature extraction model based on RoBERTa model is designed, and a coarse and fine-grained web server fingerprint recognition classification model is constructed based on random forests and feedforward neural networks.

Benefits of technology

It realizes the effective identification of web server asset fingerprints, improves the accuracy of identifying and mapping web assets, and adapts to the fast iterative web application architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017428A_ABST
    Figure CN120017428A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, in particular to a Web asset surveying and mapping method based on response structure recognition. Comprising the following steps: constructing five types and sixteen types of HTTP request loads, sending a request to a target domain name, obtaining response head data, carrying out standardization processing on the response head data, and constructing an original response head domain key sequence data set; carrying out fine tuning on the RoBERTa model through transfer learning and a fine tuning strategy, extracting deep semantic features of a response head domain key sequence, and generating a high-dimensional semantic embedding data set; carrying out coarse-grained fingerprint classification and identification on the Web asset server, and outputting coarse-grained fingerprint classification and identification data; performing fine-grained fingerprint classification and identification of the Web asset server, and outputting fine-grained fingerprint classification and identification data; by means of the method, the Web server asset fingerprint can be effectively recognized, and the Web asset recognition and surveying and mapping accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security, and in particular to a Web asset mapping method based on response structure recognition. Background Art

[0002] With the widespread use of Web application services, Web asset security faces severe challenges. These assets not only carry a large amount of personal privacy and industrial production data, but are also exposed to the Internet environment and become targets of malicious attackers. To deal with this threat, enterprise security personnel need to fully understand the number and composition of Web assets in the organization's Internet exposure, and carry out targeted reinforcement of assets with security risks.

[0003] Traditional rule-based Web asset fingerprinting methods have problems such as high false positive rate and high maintenance threshold, and are difficult to adapt to the rapidly iterating Web application architecture. In recent years, academia and industry have begun to pay attention to the identification features in the Web server response header, but these features are high-dimensional and have different text contents, making it difficult to extract effective semantic features.

[0004] Therefore, it is necessary to propose a Web asset mapping method that can effectively identify Web server asset fingerprints and improve the accuracy of identifying and mapping Web assets. Summary of the invention

[0005] The purpose of the present invention is to provide a Web asset mapping method based on response structure recognition, aiming to effectively identify Web server asset fingerprints and improve the accuracy of identifying and mapping Web assets.

[0006] To achieve the above object, the present invention adopts a Web asset mapping method based on response structure recognition, which includes the following steps: Construct five categories and sixteen types of HTTP request payloads, send requests to the target domain name, obtain response header data, and standardize it to build the original response header domain key sequence data set; Design a deep semantic feature extraction model based on RoBERTa model fine-tuning. Through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding dataset. Based on random forest, a coarse-grained Web server fingerprint recognition classification model is constructed to perform coarse-grained fingerprint classification recognition of Web asset servers and output coarse-grained fingerprint classification recognition data; A fine-grained Web server fingerprint recognition and classification model is constructed based on a feedforward neural network to perform fine-grained fingerprint classification and recognition of Web asset servers and output fine-grained fingerprint classification and recognition data.

[0007] Among them, in the steps of constructing five categories and sixteen types of HTTP request payloads, sending requests to the target domain name, obtaining response header data, and standardizing it, and constructing the original response header domain key sequence data set: The categories of request payload include basic request payload, newline variant payload, abnormal protocol and method payload, special method and header payload, and HTTP / 2 request payload.

[0008] Among them, in the step of designing a deep semantic feature extraction model based on RoBERTa model fine-tuning, fine-tuning the RoBERTa model through transfer learning and fine-tuning strategies, extracting deep semantic features of the response header domain key sequence, and generating a high-dimensional semantic embedding dataset: Use a byte-level byte pair encoding tokenizer to break down the HTTP response header field key sequence text into the smallest byte-level tokens; Calculate the frequency of character pairs, merge the character pairs with the highest frequency in each iteration, and gradually select the vocabulary to form the required high-frequency character pairs; Obtain a response dataset, learn the semantic features of the response header key sequence samples of the response dataset by fine-tuning the RoBERTa model, and construct a high-dimensional embedding that carries the semantic features of the original response header domain key sequence to generate a high-dimensional feature dataset.

[0009] Among them, in the step of building a coarse-grained Web server fingerprint identification classification model based on random forest, performing coarse-grained fingerprint classification identification of Web asset servers, and outputting coarse-grained fingerprint classification identification data: Build multiple decision trees and integrate them into an integrated classifier; The prediction results of all decision trees are integrated through majority voting method as the prediction result output; Normalize the high-dimensional semantic vectors in the high-dimensional feature dataset, search for the best hyperparameters, and output them.

[0010] Among them, in the step of building a fine-grained Web server fingerprint identification and classification model based on a feedforward neural network, performing fine-grained fingerprint classification and identification of Web asset servers, and outputting fine-grained fingerprint classification and identification data: The Softmax activation function is used to output the probability distribution of each fine-grained category fingerprint.

[0011] The present invention discloses a Web asset mapping method based on response structure recognition, which constructs five categories and sixteen types of HTTP request payloads, sends requests to target domain names, obtains response header data, and performs standardization processing on the response header data to construct an original response header domain key sequence data set; designs a deep semantic feature extraction model based on RoBERTa model fine-tuning, fine-tunes the RoBERTa model through transfer learning and fine-tuning strategies, extracts deep semantic features of the response header domain key sequence, and generates a high-dimensional semantic embedding data set; constructs a coarse-grained Web server fingerprint recognition and classification model based on random forest, performs coarse-grained fingerprint classification and recognition of Web asset servers, and outputs coarse-grained fingerprint classification and recognition data; constructs a fine-grained Web server fingerprint recognition and classification model based on a feedforward neural network, performs fine-grained fingerprint classification and recognition of Web asset servers, and outputs fine-grained fingerprint classification and recognition data; through the above-mentioned method, effective recognition of Web server asset fingerprints is achieved, and the accuracy of recognition and mapping of Web assets is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0013] Figure 1 It is a flow chart of the steps of the Web asset mapping method based on response structure recognition of the present invention.

[0014] Figure 2 It is a schematic diagram of a request payload carrying the OPTIONS method of the present invention.

[0015] Figure 3 It is a schematic diagram of a standard HEAD request and a simplified GET request using the HTTP / 1.1 protocol of the present invention.

[0016] Figure 4 It is a schematic diagram of a simple GET request of the present invention and a GET request using the HTTP / 1.0 protocol.

[0017] Figure 5 It is a schematic diagram of an HTTP / 1.0 protocol GET request with a Host header according to the present invention.

[0018] Figure 6 It is a schematic diagram of a HEAD request using \n as a line break character according to the present invention.

[0019] Figure 7 It is a schematic diagram of a HEAD request using \r as a line break character according to the present invention.

[0020] Figure 8 It is a schematic diagram of a HEAD request using the HTTP / 4.0 non-standard protocol of the present invention.

[0021] Fig. 9 It is a schematic diagram of an HTTP request using a non-existent SPOCK method according to the present invention.

[0022] Fig.10 It is a schematic diagram of the method for using the present invention and a mixed case HTTP request of the protocol.

[0023] Fig.11 It is a schematic diagram of an HTTP request using the OPTIONS method of the present invention.

[0024] Fig.12 It is a schematic diagram of a HEAD request carrying a Range header according to the present invention.

[0025] Fig.13 It is a schematic diagram of a HEAD request with a future time If-Modified-Since header according to the present invention.

[0026] Fig.14 It is a schematic diagram of a HEAD request with a past time If-Unmodified-Since header according to the present invention.

[0027] Fig.15 It is a schematic diagram of a HEAD request with a complex Accept header according to the present invention.

[0028] Fig.16 It is a schematic diagram of an HTTP request using the HTTP / 2 protocol of the present invention. DETAILED DESCRIPTION

[0029] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.

[0030] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0031] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0032] See also Figure 1 to Figure 16 The present invention provides a Web asset mapping method based on response structure recognition, comprising the following steps: S100: Construct five categories and sixteen types of HTTP request payloads, send requests to the target domain name, obtain response header data, and perform standardization on it to build an original response header domain key sequence data set.

[0033] In this embodiment, in the fingerprint identification work of the Web asset server, it is necessary to design a fuzz test request payload for the target Web server to collect the Web server response data packet information. On the one hand, considering the normal operation of the target asset, the request payload data packets are all request payloads such as GET and HEAD for the purpose of obtaining response data, which will not affect the target server. On the other hand, according to the different request payload types, a total of five categories and sixteen request payload test samples are adopted. The categories of request payloads include basic request payloads, newline variant payloads, abnormal protocol and method payloads, special methods and header payloads, and HTTP / 2 request payloads; among them, there are 4 basic request payloads, 2 newline variant payloads, 3 abnormal protocol and method payloads, 5 special methods and header payloads, and 2 HTTP / 2 request payloads. The original header field features of the target Web asset server response data packet are collected through the above-mentioned types of payloads, such as Figure 2 to Figure 16 As shown, Figure 2 This is a schematic diagram of the request payload carrying the OPTIONS method; Figure 3 This is a diagram of a standard HEAD request and a simplified GET request using the HTTP / 1.1 protocol; Figure 4 A schematic diagram of a simple GET request and a GET request using the HTTP / 1.0 protocol; Figure 5 This is a schematic diagram of an HTTP / 1.0 protocol GET request with a Host header; Figure 6 This is a diagram of a HEAD request using \n as a line break. Figure 7 This is a diagram of a HEAD request using \r as a line break. Figure 8 A schematic diagram of a HEAD request using the HTTP / 4.0 non-standard protocol; Fig. 9 This is a diagram of an HTTP request using a non-existent SPOCK method; Fig.10 A diagram of an HTTP request using mixed case for method and protocol; Fig.11 This is a schematic diagram of an HTTP request using the OPTIONS method; Fig.12 This is a schematic diagram of a HEAD request carrying a Range header; Fig.13 This is a diagram of a HEAD request with a future time If-Modified-Since header; Fig.14 This is a diagram of a HEAD request with a past time If-Unmodified-Since header; Fig.15 This is a diagram of a HEAD request with a complex Accept header. Fig.16 A schematic diagram of an HTTP request using the HTTP / 2 protocol.

[0034] In order to fully collect the original feature data of the target Web server response header, request payloads are sent to ports 80 and 443 of the target Web server respectively, and coroutines are used to accelerate concurrent request payloads. In abnormal response processing, since different types of Web servers have different response states when processing different types of request payloads, in order to standardize the original response header key sequence, the <EMPTY> tag is used as the identifier to fill in the characteristics of response failure. In sample feature processing, considering that the number of response header keys containing normal service identifiers is in the range of [5,10], when the target response header is obtained, the response header key sequence value of the current request payload feature is constructed according to the header field key sequence with spaces as intervals. In sample label processing, the content corresponding to the Server field in the Web server response design is the type identifier of the response Web server, such as "Server:Nginx" or "Server:Apache / 2.4.43". In the response header that does not contain the "Server" identifier and cannot extract the sample label, the <LABEL> tag is used as the identifier to fill in.

[0035] After the above payload request and original response feature and label processing, the detection target set is defined as ST, and the response feature set corresponding to the original request payload type is: S PT ={BRP j , LBVP k , APMP m , SMHP n , H2P r , j∈[1, 4], k∈[1, 2], m∈[1, 3],n∈[1, 5], r∈[1, 2]}; The original response header field key sequence set is S R , the label set is L RWeb asset server original response header field key sequence dataset DS RespHeaders Defined as: .

[0036] In extracting the dataset DS RespHeaders Before embedding the semantic sequence of the sample into features, DS RespHeaders The response data is preprocessed and standardized as the input of the subsequent RoBERTa model. Therefore, in the data preprocessing stage, in the sample feature value cleaning, due to the business differences of different types of Web servers, the original response header field key sequence set S R The abnormal response header field key sequence values ​​are mixed with HTML tags, Javascript code segments, etc. In the data cleaning stage, regular matching combined with <HTML> mark filling is used to process the abnormal header field key sequence. In the effective response sample filtering, if the detection target set S R If the invalid labels filling <EMPTY> and <LABEL> in the target response feature header field key sequence value in the target response feature header field key sequence value exceeds 60%, it is considered as an invalid sample. RespHeaders The response label set L for each sample in Sample The frequent items are used as the true labels of the original samples to obtain the sample label set .

[0037] In the original response dataset standardization stage, we first select common Web server types based on statistical data to construct the dataset DS. RespHeaders The coarse-grained label set of: L major ={Nginx,Apache,MicrosoftIIS,LiteSpeed,Openresty},L major ∈ .

[0038] Secondly, in L major On the basis of Construct dataset DS with labels of fine-class version RespHeaders The fine-grained label set L minor If the tag value of the detection target lacks the detailed version information, it is replaced by the <Major> identifier. major With L minor Filter DS RespHeaders Invalid response samples in the data set DS are obtained after the response data is cleaned. RH_Standardized Finally, from the dataset DS RH_Standardized Extract valid response header field key sequence sample set As the training input for the subsequent tokenizer and RoBERTa feature embedding model. Standardized dataset DS RH_Standardized The structure is defined as: .

[0039] S200: Design a deep semantic feature extraction model based on RoBERTa model fine-tuning. Through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding dataset.

[0040] In this embodiment, in order to use the RoBERTa model to standardize the dataset DS RH_Standardized The response header domain key sequence samples in the data embedding generate a high-dimensional feature dataset. It is necessary to embed the response header domain key sequence feature sets of different request types. As the training corpus for Tokenizer segmentation, a text segmenter is constructed for the Web server asset fingerprint recognition scenario. Compared with the BERT model that uses subword tags for segmentation, the RoBERTa model uses byte-level byte pair encoding segmentation to decompose the HTTP response header field key sequence text into the smallest byte-level tags, allowing the model to learn more subtle HTTP response text features. The segmentation calculation process is: ; Among them, the valid response header field key sequence sample set As the input corpus, k is the number of times the corpus characters are merged, and Freq is the merge frequency of the calculated character pair (a, b). is the character pair merged for the i-th time.

[0041] Character pair frequencies are calculated as: ; in, It is a corpus The i-th sample in (·) is the indicator function. When the character pair (a, b) appears in The value is 1 when it is in the range, otherwise it is 0.

[0042] In the character pair merging calculation, the character pairs (ai, bi) with the highest frequency in each iteration are merged, and the vocabulary is gradually selected to form the required high-frequency character pairs: ; The word list required for the final model is constructed through the above character comparison rate and merging iterative calculation. The word list iterative update process is: ; Among them, Vk is the vocabulary after the kth iteration, V k+1 It is the new vocabulary after merging character pairs.

[0043] In order to reduce the training cost of the tokenizer and subsequent embedding model, the RoBERTa model is used as a pre-trained model and migrated to the Web server asset fingerprint recognition task. From the perspective of task scale, in order to ensure that the vocabulary can cover most of the high-frequency words and avoid the presence of uncollected words, the relationship formula between the vocabulary size and the corpus coverage is introduced to make the corpus coverage reach the size of the training corpus. Furthermore, in order to adapt to the original data distribution of the response header field key sequence of different HTTP requests, the vocabulary parameter size is set to 8192 in the construction of the Tokenizer tokenizer: ; Among them, C(V) identifies the vocabulary V for the corpus The coverage rate of The ratio of the frequency of words successfully recognized by the vocabulary to the total frequency of words. Freq(w) is the frequency of word w in the corpus, reflecting the frequency of word w in the entire corpus. The coverage is calculated to evaluate the effectiveness of the vocabulary and ensure that it can fully capture the important information in the corpus.

[0044] The standardized response dataset DS obtained in the data preprocessing section RH_Standardized After that, on the one hand, it is necessary to embed the original response data samples containing the semantic features of the response key sequence into high-dimensional features to extract the text semantics of the key sequence in the response header. On the other hand, the text of the original dataset needs to be converted into data input that can be trained by the subsequent classification model. In the response data feature embedding generation stage of the proposed method, the RoBERTa model is fine-tuned to learn the dataset DS. RH_Standardized The semantic features of the response header key sequence samples are used to construct a high-dimensional embedding generation dataset DS that carries the semantic features of the original response header domain key sequence. Embedding . The RoBERTa-based Web asset response data embedding uses the RoBERTa model for migration and fine-tuning. In the model learning phase of the semantic features of the response header domain key sequence, on the one hand, in the model structure design, the embedding dimension of the RoBERTa model is first set to 256 dimensions to represent the vector length of each token in the embedding space. Specifically, the embedding layer maps the input sequence x to a high-dimensional vector representation: ; Where E is the embedding matrix, d is the embedding dimension, is the length of the input sequence.

[0045] In order to reduce the complexity of the embedding generation model and speed up training and inference, the number of Transformer encoder layers in the model is set to 4. Furthermore, in order for the model to capture more fine-grained feature relationships of the response head domain key sequence, a multi-head attention mechanism is introduced. The multi-head attention calculation process is: ; Where Q, K, V are query, key, and value matrices respectively, h is the number of attention heads, and W O is the weight output matrix, the number of attention heads in the model is set to 4, and the size of the middle layer of the model is set to 1024 to increase the nonlinear transformation capability of the middle layer.

[0046] On the other hand, in terms of training parameter selection, firstly, in order to fully learn the feature representation of the key sequence in the response header field, the number of model training iterations is set to 30 rounds. Considering the computing resources and model convergence, the training batch size is selected to be 1024. Secondly, in order to stabilize the model convergence and avoid gradient oscillation, the learning rate is set to 0.0001 during the model training process, and the Adam optimizer is used to update the parameters. The parameter update process is as follows: ; in, is the learning rate, and are the estimates of the first and second order moments, respectively.

[0047] Furthermore, in order to minimize the loss function of the masked language model and enhance the generalization and robustness of the model, the proportion of random masked tokens is set to 15% in the masked language model training. The loss function of the masked language model is: ; Finally, a weight decay coefficient of 0.01 is selected to prevent overfitting during the model training process.

[0048] To obtain a high-dimensional embedding dataset DS containing the semantic features of the original response header field key sequence Embedding In the embedding generation stage, a fine-tuned RoBERTa model is used to standardize the dataset DS RH_Standardized The semantic features of the response header domain key sequence are extracted. In the word segmentation part, the dataset DS RH_Standardized The input sequence is padded and truncated, and then the embedding vector of each token is converted into the semantic representation of the entire sequence through mean pooling to capture the global features of the sequence. In order to reduce the embedding dimension and extract key semantic features, the embedding dataset DS Embedding The principal component analysis technology is introduced in the generation process to reduce the dimension of the original embedding vector to 64 dimensions. The dimension reduction calculation process is: ; Among them, X is the original embedding matrix, U is the orthogonal matrix containing the first k principal components, and k is the target dimension. Finally, the embedding dataset DS containing the semantic features of the original response header domain key sequence is generated Embedding .

[0049] S300: constructing a coarse-grained Web server fingerprint identification classification model based on random forest, performing coarse-grained fingerprint classification identification of Web asset servers, and outputting coarse-grained fingerprint classification identification data.

[0050] In this embodiment, the original response header domain key sequence dataset DS is fine-tuned by RoBERTa model. RH_Standardized Extract semantic features and embed them to generate high-dimensional feature dataset DS Embedding , which is used to classify and identify the coarse-grained Web asset server fingerprint samples in the future. The random forest algorithm is used for ensemble learning strategy, which builds multiple decision trees and integrates them into a comprehensive classifier to improve the generalization ability of the model. The decision tree is a building block of the random forest model, and its construction process is as follows: ; Where Tb(x) is the prediction result of the bth decision tree, y i is the sample label, c is the category, (·) is the indicator function. The final prediction result of the random forest is to integrate the prediction results of all decision trees through the majority voting method as the prediction result output. The integrated prediction process is: ; Where B is the number of decision trees, is the final prediction result of random forest.

[0051] Therefore, in the downstream task stage, a coarse-grained server fingerprint recognition model based on random forest is constructed to realize the coarse-grained fingerprint classification and recognition of Web asset servers. The model first uses StandardScaler to classify the dataset DS Embedding The high-dimensional semantic vectors in are normalized to eliminate the scale differences between different features. In addition, in the training process of the proposed model, GridSearchCV is used to search for the best hyperparameters. The calculation process is: ; Among them, θ is a hyperparameter combination, θ * is the best hyperparameter combination.

[0052] At the same time, the 5-fold cross-validation method is used to verify the stability of the model during the training process, where the loss function of the cross-validation is: ; Where T k is the k-th fold training set, V k is the k-th fold validation set, is the loss value of the model on the validation set.

[0053] Furthermore, the early stopping strategy is combined to reduce the time required for model training. By combining the above training strategies and methods, the coarse-grained server fingerprint recognition model based on random forest can more effectively cope with high-dimensional vector data sets containing high-dimensional fingerprint semantic features and improve classification and recognition accuracy.

[0054] S400: constructing a fine-grained Web server fingerprint identification classification model based on a feedforward neural network, performing fine-grained fingerprint classification identification of Web asset servers, and outputting fine-grained fingerprint classification identification data.

[0055] In this embodiment, for the fine-grained fingerprint recognition task of Web asset servers containing detailed version information, considering that the number of fine-grained fingerprint categories contained in different coarse-grained Web servers far exceeds the number of coarse-grained categories, in order to capture the subtle differences of fine-grained fingerprint features in high-dimensional space and improve the classification accuracy of the multi-category Web asset server fingerprint recognition classification model, the fine-grained server fingerprint recognition model based on the feedforward neural network is used to realize the fingerprint recognition of fine-grained Web asset servers containing detailed versions.

[0056] In the classification model, the input data is a high-dimensional embedding vector of fine-grained fingerprints with a dimension size of 64. In order to make full use of the feature information of the embedded vector, first, the overall structure of the model consists of 1 input layer, 2 hidden layers and 1 output layer. The first hidden layer contains 1024 neurons and the second hidden layer contains 512 neurons. ReLU (activation function is used to introduce nonlinear characteristics to enhance the expressiveness of the model. Secondly, in order to reduce model overfitting and improve the generalization ability of the model, the model sets a Dropout rate of 0.2 after each hidden layer. By randomly discarding a certain proportion of neurons during the training process, the model is forced to learn more robust feature representations and improve performance on unseen data. Finally, the number of neurons in the output layer is set to the total number of categories, and the Softmax activation function is used to output the probability distribution of each fine-grained category fingerprint.

[0057] In the model training process, in order to accelerate model convergence and improve training efficiency, Adam is used as the model optimizer, and the early stopping strategy is used to prevent model overfitting and reduce model training time. Through optimizer selection, loss function configuration and training strategy, the training efficiency and generalization ability of the model can be effectively improved, and accurate classification of fine-grained fingerprint data can be achieved.

[0058] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0059] It should be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A Web asset mapping method based on response structure recognition, characterized in that: The steps include: Construct five categories and sixteen types of HTTP request payloads, send requests to the target domain name, obtain response header data, and standardize it to build the original response header domain key sequence data set; Design a deep semantic feature extraction model based on RoBERTa model fine-tuning. Through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding dataset. Based on random forest, a coarse-grained Web server fingerprint recognition classification model is constructed to perform coarse-grained fingerprint classification recognition of Web asset servers and output coarse-grained fingerprint classification recognition data; A fine-grained Web server fingerprint recognition and classification model is constructed based on a feedforward neural network to perform fine-grained fingerprint classification and recognition of Web asset servers and output fine-grained fingerprint classification and recognition data.

2. The Web asset mapping method based on response structure recognition according to claim 1, characterized in that: In the steps of constructing five categories and sixteen types of HTTP request payloads, sending requests to the target domain name, obtaining response header data, and standardizing it, and constructing the original response header field key sequence data set: The categories of request payload include basic request payload, newline variant payload, abnormal protocol and method payload, special method and header payload, and HTTP / 2 request payload.

3. The Web asset mapping method based on response structure recognition according to claim 1, characterized in that: In the steps of designing a deep semantic feature extraction model based on RoBERTa model fine-tuning, fine-tuning the RoBERTa model through transfer learning and fine-tuning strategies, extracting deep semantic features of the response header domain key sequence, and generating a high-dimensional semantic embedding dataset: Use a byte-level byte pair encoding tokenizer to break down the HTTP response header field key sequence text into the smallest byte-level tokens; Calculate the frequency of character pairs, merge the character pairs with the highest frequency in each iteration, and gradually select the vocabulary to form the required high-frequency character pairs; Obtain a response dataset, learn the semantic features of the response header key sequence samples of the response dataset by fine-tuning the RoBERTa model, and construct a high-dimensional embedding that carries the semantic features of the original response header domain key sequence to generate a high-dimensional feature dataset.

4. The Web asset mapping method based on response structure recognition as claimed in claim 3, characterized in that: In the steps of building a coarse-grained Web server fingerprint identification and classification model based on random forest, performing coarse-grained fingerprint classification and identification of Web asset servers, and outputting coarse-grained fingerprint classification and identification data: Build multiple decision trees and integrate them into an integrated classifier; The prediction results of all decision trees are integrated through majority voting method as the prediction result output; Normalize the high-dimensional semantic vectors in the high-dimensional feature dataset, search for the best hyperparameters, and output them.

5. The Web asset mapping method based on response structure recognition as claimed in claim 4, characterized in that: In the steps of building a fine-grained Web server fingerprint identification and classification model based on a feedforward neural network, performing fine-grained fingerprint classification and identification of Web asset servers, and outputting fine-grained fingerprint classification and identification data: The Softmax activation function is used to output the probability distribution of each fine-grained category fingerprint.

Citation Information

Patent Citations

  • Rapid and accurate flow detection method based on network security probe

    CN113904795A

  • Network asset fingerprint feature identification method and apparatus, and electronic device

    CN115830649A

  • Column semantic recognition method and system based on context awareness of GCN and RoBERTa

    CN117312989A

  • Network asset fingerprint identification method and device based on deep learning

    CN117851547A

  • Network asset fingerprint identification method and device based on Bayesian algorithm

    CN119760527A

Cited By

  • Network fingerprint identification method and system based on feature hierarchical filtering and deep learning classification

    CN121012761A