Web Asset Mapping Method Based on Response Structure Recognition
By constructing multiple HTTP request payloads and RoBERTa models fine-tuning, combined with random forests and feedforward neural networks, the problems of high false positive rate and maintenance threshold of traditional Web asset recognition methods are solved, and efficient, accurate identification and surveying of Web assets are achieved.
Patent Information
- Application Number
- CN202510502401.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The traditional web asset fingerprint recognition method based on rule matching has problems such as high false positive rate and high maintenance threshold, which is difficult to adapt to the rapidly iterative web application architecture, and it is difficult to extract effective semantic features.
Five categories of sixteen HTTP request payloads are constructed, and the deep semantic features of the response head domain key sequence are extracted through the RoBERTa model, and the coarse and fine-grained fingerprint classification and recognition are combined with random forests and feedforward neural networks to output accurate Web server asset fingerprints.
It realizes the accuracy of efficient identification and mapping of Web assets, and improves the accuracy and stability of Web assets identification.
Smart Images

Figure CN120017428B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a Web asset mapping method based on response structure recognition. Background Art
[0002] With the wide application of Web application services, Web asset security faces severe challenges. These assets not only carry a large amount of personal privacy and industrial production data, but also are exposed to the Internet environment and become targets of malicious attackers. To address this threat, enterprise security protection personnel need to comprehensively master the quantity and composition of Web assets in the organization's Internet exposure surface, and specifically reinforce the assets with security risks.
[0003] Traditional Web asset fingerprint recognition methods based on rule matching have problems such as high false alarm rates and high maintenance thresholds, and it is difficult to adapt to the rapidly iterating Web application architecture. In recent years, the academic and industrial communities have begun to pay attention to the identification features in the Web server response headers, but these features have high dimensions and different text contents, and it is difficult to extract effective semantic features.
[0004] Therefore, it is very necessary to propose a Web asset mapping method that can effectively identify Web server asset fingerprints and improve the accuracy of identifying and mapping Web assets. Summary of the Invention
[0005] The purpose of the present invention is to provide a Web asset mapping method based on response structure recognition, aiming to effectively identify Web server asset fingerprints and improve the accuracy of identifying and mapping Web assets.
[0006] To achieve the above object, a Web asset mapping method based on response structure recognition adopted by the present invention includes the following steps:
[0007] Construct five categories and sixteen kinds of HTTP request payloads, send requests to the target domain name, obtain response header data, and perform standardization processing on it to construct an original response header domain key sequence data set;
[0008] Design a deep semantic feature extraction model based on fine-tuning of the RoBERTa model, and through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding data set;
[0009] Construct a coarse-grained Web server fingerprint recognition classification model based on random forest, perform coarse-grained fingerprint classification recognition of Web asset servers, and output coarse-grained fingerprint classification recognition data;
[0010] Construct a fine-grained Web server fingerprint recognition classification model based on a feedforward neural network to perform fine-grained fingerprint classification recognition of Web asset servers and output fine-grained fingerprint classification recognition data.
[0011] Among them, in the step of constructing five categories and sixteen types of HTTP request payloads, sending requests to the target domain name, obtaining response header data, and performing standardization processing on it to construct the original response header domain key sequence data set:
[0012] The categories of request payloads include basic request payloads, newline character variant payloads, abnormal protocol and method payloads, special method and header payloads, and HTTP / 2 request payloads.
[0013] Among them, in the step of designing a deep semantic feature extraction model based on fine-tuning of the RoBERTa model, fine-tuning the RoBERTa model through transfer learning and fine-tuning strategies, extracting deep semantic features of the response header domain key sequence, and generating a high-dimensional semantic embedding data set:
[0014] Use a byte-level byte pair encoding tokenizer to decompose the HTTP response header domain key sequence text into the smallest byte-level tokens;
[0015] Calculate the character pair frequencies, merge the character pairs with the highest frequencies selected in each iteration, and gradually select the vocabulary to form the required high-frequency character pairs;
[0016] Obtain the response data set, adopt the method of fine-tuning based on the RoBERTa model, learn the semantic features of the response header key sequence samples in the response data set, and construct a high-dimensional embedding carrying the semantic features of the original response header domain key sequence to generate a high-dimensional feature data set.
[0017] Among them, in the step of constructing a coarse-grained Web server fingerprint recognition classification model based on a random forest to perform coarse-grained fingerprint classification recognition of Web asset servers and output coarse-grained fingerprint classification recognition data:
[0018] Construct multiple decision trees and integrate them into a comprehensive classifier;
[0019] Integrate the prediction results of all decision trees through the majority voting method as the prediction result output;
[0020] Perform feature standardization on the high-dimensional semantic vectors in the high-dimensional feature data set, search for the best hyperparameters, and output.
[0021] Among them, in the step of constructing a fine-grained Web server fingerprint recognition classification model based on a feedforward neural network to perform fine-grained fingerprint classification recognition of Web asset servers and output fine-grained fingerprint classification recognition data:
[0022] The probability distribution of each fine-grained category fingerprint is output using the Softmax activation function.
[0023] A Web asset mapping method based on response structure recognition in the present invention constructs five categories and sixteen types of HTTP request payloads, sends requests to the target domain name, obtains response header data, and performs normalization processing on it to construct an original response header domain key sequence data set; designs a deep semantic feature extraction model based on fine-tuning of the RoBERTa model, fine-tunes the RoBERTa model through transfer learning and fine-tuning strategies, extracts deep semantic features of the response header domain key sequence, and generates a high-dimensional semantic embedding data set; constructs a coarse-grained Web server fingerprint recognition classification model based on random forest to perform coarse-grained fingerprint classification recognition of Web asset servers and output coarse-grained fingerprint classification recognition data; constructs a fine-grained Web server fingerprint recognition classification model based on a feed-forward neural network to perform fine-grained fingerprint classification recognition of Web asset servers and output fine-grained fingerprint classification recognition data; through the above methods, effectively identify Web server asset fingerprints and improve the accuracy of identifying and mapping Web assets. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a flowchart of the steps of the Web asset mapping method based on response structure recognition of the present invention.
[0026] Figure 2 It is a schematic diagram of the request payload carrying the OPTIONS method of the present invention.
[0027] Figure 3 It is a schematic diagram of the standard HEAD request and the simplified GET request using the HTTP / 1.1 protocol of the present invention.
[0028] Figure 4 It is a schematic diagram of the simple GET request and the GET request using the HTTP / 1.0 protocol of the present invention.
[0029] Figure 5 It is a schematic diagram of the GET request using the HTTP / 1.0 protocol with the Host header of the present invention.
[0030] Figure 6 It is a schematic diagram of the HEAD request using \n as the line break character of the present invention.
[0031] Figure 7 It is a schematic diagram of the HEAD request using \r as a line break in the present invention.
[0032] Figure 8 It is a schematic diagram of the HEAD request using the non-standard HTTP / 4.0 protocol in the present invention.
[0033] Figure 9 It is a schematic diagram of the HTTP request using no SPOCK method in the present invention.
[0034] Figure 10 It is a schematic diagram of the HTTP request using a mixed case of method and protocol in the present invention.
[0035] Figure 11 It is a schematic diagram of the HTTP request using the OPTIONS method in the present invention.
[0036] Figure 12 It is a schematic diagram of the HEAD request carrying the Range header in the present invention.
[0037] Figure 13 It is a schematic diagram of the HEAD request with the future-time If-Modified-Since header in the present invention.
[0038] Figure 14 It is a schematic diagram of the HEAD request with the past-time If-Unmodified-Since header in the present invention.
[0039] Figure 15 It is a schematic diagram of the HEAD request with a complex Accept header in the present invention.
[0040] Figure 16 It is a schematic diagram of the HTTP request using the HTTP / 2 protocol in the present invention. Detailed implementation manners
[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0042] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms of "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0043] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0044] Please refer to Figures 1 to 16 , the present invention provides a Web asset mapping method based on response structure recognition, including the following steps:
[0045] S100: Construct five categories and sixteen kinds of HTTP request payloads, send requests to the target domain name, obtain response header data, and perform standardization processing on it to construct an original response header domain key sequence data set.
[0046] In this embodiment, in the Web asset server fingerprint recognition work, it is necessary to design fuzz test request payloads for the target Web server to collect Web server response packet information. On the one hand, considering the normal operation of the target asset, the request payload packets are all requests such as GET and HEAD for the purpose of obtaining response data, and will not affect the target server. On the other hand, according to the different types of request payloads, a total of five categories and sixteen kinds of request payload test cases are adopted. The categories of request payloads include basic request payloads, newline variant payloads, abnormal protocol and method payloads, special method and header payloads, and HTTP / 2 request payloads; among them, there are 4 basic request payloads, 2 newline variant payloads, 3 abnormal protocol and method payloads, 5 special method and header payloads, and 2 HTTP / 2 request payloads. The original header domain characteristics of the target Web asset server response packets are collected through the above types of payloads, such as Figures 2 to 16 shown Figure 2 is a schematic diagram of a request payload carrying the OPTIONS method; Figure 3 is a schematic diagram of a standard HEAD request and a simplified version of a GET request using the HTTP / 1.1 protocol; Figure 4 is a schematic diagram of a simple GET request and a GET request using the HTTP / 1.0 protocol; Figure 5 is a schematic diagram of a GET request using the HTTP / 1.0 protocol with a Host header; Figure 6 is a schematic diagram of a HEAD request using \n as the newline character; Figure 7 is a schematic diagram of a HEAD request using \r as the newline character; Figure 8 is a schematic diagram of a HEAD request using the non-standard HTTP / 4.0 protocol;Figure 9 Schematic diagram of an HTTP request using a non-existent SPOCK method; Figure 10 Schematic diagram of an HTTP request using a method and protocol with mixed case; Figure 11 Schematic diagram of an HTTP request using the OPTIONS method; Figure 12 Schematic diagram of a HEAD request carrying a Range header; Figure 13 Schematic diagram of a HEAD request with a future-time If-Modified-Since header; Figure 14 Schematic diagram of a HEAD request with a past-time If-Unmodified-Since header; Figure 15 Schematic diagram of a HEAD request with a complex Accept header; Figure 16 Schematic diagram of an HTTP request using the HTTP / 2 protocol.
[0047] To fully collect the original feature data of the response header of the target web server, request payloads are sent to ports 80 and 443 of the target web server respectively, and coroutines are used to accelerate the concurrent request payloads. In the handling of abnormal responses, since the response statuses of different types of web servers for different types of request payloads are different, to standardize the processing of the original response header key sequence, the <EMPTY> tag is used as an identifier to fill in the features of failed responses. In the processing of sample features, considering that the number of response header keys containing normal service identifiers is in the range of [5, 10], when the target response header is obtained, the response header key sequence value of the current request payload feature is constructed in the order of the header field keys with spaces as separators. In the processing of sample tags, according to the content corresponding to the Server field in the web server response design, it is the type identifier of the response web server, such as "Server:Nginx" or "Server:Apache / 2.4.43", etc. In responses where sample tags cannot be extracted from the response header without the "Server" identifier, the <LABEL> tag is used as an identifier to fill in.
[0048] Through the above load request, original response feature, and label processing, the set of detection targets is defined as ST, and the response feature sets corresponding to the original request payload types are:
[0049] S PT ={BRP j , LBVP k , APMP m , SMHP n , H2P r , j ∈ [1, 4], k ∈ [1, 2], m ∈ [1, 3], n ∈ [1, 5], r ∈ [1, 2]};
[0050] The set of keys in the original response header fields is S R , and the set of tags is L R . The dataset DS of the keys in the original response header fields of the Web asset server RespHeaders is defined as:
[0051] .
[0052] Before extracting the semantic sequence embedding features of the samples in the dataset DS RespHeaders , it is necessary to perform response data preprocessing and standardization on DS RespHeaders as the input for the subsequent RoBERTa model. Therefore, in the data preprocessing stage, in terms of cleaning the sample feature values, due to the business differences of different types of Web servers, in the set of keys in the original response header fields S R , there are abnormal response header field key sequence values such as HTML tags and JavaScript code segments. In the data cleaning stage, abnormal header field key sequences are processed by regular matching combined with the <HTML> identifier filling. In terms of filtering valid response samples, if the invalid identifier fillings <EMPTY> and <LABEL> in the target response feature header field key sequence values in the detection target set S R exceed 60%, the sample is regarded as invalid. In terms of sample label selection, the frequent items in the response label set L RespHeaders of each sample in the dataset DS Sample are used as the true labels of the original samples, and the sample label set is obtained.
[0053] In the stage of standardizing the original response dataset, first, according to the statistical data, a general Web server type is selected to construct the coarse-grained label set of the dataset DS RespHeaders :
[0054] L major = {Nginx, Apache, Microsoft IIS, LiteSpeed, Openresty}, L major ∈ .
[0055] Secondly, on the basis of L major , the tags carrying the fine-grained version in are used to construct the fine-grained label set L RespHeaders of the dataset DS minor . If the tag value of the detection target lacks the fine-grained version information, it is replaced by the <Major> identifier. Further, after determining the label category range, the invalid response samples in DS major are filtered by the label sets L minor and L RespHeaders , and the standardized dataset DS after cleaning the response data is obtainedRH_Standardized Finally, extract the valid response header field key sequence sample set RH_Standardized from the data set DS as the training input for the subsequent Tokenizer and RoBERTa feature embedding model. Normalize the data set DS RH_Standardized The structure is defined as:
[0056] .
[0057] S200: Design a deep semantic feature extraction model based on fine-tuning of the RoBERTa model. Through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract the deep semantic features of the response header field key sequence and generate a high-dimensional semantic embedding data set.
[0058] In this embodiment, in order to be able to use the RoBERTa model to perform data embedding on the response header field key sequence samples in the normalized data set DS RH_Standardized and generate a high-dimensional feature data set, it is necessary to use the response header field key sequence feature sets of different request types as the training corpus for Tokenizer tokenization and construct a text tokenizer in the Web server asset fingerprint recognition scenario. Compared with the BERT model that uses sub-word tokens for tokenization, the RoBERTa model uses a byte-level byte pair encoding tokenizer that can decompose the HTTP response header field key sequence text into the smallest byte-level tokens, enabling the model to learn more subtle HTTP response text features. The tokenization calculation process is as follows:
[0059] ;
[0060] Among them, the valid response header field key sequence sample set is used as the input corpus, k is the number of times of corpus character merging, Freq is the merged frequency of calculating the character pair (a, b), is the character pair merged in the i-th merge.
[0061] The character pair frequency calculation is:
[0062] ;
[0063] Among them, is the i-th sample in the corpus , (·) is the indicator function, which takes the value of 1 when the character pair (a, b) appears in and 0 otherwise.
[0064] In the character pair merging calculation, the character pair (ai, bi) with the highest frequency is selected for merging in each iteration, and the vocabulary is gradually selected to form the required high-frequency character pairs:
[0065] ;
[0066] Construct the vocabulary required for the final model through the above character pair frequency evaluation and combined iterative calculation. The process of iterative update of the vocabulary is as follows:
[0067] ;
[0068] Among them, V k is the vocabulary after the k-th iteration, and V k+1 is the new vocabulary after merging character pairs.
[0069] To reduce the training cost of the tokenizer and the subsequent embedding model, the RoBERTa model is used as the pre-trained model and migrated to the Web server asset fingerprint recognition task. From the perspective of the task scale, to ensure that the vocabulary can cover most high-frequency words and avoid the occurrence of uncollected words, a relationship formula between the vocabulary size and the corpus coverage rate is introduced to make the corpus coverage rate reach the size scale that meets the training corpus. Further, to adapt to the original data distribution of the response header field key sequences in different HTTP requests, the size of the vocabulary parameter in the construction of the Tokenizer is set to 8192:
[0070] ;
[0071] Among them, C(V) identifies the coverage rate of the vocabulary V for the corpus , which represents the proportion of the vocabulary frequency successfully recognized by the vocabulary in the corpus to the total vocabulary frequency. Freq(w) is the frequency of the word w in the corpus, reflecting the number of times the word w appears in the entire corpus . The effectiveness of the vocabulary is evaluated by calculating the coverage rate to ensure that it can fully capture the important information in the corpus.
[0072] After obtaining the standardized response data set DS RH_Standardized in the data preprocessing part, on the one hand, it is necessary to perform high-dimensional feature embedding on the original response data samples containing the semantic features of the response key sequence to extract the text semantics of the response header field key sequence. On the other hand, the text of the original data set needs to be converted into the data input that the subsequent classification model can train. In the response data feature embedding generation stage of the proposed method, by adopting the method of fine-tuning based on the RoBERTa model, learn the semantic features of the response header key sequence samples in the data set DS RH_Standardized to construct a high-dimensional embedding generation data set DS Embedding. The RoBERTa-based Web asset response data embedding uses the RoBERTa model for transfer fine-tuning. In the semantic feature learning stage of the model for the response header field key sequence, on the one hand, in terms of model structure design, first, the embedding dimension of the RoBERTa model is set to 256 dimensions, representing the vector length of each token in the embedding space. Specifically, the embedding layer maps the input sequence x to a high-dimensional vector representation:
[0073] ;
[0074] where E is the embedding matrix and d is the embedding dimension, is the length of the input sequence.
[0075] To reduce the complexity of the embedding generation model and accelerate the training and inference speed, the number of Transformer encoder layers in the model is set to 4 layers. Further, to enable the model to capture the finer-grained feature relationships of the response header field key sequence, the multi-head attention mechanism is introduced. The multi-head attention calculation process is:
[0076] ;
[0077] where Q, K, and V are the query, key, and value matrices respectively, h is the number of attention heads, and W O is the weight output matrix. The number of attention heads in the model is set to 4. To increase the non-linear transformation ability of the intermediate layer, the size of the intermediate layer of the model is set to 1024.
[0078] On the other hand, in terms of training parameter selection, first, to fully learn the feature representation of the response header field key sequence, the number of model training iterations is set to 30 rounds. Considering the computing resources and model convergence, the training batch size is selected to be 1024. Second, to ensure the stable convergence of the model and avoid gradient oscillation, the learning rate is set to 0.0001 during the model training process, and the Adam optimizer is used to update the parameters. The parameter update process is as follows:
[0079] ;
[0080] where, is the learning rate, and are the estimated values of the first moment and the second moment respectively.
[0081] Furthermore, to minimize the loss function of the masked language model and enhance the generalization ability and robustness of the model, the random masking token ratio in the masked language model training is set to 15%. Among them, the loss function of the masked language model is:
[0082] ;
[0083] Finally, a weight decay coefficient of 0.01 is selected to prevent overfitting during the model training process.
[0084] To obtain a high-dimensional embedding dataset DS containing the semantic features of the original response header field key sequence Embedding , in the embedding generation stage, a fine-tuned RoBERTa model is used to extract semantic features from the standardized dataset DS RH_Standardized of the response header field key sequence. In the tokenization part, the input sequence of the dataset DS RH_Standardized is first padded and truncated, and then through mean pooling, the embedding vectors of each token are converted into the semantic representation of the overall sequence to capture the global features of the sequence. To reduce the embedding dimension and extract key semantic features, in the process of generating the embedding dataset DS Embedding , the principal component analysis technique is introduced to reduce the dimension of the original embedding vector to 64 dimensions. The dimension reduction calculation process is as follows:
[0085] ;
[0086] where X is the original embedding matrix, U is the orthogonal matrix containing the first k principal components, and k is the target dimension. Finally, an embedding dataset DS containing the semantic features of the original response header field key sequence is generated Embedding .
[0087] S300: Construct a coarse-grained Web server fingerprint recognition classification model based on random forest to perform coarse-grained fingerprint classification recognition of Web asset servers and output coarse-grained fingerprint classification recognition data.
[0088] In this embodiment, semantic features are extracted from the original response header field key sequence dataset DS RH_Standardized through RoBERTa model fine-tuning, and a high-dimensional feature dataset DS Embedding is embedded and generated for subsequent classification and recognition of coarse-grained Web asset server fingerprint samples. The random forest algorithm is used for the ensemble learning strategy. By constructing multiple decision trees and integrating them into a comprehensive classifier, the generalization ability of the model is improved. The decision tree, as a building block of the random forest model, its construction process is as follows:
[0089] ;
[0090] where Tb(x) is the prediction result of the bth decision tree, y i is the sample label, c is the category, (·) is the indicator function. And the final prediction result of the random forest is to integrate the prediction results of all decision trees by the majority voting method as the prediction result output. The ensemble prediction process is as follows:
[0091] ;
[0092] where B is the number of decision trees, which is the final prediction result of the random forest.
[0093] Therefore, in the downstream task stage, a coarse-grained server fingerprint recognition model based on random forest is constructed to achieve the coarse-grained fingerprint classification and recognition of Web asset servers. The model first uses StandardScaler to standardize the features of the high-dimensional semantic vectors in the dataset DS Embedding to eliminate the scale differences between different features. In addition, during the training process of the proposed model, GridSearchCV is used to search for the best hyperparameters, and the calculation process is as follows:
[0094] ;
[0095] where θ is the hyperparameter combination, and θ * is the best hyperparameter combination.
[0096] At the same time, the 5-fold cross-validation method is used to verify the stability of the model during the training process. The loss function of the cross-validation is:
[0097] ;
[0098] where T k is the training set of the k-th fold, V k is the validation set of the k-th fold, and is the loss value of the model on the validation set.
[0099] Furthermore, the early stopping strategy is combined to reduce the time required for model training. By combining the above training strategies and methods, the coarse-grained server fingerprint recognition model based on random forest can more effectively handle the high-dimensional vector dataset containing high-dimensional fingerprint semantic features and improve the classification and recognition accuracy.
[0100] S400: Construct a fine-grained Web server fingerprint recognition classification model based on a feedforward neural network to perform fine-grained fingerprint classification and recognition of Web asset servers, and output fine-grained fingerprint classification and recognition data.
[0101] In this embodiment, for the fine-grained fingerprint recognition task of Web asset servers with detailed version information, considering that the number of fine-grained fingerprint categories contained in different coarse-grained Web servers far exceeds the number of coarse-grained categories, in order to capture the subtle differences of fine-grained fingerprint features in the high-dimensional space and improve the classification accuracy of the multi-category Web asset server fingerprint recognition classification model, a fine-grained server fingerprint recognition model based on a feedforward neural network is used to achieve the fine-grained fingerprint recognition of Web asset servers with detailed versions.
[0102] In the classification model, the input data is a high-dimensional embedding vector of fine-grained fingerprints, with a dimension size of 64 dimensions. To make full use of the feature information of the embedding vector, first, the overall structure of the model consists of 1 input layer, 2 hidden layers, and 1 output layer. Among them, the first hidden layer contains 1024 neurons, and the second hidden layer contains 512 neurons. The ReLU activation function is used to introduce non-linearity to enhance the expressive power of the model. Secondly, to reduce model overfitting and improve the generalization ability of the model, a Dropout rate of 0.2 is set after each hidden layer. By randomly discarding a certain proportion of neurons during the training process, the model is forced to learn more robust feature representations and improve its performance on unseen data. Finally, the number of neurons in the output layer is set to the total number of categories, and the Softmax activation function is used to output the probability distribution of each fine-grained category fingerprint.
[0103] In the part of the model training process, on the one hand, to accelerate model convergence and improve training efficiency, Adam is used as the model optimizer, and at the same time, an early stopping strategy is adopted to prevent model overfitting and reduce the training time of the model. Through optimizer selection, loss function configuration, and training strategies, the training efficiency and generalization ability of the model can be effectively improved, and accurate classification of fine-grained fingerprint data can be achieved.
[0104] After considering the specification and the content disclosed herein, those skilled in the art will readily think of other embodiments of the present application. The present application aims to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0105] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A Web asset mapping method based on response structure recognition, characterized in that It includes the following steps: Construct sixteen HTTP request payloads of five categories, send requests to the target domain name, obtain response header data, and perform standardization processing on it to construct an original response header domain key sequence dataset; Design a deep semantic feature extraction model based on fine-tuning of the RoBERTa model. Through transfer learning and fine-tuning strategies, fine-tune the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding dataset; Construct a coarse-grained Web server fingerprint recognition classification model based on random forest, perform coarse-grained fingerprint classification recognition of Web asset servers, and output coarse-grained fingerprint classification recognition data; Construct a fine-grained Web server fingerprint recognition classification model based on a feed-forward neural network, perform fine-grained fingerprint classification recognition of Web asset servers, and output fine-grained fingerprint classification recognition data; In the step of designing a deep semantic feature extraction model based on fine-tuning of the RoBERTa model, through transfer learning and fine-tuning strategies, fine-tuning the RoBERTa model to extract deep semantic features of the response header domain key sequence and generate a high-dimensional semantic embedding dataset: Use a byte-level byte pair encoding tokenizer to decompose the HTTP response header domain key sequence text into the smallest byte-level tokens; Calculate the character pair frequencies, merge the character pairs with the highest frequencies selected in each iteration, and gradually select the vocabulary to form the required high-frequency character pairs; Obtain the response dataset, adopt the method of fine-tuning based on the RoBERTa model, learn the semantic features of the response header key sequence samples in the response dataset, and construct a high-dimensional embedding carrying the semantic features of the original response header domain key sequence to generate a high-dimensional feature dataset.
2. The Web asset mapping method based on response structure recognition according to claim 1, wherein In the step of constructing sixteen HTTP request payloads of five categories, sending requests to the target domain name, obtaining response header data, and performing standardization processing on it to construct an original response header domain key sequence dataset: The categories of request payloads include basic request payloads, newline character variant payloads, abnormal protocol and method payloads, special method and header payloads, and HTTP / 2 request payloads.
3. The Web asset mapping method based on response structure recognition according to claim 1, wherein In the step of constructing a coarse-grained Web server fingerprint recognition classification model based on random forest, performing coarse-grained fingerprint classification recognition of Web asset servers, and outputting coarse-grained fingerprint classification recognition data: Construct multiple decision trees and integrate them into a comprehensive classifier; Integrate the prediction results of all decision trees as the prediction result output through the majority voting method; Perform feature standardization on the high-dimensional semantic vectors in the high-dimensional feature dataset, search for the best hyperparameters, and output.
4. The Web asset mapping method based on response structure recognition according to claim 3, wherein In the step of constructing a fine-grained Web server fingerprint recognition classification model based on a feed-forward neural network, performing fine-grained fingerprint classification recognition of Web asset servers, and outputting fine-grained fingerprint classification recognition data: Use the Softmax activation function to output the probability distribution of each fine-grained category fingerprint.
Citation Information
Patent Citations
Rapid and accurate flow detection method based on network security probe
CN113904795A
Network asset fingerprint feature identification method and apparatus, and electronic device
CN115830649A
Column semantic recognition method and system based on context awareness of GCN and RoBERTa
CN117312989A
Network asset fingerprint identification method and device based on Bayesian algorithm
CN119760527A