XSS attack detection method based on attention mechanism
By using the ResidualBiLSTM-Attention model in XSS attack detection, using bidirectional LSTM, multi-head self-attention mechanism and residual connection, the problems of high false alarm rate and insufficient context-related capture capabilities in the existing technology are solved, and high-precision and low false alarm rate XSS attack detection is achieved.
Patent Information
- Application Number
- CN202510271320.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-08
- Publication Date
- 2025-06-06
AI Technical Summary
The existing XSS attack detection technology has high false positive rate, making it difficult to deal with new attack variants, and deep learning models lack the context correlation ability of long-sequence attack text.
The ResidualBiLSTM-Attention model based on attention mechanism is adopted to capture the front and back dependencies of traffic characteristics through bidirectional long and short-term memory networks, and dynamically allocate weights with the multi-head self-attention mechanism, strengthen the semantic representation of key fragments of XSS attacks, and solve the problem of deep network gradient disappearance through residual connections.
It realizes high-precision detection of XSS attacks, low false alarm rate and strong anti-obfuscation ability, and improves the accuracy and stability of detection.
Smart Images

Figure CN120110758A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and in particular relates to an XSS attack detection method based on an attention mechanism. Background Art
[0002] With the widespread popularity of Web applications, cross-site scripting (XSS) attacks have become a core threat in the field of network security due to their concealment, diversity of variants and harmfulness.
[0003] Traditional XSS detection technology mainly relies on static analysis (such as grammatical rule matching) and dynamic analysis (such as simulated attack injection), but its false positive rate is high and it is difficult to deal with new attack variants such as obfuscated coding. Although the detection method based on machine learning can automatically learn attack patterns, it relies on manual feature engineering and has insufficient ability to capture contextual associations of long-sequence attack texts (such as nested malicious tags).
[0004] In recent years, although deep learning models (such as LSTM and CNN) have performed well in feature extraction, unidirectional LSTM cannot model bidirectional dependencies, the standard attention mechanism has limited ability to focus on key segments of long sequences, and deep networks are prone to model degradation due to gradient disappearance, which restricts detection accuracy and generalization ability. Summary of the invention
[0005] Based on this, it is necessary to provide an XSS attack detection method based on the attention mechanism to address the above technical problems.
[0006] In a first aspect, the present application provides an XSS attack detection method based on an attention mechanism, comprising:
[0007] S1: Clean the original network traffic data to obtain redundant traffic data;
[0008] S2: Perform structured word segmentation processing on the de-redundant traffic data to obtain a word segmentation sequence;
[0009] S3: Use the pre-trained Word2Vec model to vectorize the word segmentation sequence and generate a structured feature vector;
[0010] S4: Input the structured feature vector into the ResidualBiLSTM-Attention model for detection to obtain the classification results of the original network traffic data; the ResidualBiLSTM-Attention model consists of a bidirectional long short-term memory network, a multi-head self-attention mechanism, and a residual structure; the classification results are normal traffic, XSS attack traffic, or non-XSS attack traffic;
[0011] S5: Generate an alarm instruction or an interception instruction according to the classification result.
[0012] In the second aspect, the present application also provides an XSS attack detection system based on an attention mechanism, including:
[0013] The flow data cleaning module is used to clean the original network flow data to obtain redundant flow data;
[0014] The structured word segmentation module is used to receive the de-redundant traffic data, perform structured word segmentation processing on the de-redundant traffic data, and obtain a word segmentation sequence;
[0015] The feature vectorization module is used to receive the word segmentation sequence, perform vectorization processing on the word segmentation sequence through the pre-trained Word2Vec model, and generate a structured feature vector;
[0016] The traffic detection module is used to receive the structured feature vector, input the structured feature vector into the ResidualBiLSTM-Attention model for detection, and obtain the classification result of the original network traffic data; wherein the ResidualBiLSTM-Attention model is composed of a bidirectional long short-term memory network, a multi-head self-attention mechanism, and a residual structure; the classification result is normal traffic, XSS attack traffic, or non-XSS attack traffic;
[0017] The instruction generation module is used to receive the classification results and generate an alarm instruction or an interception instruction according to the classification results.
[0018] In a third aspect, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, an XSS attack detection method based on an attention mechanism as in the first aspect is implemented.
[0019] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the XSS attack detection method based on an attention mechanism as in the first aspect is implemented.
[0020] The above-mentioned XSS attack detection method based on attention mechanism constructs RBA model (ResidualBiLSTM-Attention model) by integrating bidirectional long short-term memory network (BiLSTM), multi-head self-attention mechanism and residual connection, uses BiLSTM to capture the forward and backward bidirectional dependency of traffic features, dynamically allocates weights through multi-head self-attention to strengthen the semantic representation of key fragments of XSS attack, and combines residual connection to solve the gradient disappearance problem of deep network, finally achieving high-precision detection of XSS attacks, low false alarm rate and strong anti-confusion ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 A flowchart of an XSS attack detection method based on an attention mechanism provided by the present invention;
[0023] Figure 2 A structural schematic diagram of an XSS attack detection system based on an attention mechanism provided by the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0025] refer to Figure 1 , which shows a flow chart of an XSS attack detection method based on an attention mechanism provided by the present application, the method comprising the following steps:
[0026] S1: Clean the original network traffic data to obtain redundant traffic data.
[0027] Specifically, the cleaning process may include: (1) Removing irrelevant fields: Based on the characteristics of the dataset of the original network traffic data, only the fields related to XSS attack detection, such as HTTP method, path, and parameter parts, are retained. For example, in the HTTPDATASET CSIC 2010 dataset, other HTTP header information is removed and only the URL and parameter parts are retained. (2) Formatting: Format the retained data to ensure the consistency and standardization of the data. For example, special characters in the URL are uniformly processed to ensure the accuracy of subsequent word segmentation and feature extraction. (3) Removing duplicate data: Check whether there are duplicate request records in the dataset and remove duplicate items to reduce redundancy.
[0028] The purpose of this step is to remove irrelevant information and redundant data and improve data quality and processing efficiency.
[0029] S2: Perform structured word segmentation processing on the de-redundant traffic data to obtain a word segmentation sequence.
[0030] Specifically, structured word segmentation processing may include: (1) Special character segmentation: Use special characters (such as ":", "?", "&", "=", "<", ">", " / ", ".", ";", "@", "~", "*", "%", "-", "+", etc.) to segment the URL. During the word segmentation process, the above special characters are replaced with spaces on the left and right to ensure the accuracy of the word segmentation results. (2) Number conversion: Convert the pure numeric part in the URL into a unified "Numbers" identifier to reduce the complexity of the data and retain the digital features. (3) Word conversion: Perform a simple word conversion on the results after word segmentation, such as converting "http", "https", etc. to lowercase to ensure the consistency of the features.
[0031] The purpose of this step is to separate the keywords and symbols in the URL to better extract features and perform subsequent processing.
[0032] S3: Use the pre-trained Word2Vec model to vectorize the word segmentation sequence and generate a structured feature vector.
[0033] Specifically, a pre-trained Word2Vec model (such as CBOW or Skip-Gram model) is used to vectorize the word sequence to generate a vector representation of each word. The word vector dimension of the model can be adjusted according to the scale and feature complexity of the dataset, for example, set to 100 or 256.
[0034] Vectorization is to convert each word in the word segmentation sequence into a corresponding vector through the Word2Vec model to generate a structured feature vector. This vector can capture the semantic relationship and contextual information between words, providing a basis for subsequent model detection.
[0035] The purpose of this step is to convert text data into numerical vectors so that the model can learn and process them.
[0036] S4: Input the structured feature vector into the ResidualBiLSTM-Attention model for detection to obtain the classification results of the original network traffic data; the ResidualBiLSTM-Attention model consists of a bidirectional long short-term memory network, a multi-head self-attention mechanism, and a residual structure; the classification results are normal traffic, XSS attack traffic, or non-XSS attack traffic.
[0037] Specifically, the structured feature vector is input into the ResidualBiLSTM-Attention model. The input layer of the model receives the feature vector and passes it to the subsequent BiLSTM layer.
[0038] The BiLSTM layer captures the contextual information in the sequence through a bidirectional learning mechanism and generates a hidden state, which contains the dependencies between the previous and next sequences.
[0039] The attention mechanism calculates the importance weight of each position based on the hidden state to highlight key information. The multi-head self-attention mechanism can extract information from different feature subspaces and enhance the expressiveness of the model.
[0040] Through residual connections, the model can better learn deep features and solve the gradient vanishing problem. Residual connections pass the input directly to subsequent layers and superimpose it with the learned features, enhancing the training stability of the model.
[0041] The feature vector processed by the attention mechanism and residual connection is input to the output layer for classification prediction. The output layer uses the softmax function to calculate the probability of each category and finally obtains the classification results, including normal traffic, XSS attack traffic or non-XSS attack traffic.
[0042] S5: Generate an alarm instruction or an interception instruction according to the classification result.
[0043] Specifically, if the detection result is XSS attack traffic, an alarm instruction is generated to notify the security management personnel for further analysis and processing. The alarm instruction may include detailed information of the attack, such as the attack source, attack type, and attack time.
[0044] According to the preset security policy, interception instructions can be automatically generated to prevent attack traffic from entering the network system. The interception instructions can be sent to the firewall or other security devices to block the attack in real time.
[0045] The detection results and related instructions can be recorded in log files for subsequent audit and analysis. Log files can be used to track attack behaviors, evaluate system performance, and optimize security policies.
[0046] The above-mentioned XSS attack detection method based on attention mechanism constructs RBA model (ResidualBiLSTM-Attention model) by integrating bidirectional long short-term memory network (BiLSTM), multi-head self-attention mechanism and residual connection, uses BiLSTM to capture the forward and backward bidirectional dependency of traffic features, dynamically allocates weights through multi-head self-attention to strengthen the semantic representation of key fragments of XSS attack, and combines residual connection to solve the gradient disappearance problem of deep network, finally achieving high-precision detection of XSS attacks, low false alarm rate and strong anti-confusion ability.
[0047] In an optional embodiment, the S3 includes the following steps:
[0048] S31: Input the structured feature vector into a bidirectional long short-term memory network for bidirectional context feature extraction to obtain a bidirectional context feature.
[0049] Specifically, the traditional unidirectional LSTM can only process sequence data from a single direction (such as from front to back), while the bidirectional LSTM can capture richer contextual information in the sequence by processing sequence data from two directions at the same time (from front to back and from back to front). In XSS attack detection, the above-mentioned bidirectional processing method can more accurately capture the possible front-to-back correlation features in the network traffic sequence, such as the key character combination at different positions of the attack payload.
[0050] Each data packet or feature element in network traffic can be regarded as an element in a sequence. Bidirectional LSTM can extract contextual features around each element by bidirectionally processing the sequence. For example, when processing an HTTP request containing a potential XSS attack payload, the bidirectional LSTM can analyze the relationship between elements such as parameters and paths in the request and the previous and next elements to capture possible attack patterns. The above contextual features are crucial for identifying key features of XSS attacks (such as specific HTML tags, JavaScript code snippets, etc.).
[0051] S32: Input the bidirectional context features into the multi-head self-attention mechanism for attention weighted processing to obtain a weighted feature vector.
[0052] Specifically, this mechanism allows the model to focus on different features at different positions when processing the input sequence. Compared with the single-head attention mechanism, the multi-head attention mechanism can capture information from multiple feature subspaces, thereby providing a richer feature representation. In XSS attack detection, the multi-head self-attention mechanism can simultaneously focus on multiple key features in network traffic, such as specific character combinations, abnormal parameter values, etc., thereby improving the ability to identify XSS attack patterns.
[0053] Through the attention mechanism, the model can assign a weight to each element in the input sequence, indicating the importance of the element in the current task. In XSS attack detection, this step helps the model automatically identify features related to XSS attacks in network traffic and give them higher attention. For example, if a URL parameter contains a large number of angle brackets < and >, these characters may be part of the XSS attack payload, and the attention mechanism will assign them a higher weight, thereby highlighting the key features in the subsequent feature fusion and classification process.
[0054] S33: The weighted feature vector is input into the residual structure for residual optimization processing, and the shallow and deep features are fused by combining the superposition operation to obtain the classification result.
[0055] Specifically, the residual network introduces a shortcut connection to pass the input directly to the next layer, so that the network can learn the residual of the input. This step helps solve the gradient vanishing problem in deep neural networks and promotes the effective transmission of information. In XSS attack detection, the residual structure enables the model to better retain shallow features in the deep network while learning new deep features. In this way, the model can not only use shallow features (such as simple character matching) to detect XSS attacks, but also combine deep features (such as complex semantic patterns) to improve the accuracy of detection.
[0056] Through the superposition operation in the residual structure, shallow features and deep features can be fused. This method helps the model to comprehensively analyze the input data at different levels. For example, in XSS attack detection, shallow features may include specific characters or phrases in the URL, while deep features may involve complex association patterns between parameters. Through feature fusion, the model can understand the network traffic data more comprehensively, thereby more accurately determining whether there is an XSS attack.
[0057] In an optional embodiment, S31 includes the following steps:
[0058] S311: Perform forward LSTM processing on the structured feature vector to generate a forward hidden state sequence.
[0059] Specifically, the forward LSTM starts from the start position of the sequence and processes the structured feature vector sequentially along the time step. At each time step, the LSTM unit receives the current feature vector and the hidden state of the previous time step, and updates the hidden state through the control of the forget gate, input gate, and output gate. The forget gate determines whether to retain or discard the hidden state information of the previous time step, the input gate controls the degree of update of the hidden state by the current feature vector, and the output gate determines how the hidden state of the current time step is used as the output. After the above series of operations, the forward LSTM generates a forward hidden state sequence, which contains the context information from the start position of the sequence to the current position.
[0060] Each element in the forward hidden state sequence represents the contextual features from the start of the sequence to the current time step. For example, when processing a sequence of URL parameters, the forward LSTM can capture the semantic connections and pattern changes of the parameters from left to right. The above information is very useful for identifying common leading features in XSS attacks (such as <script>标签)具有重要意义,因为该标签通常出现在序列的前部。
[0061] S312:对结构化特征向量进行后向LSTM处理,生成后向隐藏状态序列。
[0062] 具体的,后向LSTM与前向LSTM的方向相反,从序列的结束位置开始,反向处理结构化特征向量。同样地,在每个时间步,LSTM单元接收当前特征向量和前一时间步(在反向处理中的后一时间步)的隐藏状态,通过遗忘门、输入门和输出门的控制,更新隐藏状态。经过处理,后向LSTM生成一个后向隐藏状态序列,该序列包含从序列结束位置到当前位置的上下文信息。
[0063] 后向隐藏状态序列中的每个元素代表从序列结束位置到当前时间步的上下文特征。在XSS攻击检测中,该步骤有助于捕捉到攻击载荷可能存在的尾部特征,如< / script> Tags or some specific end symbols. By processing the sequence in reverse, the model can better understand the complete structure and semantic pattern of the attack payload.
[0064] S313: Concatenate the forward hidden state sequence and the backward hidden state sequence to generate a bidirectional context feature.
[0065] Specifically, the concatenation operation merges the hidden states of the forward hidden state sequence and the backward hidden state sequence at each time step to generate a higher-dimensional bidirectional context feature. This concatenation operation enables the model to utilize both forward and backward information to better understand the context of each position in the sequence. For example, for a specific input feature vector, its corresponding bidirectional context feature contains the forward information from the start position of the sequence to that position, and the backward information from the end position of the sequence to that position. This step provides the model with a more comprehensive semantic representation.
[0066] Bidirectional context features can capture the dependencies between the previous and next characters in a sequence, which is crucial for XSS attack detection. Many XSS attack payloads do not exist in isolation, and they may depend on a specific context. By extracting bidirectional context features, the model can more accurately identify the above dependencies and use them as an important basis for judging whether an attack has occurred. For example, when detecting XSS attacks, bidirectional context features can help the model identify whether a key attack character is in a specific context, such as whether a < symbol is adjacent to a / symbol, which may be part of a tag.
[0067] In an optional embodiment, S32 includes the following steps:
[0068] S321: Split the bidirectional context features into multiple groups of subspaces, and generate corresponding query matrices, key matrices, and value matrices based on each group of subspaces.
[0069] Specifically, the multi-head self-attention mechanism can capture the dependencies in the input sequence from multiple angles by mapping the input features into multiple different subspaces. In XSS attack detection, multi-angle observation helps the model capture different features in the attack payload, such as specific character combinations, abnormal parameter values, etc.
[0070] The bidirectional context features are split into multiple groups of subspaces, each of which corresponds to an attention head. Each attention head generates a query matrix (Query), a key matrix (Key), and a value matrix (Value) through linear transformation. The generation process of the above matrices can be regarded as a projection of the input features, so that each attention head can focus on different aspects of the input sequence. For example, one attention head may focus on a specific character combination in the URL, while another attention head may focus on the relationship between parameters.
[0071] The query matrix (Query) represents the focus of the current attention head and is used to match the key matrix to determine which parts of the input sequence are relevant. The key matrix (Key) represents the features in the input sequence and is used to match the query matrix to determine which parts of the input sequence are relevant. The value matrix (Value) represents the actual content in the input sequence and is used to perform weighted summation in the attention calculation results to generate the final weighted feature vector.
[0072] S322: Perform a scaled dot product attention calculation on each group of subspaces according to the query matrix, the key matrix, and the value matrix to obtain an attention calculation result; the formula for the scaled dot product attention calculation is:
[0073]
[0074] Among them, d k is the dimension of the subspace, Q i is the query matrix of the i-th group of subspace, K i is the bond matrix of the ith group of subspaces, V i is the value matrix of the i-th group of subspace.
[0075] Specifically, the attention calculation result represents the importance weight of each position in the input sequence, which reflects the importance of the position in the current attention head. In this way, the model can automatically focus on the key parts of the input sequence, thereby better capturing the characteristics of XSS attacks. For example, when detecting XSS attacks, the model may give a higher attention weight to URL parameters containing malicious code, thereby more accurately identifying attack behaviors.
[0076] S323: Concatenate and linearly transform the attention calculation results of multiple groups of subspaces to generate a weighted feature vector.
[0077] Specifically, the attention calculation results of multiple groups of subspaces are concatenated to form a higher-dimensional feature vector. The concatenation operation enables the model to integrate information from different subspaces to obtain a more comprehensive feature representation. For example, if each attention head focuses on different aspects of the input sequence, the concatenated feature vector will contain information from different aspects, thereby more comprehensively describing the characteristics of the input sequence.
[0078] Perform a linear transformation on the concatenated feature vectors to generate the final weighted feature vector. The linear transformation is implemented through a fully connected layer, and the weight matrix of this layer can be learned through training. The role of the linear transformation is to map the concatenated feature vectors to a more suitable feature space, so as to better adapt to subsequent classification tasks. For example, the linear transformation can compress the redundant information in the concatenated feature vectors and extract more useful features, thereby improving the classification performance of the model.
[0079] The weighted feature vector contains the weighted information of each position in the input sequence, and the weight reflects the importance of the position in different attention heads. In this way, the model can more comprehensively understand the characteristics of the input sequence, so as to detect XSS attacks more accurately. For example, when detecting XSS attacks, the weighted feature vector can highlight the URL parameters containing malicious code, thus helping the model to more accurately identify the attack behavior.
[0080] In an optional embodiment, S33 includes the following steps:
[0081] S331: The weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model is input into the combined operation layer of the bidirectional long short-term memory network and the multi-head self-attention mechanism to generate an intermediate feature vector.
[0082] Specifically, the combined operation layer combines the bidirectional long short-term memory network (BiLSTM) and the multi-head self-attention mechanism to fully utilize the advantages of both. BiLSTM can capture long-term dependencies in the sequence, while the multi-head self-attention mechanism can extract rich feature information from multiple subspaces. By inputting the weighted feature vector into the combined operation layer, the model can further extract and fuse features to generate richer intermediate feature vectors.
[0083] In the combination operation layer, the weighted feature vector first passes through the BiLSTM layer to generate a hidden state sequence containing long-term dependency information. Then, the hidden state sequence is processed by the multi-head self-attention mechanism to generate a weighted feature vector. Finally, the weighted feature vector is integrated into an intermediate feature vector that contains more comprehensive context information and key features.
[0084] S332: perform element-by-element addition processing on the intermediate feature vector and the weighted feature vector outputted from the lth layer to generate an optimized feature vector of the l+1th layer; the formula for element-by-element addition processing is:
[0085] H (l+1) =H (l) +F(H (l) );
[0086] Among them, H (l+1) is the optimized feature vector output by the l+1th layer of the ResidualBiLSTM-Attention model, H (l) is the weighted feature vector output by the lth layer of the Residua lBiLSTM-Attention model, and F(·) represents the combined operation of the bidirectional long short-term memory network and the multi-head self-attention mechanism.
[0087] Specifically, element-by-element addition is one of the core operations of the residual structure. By adding the intermediate feature vector to the weighted feature vector output by the lth layer element by element, the model can fuse the shallow features with the deep features. This fusion method not only retains the important information in the shallow features, but also utilizes the complex patterns in the deep features, thereby improving the expressiveness and generalization capabilities of the model. For example, in XSS attack detection, shallow features may include simple character matching information, while deep features may include complex semantic patterns. By adding element by element, the model can better utilize the above information and improve the accuracy of detection.
[0088] S333: Perform full connection layer processing on the optimized feature vector to generate classification probability distribution.
[0089] Specifically, the fully connected layer maps the optimized feature vector to a higher dimensional space to generate a classification probability distribution. This process is implemented through one or more fully connected layers, each of which contains multiple neurons, and the neurons are connected to the neurons in the previous layer through a weight matrix. The output of the fully connected layer is a vector representing the score of each category.
[0090] The output vector of the fully connected layer represents the score of each category, which is converted into a probability distribution through the Softmax function. The Softmax function normalizes the score of each category to a probability value, and the sum of the probability values is 1. The classification probability distribution represents the probability that the input data belongs to each category, providing a basis for the final classification result.
[0091] S334: Perform Softmax normalization processing on the classification probability distribution to obtain a classification result.
[0092] Specifically, the Softmax function normalizes the output vector of the fully connected layer into a probability distribution. Specifically, the Softmax function converts the score of each category into a probability value, and the sum of the probability values is 1. The formula of the Softmax function is:
[0093]
[0094] Among them, z i is the score of the ith category, and n is the total number of categories.
[0095] Through Softmax normalization, the model obtains the probability value of each category. The final classification result is the category with the highest probability value. For example, in XSS attack detection, if the probability value of normal traffic is the highest, the classification result is normal traffic; if the probability value of XSS attack traffic is the highest, the classification result is XSS attack traffic.
[0096] The above-mentioned XSS attack detection method based on attention mechanism constructs a ResidualBiLSTM-Attention model that integrates multi-head self-attention mechanism, residual connection and bidirectional long short-term memory network, extracts and classifies features of network traffic data that has been cleaned, segmented and vectorized, extracts context features through bidirectional LSTM, performs feature weighting through multi-head self-attention mechanism, and fuses shallow and deep features through residual structure, and finally generates classification results through fully connected layer and Softmax function, thereby realizing efficient detection of XSS attack traffic, improving detection accuracy and stability, effectively reducing false alarm rate and missed alarm rate, and providing an effective technical means for network security protection.
[0097] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0098] Based on the same inventive concept, the embodiment of the present application also provides a system for implementing the XSS attack detection method based on the attention mechanism involved above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more XSS attack detection system embodiments based on the attention mechanism provided below can be found in the above limitations on an XSS attack detection method based on the attention mechanism, and will not be repeated here.
[0099] In an exemplary embodiment, Figure 2 As shown, an XSS attack detection system 20 based on an attention mechanism is provided, comprising:
[0100] The flow data cleaning module 21 is used to clean the original network flow data to obtain redundancy-free flow data.
[0101] The structured word segmentation module 22 is used to receive the de-redundant traffic data, perform structured word segmentation processing on the de-redundant traffic data, and obtain a word segmentation sequence.
[0102] The feature vectorization module 23 is used to receive the word segmentation sequence, perform vectorization processing on the word segmentation sequence through the pre-trained Word2Vec model, and generate a structured feature vector.
[0103] The traffic detection module 24 is used to receive the structured feature vector, input the structured feature vector into the ResidualBiLSTM-Attention model for detection, and obtain the classification result of the original network traffic data; wherein the ResidualBiLSTM-Attention model is composed of a bidirectional long short-term memory network, a multi-head self-attention mechanism and a residual structure; the classification result is normal traffic, XSS attack traffic or non-XSS attack traffic.
[0104] The instruction generation module 25 is used to receive the classification result and generate an alarm instruction or an interception instruction according to the classification result.
[0105] Optionally, the flow detection module 24 includes:
[0106] The bidirectional long short-term memory network submodule 241 is used to receive the structured feature vector, input the structured feature vector into the bidirectional long short-term memory network to perform bidirectional context feature extraction processing, and obtain the bidirectional context feature.
[0107] The multi-head self-attention mechanism submodule 242 is used to receive bidirectional context features, input the bidirectional context features into the multi-head self-attention mechanism for attention weighting processing, and obtain a weighted feature vector.
[0108] The residual structure submodule 243 is used to receive the weighted feature vector, input the weighted feature vector into the residual structure for residual optimization processing, and combine the shallow and deep features with the superposition operation to obtain the classification result.
[0109] Optionally, the bidirectional long short-term memory network submodule 241 includes:
[0110] The forward LSTM processing unit 2411 is used to receive the structured feature vector, perform forward LSTM processing on the structured feature vector, and generate a forward hidden state sequence.
[0111] The backward LSTM processing unit 2412 is used to receive the structured feature vector, perform backward LSTM processing on the structured feature vector, and generate a backward hidden state sequence.
[0112] The concatenation processing unit 2413 is used to receive the forward hidden state sequence and the backward hidden state sequence, concatenate the forward hidden state sequence and the backward hidden state sequence, and generate a bidirectional context feature.
[0113] Optionally, the multi-head self-attention mechanism submodule 242 includes:
[0114] The subspace splitting and matrix generating unit 2421 is used to receive the bidirectional context features, split the bidirectional context features into multiple groups of subspaces, and generate corresponding query matrices, key matrices and value matrices based on each group of subspaces.
[0115] The scaled dot product attention calculation unit 2422 is used to receive the query matrix, the key matrix and the value matrix, and perform a scaled dot product attention calculation on each group of subspaces according to the query matrix, the key matrix and the value matrix to obtain an attention calculation result; the formula for the scaled dot product attention calculation is:
[0116]
[0117] Among them, d k is the dimension of the subspace, Q i is the query matrix of the i-th group of subspace, K i is the bond matrix of the ith group of subspaces, V i is the value matrix of the i-th group of subspace.
[0118] The concatenation and linear transformation unit 2423 is used to receive the attention calculation results, concatenate and linearly transform the attention calculation results of multiple groups of subspaces, and generate a weighted feature vector.
[0119] Optionally, the residual structure submodule 243 includes:
[0120] The feature vector input and combination operation unit 2431 is used to receive the weighted feature vector, input the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model into the combination operation layer of the bidirectional long short-term memory network and the multi-head self-attention mechanism, and generate an intermediate feature vector.
[0121] The residual connection and feature optimization unit 2432 is used to receive the intermediate feature vector and the weighted feature vector, perform element-by-element addition processing on the intermediate feature vector and the weighted feature vector outputted by the lth layer, and generate the optimized feature vector of the l+1th layer; the formula for element-by-element addition processing is:
[0122] H (l+1) =H (l) +F(H (l) );
[0123] Among them, H (l+1) is the optimized feature vector output by the l+1th layer of the ResidualBiLSTM-Attention model, H (l) is the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model, and F(·) represents the combined operation of the bidirectional long short-term memory network and the multi-head self-attention mechanism.
[0124] The fully connected layer processing unit 2433 is used to receive the optimized feature vector, perform fully connected layer processing on the optimized feature vector, and generate a classification probability distribution.
[0125] The Softmax normalization processing unit 2434 is used to receive the classification probability distribution, perform Softmax normalization processing on the classification probability distribution, and obtain a classification result.
[0126] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0127] The embodiments of the present application further provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0128] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial description of the method embodiments. The device embodiments described above are only schematic, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0129] The above-mentioned embodiments only express several implementation methods of the embodiments of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent of the embodiments of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the embodiments of the present application, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. A XSS attack detection method based on attention mechanism, characterized in that: The method comprises: S1: Clean the original network traffic data to obtain redundant traffic data; S2: performing structured word segmentation processing on the de-redundant traffic data to obtain a word segmentation sequence; S3: vectorizing the word segmentation sequence through a pre-trained Word2Vec model to generate a structured feature vector; S4: Input the structured feature vector into the ResidualBiLSTM-Attention model for detection to obtain the classification result of the original network traffic data; wherein the ResidualBiLSTM-Attention model is composed of a bidirectional long short-term memory network, a multi-head self-attention mechanism and a residual structure; the classification result is normal traffic, XSS attack traffic or non-XSS attack traffic; S5: Generate an alarm instruction or an interception instruction according to the classification result.
2. The method according to claim 1, characterized in that: The S3 includes: S31: Inputting the structured feature vector into the bidirectional long short-term memory network to perform bidirectional context feature extraction processing to obtain bidirectional context features; S32: Inputting the bidirectional context feature into the multi-head self-attention mechanism for attention weighting processing to obtain a weighted feature vector; S33: Input the weighted feature vector into the residual structure for residual optimization processing, combine the superposition operation to fuse the shallow and deep features, and obtain the classification result.
3. The method according to claim 2, characterized in that The S31 includes: S311: Perform forward LSTM processing on the structured feature vector to generate a forward hidden state sequence; S312: Performing backward LSTM processing on the structured feature vector to generate a backward hidden state sequence; S313: Concatenate the forward hidden state sequence and the backward hidden state sequence to generate the bidirectional context feature.
4. The method according to claim 2, characterized in that: The S32 includes: S321: Split the bidirectional context features into multiple groups of subspaces, and generate corresponding query matrices, key matrices, and value matrices based on each group of subspaces; S322: Performing a scaled dot product attention calculation on each group of the subspaces according to the query matrix, the key matrix and the value matrix to obtain an attention calculation result; the formula for the scaled dot product attention calculation is: Among them, d k is the dimension of the subspace, Q i is the query matrix of the i-th group of subspace, K i is the bond matrix of the ith group of subspaces, V i is the value matrix of the i-th group of subspace; S323: Concatenate and linearly transform the attention calculation results of multiple groups of the subspaces to generate the weighted feature vector.
5. The method according to any one of claims 2 to 4, characterized in that: The S33 includes: S331: Input the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model into the combined operation layer of the bidirectional long short-term memory network and the multi-head self-attention mechanism to generate an intermediate feature vector; S332: performing element-by-element addition processing on the intermediate feature vector and the weighted feature vector outputted from the lth layer to generate an optimized feature vector of the l+1th layer; the formula for the element-by-element addition processing is: H (l+1) =H (l) +F(H (l) ); Among them, H (l+1) is the optimized feature vector output by the l+1th layer of the ResidualBiLSTM-Attention model, H (l) is the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model, and F(·) represents the combined operation of the bidirectional long short-term memory network and the multi-head self-attention mechanism; S333: Performing full connection layer processing on the optimized feature vector to generate classification probability distribution; S334: Perform Softmax normalization processing on the classification probability distribution to obtain the classification result.
6. An XSS attack detection system based on attention mechanism, characterized in that: The system comprises: The flow data cleaning module is used to clean the original network flow data to obtain redundant flow data; A structured word segmentation module, used for receiving the de-redundant traffic data, performing structured word segmentation processing on the de-redundant traffic data, and obtaining a word segmentation sequence; A feature vectorization module is used to receive the word segmentation sequence, perform vectorization processing on the word segmentation sequence through a pre-trained Word2Vec model, and generate a structured feature vector; A traffic detection module, used for receiving the structured feature vector, inputting the structured feature vector into the ResidualBiLSTM-Attention model for detection, and obtaining the classification result of the original network traffic data; wherein the ResidualBiLSTM-Attention model is composed of a bidirectional long short-term memory network, a multi-head self-attention mechanism and a residual structure; and the classification result is normal traffic, XSS attack traffic or non-XSS attack traffic; The instruction generation module is used to receive the classification result and generate an alarm instruction or an interception instruction according to the classification result.
7. The system according to claim 6, characterized in that The flow detection module comprises: A bidirectional long short-term memory network submodule is used to receive the structured feature vector, input the structured feature vector into the bidirectional long short-term memory network to perform bidirectional context feature extraction processing, and obtain a bidirectional context feature; A multi-head self-attention mechanism submodule, used for receiving the bidirectional context feature, inputting the bidirectional context feature into the multi-head self-attention mechanism for attention weighting processing, and obtaining a weighted feature vector; The residual structure submodule is used to receive the weighted feature vector, input the weighted feature vector into the residual structure for residual optimization processing, and fuse the shallow and deep features in combination with the superposition operation to obtain the classification result.
8. The system according to claim 7, characterized in that The bidirectional long short-term memory network submodule includes: A forward LSTM processing unit, used for receiving the structured feature vector, performing forward LSTM processing on the structured feature vector, and generating a forward hidden state sequence; A backward LSTM processing unit, used for receiving the structured feature vector, performing backward LSTM processing on the structured feature vector, and generating a backward hidden state sequence; The concatenation processing unit is used to receive the forward hidden state sequence and the backward hidden state sequence, and concatenate the forward hidden state sequence and the backward hidden state sequence to generate the bidirectional context feature.
9. The system according to claim 7, characterized in that The multi-head self-attention mechanism submodule includes: A subspace splitting and matrix generating unit, configured to receive the bidirectional context features, split the bidirectional context features into a plurality of groups of subspaces, and generate a corresponding query matrix, a key matrix and a value matrix based on each group of the subspaces; A scaled dot product attention calculation unit is used to receive the query matrix, the key matrix and the value matrix, and perform a scaled dot product attention calculation on each group of the subspaces according to the query matrix, the key matrix and the value matrix to obtain an attention calculation result; the formula for the scaled dot product attention calculation is: Among them, d k is the dimension of the subspace, Q i is the query matrix of the i-th group of subspace, K i is the bond matrix of the ith group of subspaces, V i is the value matrix of the i-th group of subspace; The concatenation and linear transformation unit is used to receive the attention calculation result, concatenate and linearly transform the attention calculation results of multiple groups of the subspaces, and generate the weighted feature vector.
10. The system according to claim 7, characterized in that The residual structure submodule includes: A feature vector input and combination operation unit, used to receive the weighted feature vector, input the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model into the combination operation layer of the bidirectional long short-term memory network and the multi-head self-attention mechanism, and generate an intermediate feature vector; The residual connection and feature optimization unit is used to receive the intermediate feature vector and the weighted feature vector, perform element-by-element addition processing on the intermediate feature vector and the weighted feature vector outputted by the lth layer, and generate an optimized feature vector of the l+1th layer; the formula for the element-by-element addition processing is: H (l+1) =H (l) +F(H (l) ); Among them, H (l+1) is the optimized feature vector output by the l+1th layer of the ResidualBiLSTM-Attention model, H (l) is the weighted feature vector output by the lth layer of the ResidualBiLSTM-Attention model, and F(@) represents the combined operation of the bidirectional long short-term memory network and the multi-head self-attention mechanism; A fully connected layer processing unit, used to receive the optimized feature vector, perform fully connected layer processing on the optimized feature vector, and generate a classification probability distribution; The Softmax normalization processing unit is used to receive the classification probability distribution, perform Softmax normalization processing on the classification probability distribution, and obtain the classification result.