Multi-feature extraction method based on TBA-GRU dual-channel model

By adopting the TBA-GRU dual-channel model in the XSS detection model, combining TextCNN and BiGRU-multi-header attention channel, the shortcomings of the existing models in processing local features and global context information are solved, and higher detection accuracy and robustness are achieved.

CN120197169APending Publication Date: 2025-06-24SHANXI JINXINAN TECH CO LTD

Patent Information

Application Number
CN202510391717.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing XSS detection model has insufficient in processing local features and global context information, resulting in low detection accuracy and robustness.

Method used

Using a multi-feature extraction method based on the TBA-GRU dual-channel model, local features and BiGRU-multi-header attention channel are extracted through TextCNN channels to capture timing features, and multi-level deep feature expression is achieved through splicing and fusion output.

Benefits of technology

It significantly improves the accuracy and robustness of XSS attack detection, and is better than multiple evaluation indicators of traditional single-structure deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197169A_ABST
    Figure CN120197169A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-feature extraction method based on a TBA-GRU dual-channel model, and belongs to the technical field, and the method specifically comprises two parallel channels, namely a TextCNN channel and a BiGRU channel, which are used for extracting local features and time sequence features respectively, and further enhancing the attention to important features in combination with a multi-head attention mechanism. According to the XSS attack detection method, the output features of the TextCNN and BiGRU channels are spliced and fused through the double-channel structure, multi-level depth feature expression is achieved, feature mapping and classification tasks are completed through a full connection layer, the output of the two channels is spliced and fused, the model achieves multi-level depth feature expression, and therefore the accuracy and robustness of XSS attack detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-feature extraction method based on a TBA-GRU dual-channel model, belonging to the technical field. Background Art

[0002] Currently, deep learning models in XSS detection generally adopt structures such as CNN or GRU. Although these models have certain advantages in extracting text features and processing temporal information, they also have obvious deficiencies: Although CNN performs excellently in extracting local features and capturing short-distance dependence relationships, it has limitations in dealing with long-distance dependence and global context information; while recurrent neural networks such as GRU can effectively model the temporal information in sequence data, they are relatively weak in capturing complex local features and efficient parallel computing. These models usually lack an effective local and global feature fusion mechanism, resulting in deficiencies in capturing long-distance dependence or global context information, thus affecting the accuracy and robustness of XSS attack detection. Summary of the Invention

[0003] To solve the technical problems existing in the prior art, the present invention provides a multi-feature extraction method based on a TBA-GRU dual-channel model, which realizes multi-level deep feature expression by splicing and fusing the outputs of two channels, thereby improving the accuracy and robustness of XSS attack detection.

[0004] To achieve the above object, the technical solution adopted by the present invention is a multi-feature extraction method based on a TBA-GRU dual-channel model, including an input layer, a word embedding layer, a TextCNN channel, a BIGRU-multi-head attention channel, a connection layer, a Dropout layer, a fully connected layer, and an output layer. The input layer is used to receive the original XSS attack sample data; The word embedding layer is used to extract features from the XSS attack sample data and convert the input text data into a low-dimensional dense vector representation, thereby capturing the potential semantic information of the text. The generated weighted word vectors will be used as the inputs of the TextCNN channel and the BIGRU-multi-head attention channel. The TextCNN channel is used to continue to extract local features from the extracted features using convolutional kernels of different sizes and capture n-gram patterns through convolutional operations, and assign dynamic weights to each local feature through attention pooling. The BIGRU-multi-head attention channel adopts a combined structure of BiGRU and multi-head attention mechanism, which is used to identify long-distance dependence relationships and context information, and assign different weights to the temporal features. The output results of the TextCNN channel and the BIGRU-multihead attention channel are fused by concatenation; The connection layer is used to receive the fused feature data and deliver the fused feature data to the Dropout layer; The Dropout layer is used to prevent overfitting of the fused feature data; The fully connected layer is used to map and classify the fused feature data; The output layer is used to output the classification result of the XSS attack through the Sigmoid activation function.

[0005] Preferably, the TextCNN channel includes a convolutional layer and an attention pooling layer, The convolution kernel operation formula of the convolutional layer is as follows: , where, is the activation function, which performs non-linear transformation. represents the parameters of the convolution kernel, represents the word vector from the th row to the th row, is the bias term, and the extracted local feature vector is expressed as ; In the convolutional layer, the ReLU activation function is used, and its calculation formula is as follows: ; Assume that the feature matrix extracted in the convolutional layer is , where represents the th feature vector, the weight vector is , then the attention score is obtained by normalizing the weight vector. The specific calculation formula is as follows: , where, is the attention score of the th feature, indicating the importance of this feature; , where, is the final feature vector, which is obtained by weighted summation. The weight of each feature is determined by its corresponding attention score , and then the feature vectors obtained by 3 different convolution kernels are concatenated to obtain the final output of the local feature vector .

[0006] Preferably, the BIGRU-multihead attention channel includes a BIGRU module and a multihead attention mechanism module, The calculation process of the BiGRU is as follows: , , , Among them, is the state information of forward propagation, is the state information of backward propagation, is the output weight of the forward propagation hidden layer, is the output weight of the backward propagation hidden layer, is the hidden layer output bias, and G() is the operation process of the GRU; In the multi-head attention mechanism module, the input query , key and value are first linearly transformed to obtain representations in multiple subspaces: , Among them, , and are the weight matrices of each head. Then, each head independently calculates the attention output: , Among them, is the dimension of the key vector, is used to calculate the attention weight. Then, the outputs of all heads are concatenated together and passed through a linear transformation to obtain the final result, , Among them, is the number of heads, is the output linear transformation matrix. In this way, the multi-head attention can process the input sequence in parallel, capture dependencies at different levels and directions, and thus enhance the expression ability of the model.

[0007] Compared with the prior art, the present invention has the following technical effects: The present invention is composed of two parallel channels, namely the TextCNN channel and the BiGRU channel, which are respectively used to extract local features and temporal features, and the multi-head attention mechanism is combined to further enhance the attention to important features. The dual-channel structure realizes multi-level deep feature expression by splicing and fusing the output features of the TextCNN and BiGRU channels, and completes the feature mapping and classification tasks through the fully connected layer, and is significantly superior to the traditional single-structure deep learning model in multiple evaluation indexes, verifying the effectiveness in dealing with complex XSS attack scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1Schematic diagram of the framework structure of the present invention. Detailed implementation manners

[0009] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0010] As Figure 1 shown, a multi-feature extraction method based on a TBA-GRU dual-channel model includes an input layer, a word embedding layer, a TextCNN channel, a BIGRU-multi-head attention channel, a connection layer, a Dropout layer, a fully connected layer, and an output layer. The input layer is used to receive the original XSS attack sample data; The word embedding layer is used to extract features from the XSS attack sample data and convert the input text data into a low-dimensional dense vector representation, so as to capture the potential semantic information of the text. The generated weighted word vectors will be used as the input of the TextCNN channel and the BIGRU-multi-head attention channel; The TextCNN channel is used to continue to extract local features from the extracted features using convolution kernels of different sizes and capture n-gram patterns through convolution operations, and assign dynamic weights to each local feature through attention pooling; The BIGRU-multi-head attention channel adopts a combined structure of BiGRU and multi-head attention mechanism, which is used to identify long-distance dependence relationships and context information, and assign different weights to the temporal features; The output results of the TextCNN channel and the BIGRU-multi-head attention channel are fused through concatenation; The connection layer is used to receive the fused feature data and send the fused feature data to the Dropout layer; The Dropout layer is used to prevent overfitting of the fused feature data; The fully connected layer is used to map and classify the fused feature data; The output layer is used to output the classification result of the XSS attack through the Sigmoid activation function.

[0011] In view of the problems of the traditional XSS detection model, such as the lack of effective integration of local and global information, insufficient attention to key information, and insufficient ability to model temporal features, the TBA-GRU model is proposed. This model adopts a dual-channel architecture and a multi-head attention mechanism. The first channel is based on the optimized TextCNN, focusing on local feature extraction and enhancing the attention to key features through attention pooling. The second channel uses the BiGRU structure to capture the temporal information dependence, and the multi-head attention mechanism enhances the sensitivity of the model to important temporal features and reduces the interference of redundant information. Finally, the outputs of the two channels are fused by concatenation, and feature mapping and classification are completed through a fully connected layer.

[0012] The input layer of the model first receives the original XSS attack sample data and extracts features through the word embedding layer, converting the input text data into a low-dimensional dense vector representation to capture the potential semantic information of the text. The generated weighted word vectors will be used as the input of two parallel channels to further extract features from the text. The first channel is based on the optimized TextCNN, using convolutional kernels of different sizes (3, 4, 5) to extract local features from the text and capture n-gram patterns through convolutional operations. In the pooling layer, the traditional max pooling is replaced by attention pooling, which can focus more on the features crucial for XSS attack detection by assigning dynamic weights to each local feature, thus avoiding information loss or interference from irrelevant features. The second channel adopts a combined structure of BiGRU and multi-head attention mechanism. BiGRU is good at modeling long-distance dependencies and context information in the text, while the multi-head attention mechanism enhances the model's ability to focus on key information by assigning different weights to temporal features, reducing the impact of redundant information, and thus improving the processing ability and understanding depth of temporal data. Finally, the output results of the two channels are fused by concatenation, integrating the advantages of local features and temporal features. The fused features are mapped and classified through a fully connected layer, and finally the classification result of the XSS attack is output through the Sigmoid activation function.

[0013] Among them, for the TextCNN channel: The Convolutional Neural Network (CNN) was initially applied to the field of images. By performing convolutional operations on images, it can effectively extract local features to achieve tasks such as monitoring and recognition. The core idea of TextCNN is to extract local features in the text sequence through the convolutional neural network, and use convolutional kernels of different sizes to capture context information at multiple granularities. The main structure of the TextCNN model includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer. First, the input layer converts the original text into a word vector matrix of a fixed dimension through the word embedding method. Then, the convolutional layer slides along the text sequence direction with multiple convolutional kernels of different sizes to capture n-gram features and generate multiple feature maps. The pooling layer uses the attention pooling mechanism to weight the convolutional feature maps. By assigning different weights to different features, the model can pay more attention to important features, thereby improving the ability to extract key information and reducing the interference of redundant information. Finally, the fully connected layer integrates the weighted features output by the pooling layer and passes them to the classifier for final prediction.

[0014] Convolutional layer: After converting the text into a vector sequence through word embedding, the convolutional layer performs convolutional operations on the input sequence using convolutional kernels of different sizes. In this model, convolutional kernels of sizes 3, 4, and 5 are used respectively. The role of these convolutional kernels is to extract features from different context windows to identify semantic patterns at different granularities. Each convolutional kernel slides in the text sequence to perform convolutional operations, generate feature maps, and can effectively capture the local dependencies between words. The operation process of the convolutional kernel is shown in formula (4.1).

[0015] , Among them, is the activation function, which performs a non-linear transformation. represents the parameters of the convolutional kernel, represents from the th row to the th row of the word vectors, is the bias term, and the extracted local feature vector is represented as .

[0016] Activation functions: Common non-linear activation functions include Sigmoid, Tanh, and ReLU. The Sigmoid function compresses the input value between 0 and 1, while Tanh maps the input to the range from -1 to 1. Although these two functions can alleviate the vanishing gradient problem to some extent, they still face the problem of gradient dispersion, resulting in unsatisfactory training effects. In contrast, the ReLU function directly returns the input value when it is positive, and outputs zero when the input is negative. The ReLU function can effectively accelerate convergence because its gradient in the positive interval is always 1, avoiding large vanishing gradient phenomena. Therefore, this model selects to use the ReLU activation function in the convolutional layer, and the specific formula is as follows: .

[0017] Attention pooling layer: The role of the pooling layer is to reduce the dimension and aggregate features of the output of the convolutional layer, reducing redundant features and computational complexity. However, the traditional max pooling method extracts features by selecting the maximum value in the region. Although it is simple and efficient, it ignores other potentially useful information and cannot effectively capture the global context, resulting in insufficient recognition ability of the model for fine-grained features. Therefore, attention pooling is proposed to replace max pooling. The attention pooling mechanism dynamically weights the features, enabling the model to automatically identify the most important features according to the context, thereby more accurately extracting the detailed information in XSS attacks and improving the detection ability of the model in complex and variant attack scenarios. Suppose the feature matrix extracted in the convolutional layer is , where represents the th feature vector, and the weight vector is , then the attention score is obtained by normalizing the weight vector. The specific calculation formula is as follows: , where, is the attention score of the th feature, indicating the importance of this feature; , where, is the final feature vector, obtained by weighted summation. The weight of each feature is determined by its corresponding attention score . Then, the feature vectors obtained from 3 different convolutional kernels are concatenated to obtain the final output of the local feature vector .

[0018] BiGRU-Multi-Head Attention Channel: The BiGRU-multi-head attention channel combines BiGRU and the multi-head attention mechanism. BiGRU processes the context information of the sequence simultaneously through a bidirectional structure to capture long-range dependencies. The multi-head attention mechanism weights the extracted temporal features to enhance the attention to key information and reduce the interference of redundant information.

[0019] A unidirectional GRU can only update its state using the historical information at the current time step and previous steps, making it difficult to comprehensively capture the global dependencies in the input sequence. Relying solely on unidirectional information propagation may lead to limitations in feature extraction. In BiGRU, the network structure is extended to be bidirectional. That is, when processing the input sequence, it considers not only the information flow from front to back but also the information flow from back to front. Therefore, BiGRU enhances the model's performance in sequence data processing by introducing two independent neural network structures, the forward GRU and the backward GRU, and concatenating or merging the hidden states in both directions as the final output. Bi-GRU can capture the dependencies in the context simultaneously, thus improving the ability to model temporal data. The computational process of BiGRU is as follows: , , , where, is the state information of forward propagation, is the state information of backward propagation, is the output weight of the forward propagation hidden layer, is the output weight of the backward propagation hidden layer, is the hidden layer output bias, and G() is the operation process of GRU.

[0020] The multi-head attention mechanism has been widely used in deep learning. Its core idea is to parallelly compute multiple independent attention heads to focus on various features at different positions in the input sequence. This mechanism enables the model to capture different dependencies in the input data from multiple perspectives, thus improving the ability to model complex patterns and long-term and short-term dependencies. Different from the traditional single-head attention mechanism, which usually only focuses on a specific feature or local feature of the data, the multi-head attention can process information from multiple dimensions through parallel computing and capture richer context information. This design enables the model to better capture the potential complex relationships in the sequence and significantly enhances the model's representation ability.

[0021] In the multi-head attention mechanism, the input query , key and value First, it is mapped to different subspaces through multiple independent linear transformations to obtain multiple different representations. Each subspace corresponds to an independent attention head. When calculating, each head calculates the attention weights according to its own queries, keys, and values, focusing on different aspects of the input data. The outputs of each head are finally concatenated together and passed through a linear transformation to obtain the final output. This process enables the model to capture both long-distance dependencies and local detailed features in the sequence simultaneously, thus providing a richer representation for tasks that require considering both global context and local information. The multi-head attention mechanism not only improves the performance of the model but also significantly enhances the computational efficiency through parallel computing. Specifically, the input queries , keys and values are first linearly transformed to obtain representations in multiple subspaces: , where , and are the weight matrices for each head. Then each head independently calculates the attention output: , where is the dimension of the key vector, is used to calculate the attention weights. Then the outputs of all heads are concatenated together and passed through a linear transformation to obtain the final result, , where is the number of heads, is the linear transformation matrix of the output. In this way, the multi-head attention can process the input sequence in parallel, capture dependencies at different levels and directions, and thus enhance the expressive power of the model.

[0022] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the scope of the present invention.

Claims

1. A multi-feature extraction method based on a TBA-GRU dual-channel model, characterized in that: It includes input layer, word embedding layer, TextCNN channel, BIGRU-multi-head attention channel, connection layer, Dropout layer, fully connected layer and output layer. The input layer is used to receive original XSS attack sample data; The word embedding layer is used to extract features from XSS attack sample data and convert the input text data into a low-dimensional dense vector representation, thereby capturing the potential semantic information of the text. The generated weighted word vector will be used as the input of the TextCNN channel and the BIGRU-multi-head attention channel; The TextCNN channel is used to continue extracting local features from the extracted features using convolution kernels of different sizes and capture n-gram patterns through convolution operations, and assign dynamic weights to each local feature in the attention pooling; The BIGRU-multi-head attention channel adopts a combination structure of BiGRU and multi-head attention mechanism to identify long-distance dependencies and contextual information, and assigns different weights to temporal features; The output results of the TextCNN channel and the BIGRU-multi-head attention channel are concatenated for feature fusion; The connection layer is used to receive the fused feature data and transmit the fused feature data to the Dropout layer; The Dropout layer is used to prevent overfitting of the fused feature data; The fully connected layer is used to map and classify the fused feature data; The output layer is used to output the classification result of the XSS attack through the Sigmoid activation function.

2. The multi-feature extraction method based on the TBA-GRU dual-channel model according to claim 1, characterized in that: The TextCNN channel includes convolutional layers and attention pooling layers. The convolution kernel calculation formula of the convolution layer is as follows: , in, is the activation function that performs a nonlinear transformation. represents the parameters of the convolution kernel, Indicates that from Go to The word vector of the row, is the bias term, and the extracted local feature vector Expressed as ; The ReLU activation function is used in the convolutional layer, and its calculation formula is as follows: ; Assume that the feature matrix extracted in the convolutional layer is ,in Indicates feature vectors, and the weight vector is , then the attention score It is obtained by normalizing the weight vector. The specific calculation formula is as follows: , in, It is The attention score of a feature, indicating the importance of the feature; , in, is the final feature vector, obtained by weighted summation, where the weight of each feature is determined by its corresponding attention score Decide and then take the feature vectors obtained by three different convolution kernels Splice to get the final local feature vector output .

3. The multi-feature extraction method based on the TBA-GRU dual-channel model according to claim 1, characterized in that: The BIGRU-multi-head attention channel includes a BIGRU module and a multi-head attention mechanism module. The calculation process of BIGRU is as follows: , , , in, is the state information of the forward propagation, is the state information of the back propagation, is the output weight of the hidden layer in the forward propagation, is the output weight of the back-propagation hidden layer, is the hidden layer output bias, G() is the operation process of GRU; In the multi-head attention mechanism module, the input query ,key Sum First, we obtain the representation of multiple subspaces through linear transformation: , in, , and is the weight matrix of each head, and then each head calculates the attention output independently: , in, is the dimension of the key vector, It is used to calculate the attention weights, and then the outputs of all heads are concatenated together and the final result is obtained through a linear transformation. , in, is the number of heads, is the linear transformation matrix of the output. In this way, multi-head attention can process the input sequence in parallel and capture dependencies at different levels and directions, thereby enhancing the expressiveness of the model.

Citation Information

Patent Citations

  • Method and device for identifying XSS attack, and computer readable storage medium

    CN109388943A

  • Construction method, device and application of XSS attack detection model

    CN116910749A

  • Text sentiment classification method based on syntactic dependency relationship and attention mechanism

    CN117951304A

  • Parallel dual-channel medical consultation supervision detection method and model in combination with theme information

    CN118230974A

  • Chinese fraudulent language detection method and system based on two channels

    CN119378556A

Cited By

  • Intelligent modularized prefabricated cabin type tunnel fire station terminal dynamic state based on edge calculation and multi-source perception

    CN120977063A