SQL injection attack detection method based on iPGRU model

By adopting the iPGRU model in SQL injection attack detection, combining SSA-Feature feature extraction, Bi-GRU timing modeling and iPNN incremental learning mechanism, the efficiency and accuracy of SQL injection attack detection in the existing technology are solved, and excellent performance on multiple evaluation indicators is achieved.

CN120185897APending Publication Date: 2025-06-20SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510390799.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to detect SQL injection attacks efficiently and accurately, especially when dealing with timing characteristics and new attack patterns.

Method used

Using the detection method based on the iPGRU model, the SQL injection attack dataset is featured by SSA-Feature feature extraction method, and combined with the timing modeling of Bi-GRU and the incremental learning mechanism of iPNN, efficient detection of SQL injection attacks is achieved.

Benefits of technology

Excellent in accuracy, accuracy, recall and false alarm rate, achieving efficient and accurate detection of SQL injection attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_16
    Figure SMS_16
  • Figure SMS_17
    Figure SMS_17
  • Figure SMS_18
    Figure SMS_18
Patent Text Reader

Abstract

The invention relates to an SQL injection attack detection method based on an iPGRU model, and belongs to the technical field of network security, and the method specifically comprises the following steps: a, constructing an SQL injection attack data set, and carrying out decoding, redundant column deletion and normalization processing on the data set to obtain a target data set; b, feature extraction is conducted on the target data set through an SSA-Feature feature extraction method, features of an SQL injection attack sample are comprehensively captured, SQL injection attack detection can be efficiently and accurately achieved, and the accuracy rate, the precision rate, the recall rate and the false alarm rate are excellent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting SQL injection attacks based on the iPGRU model, belonging to the technical field of network security. Background Art

[0002] SQL injection is to insert SQL commands into the query string submitted in a Web form or input domain name or page request, and finally deceive the server into executing malicious SQL commands. It takes advantage of existing applications to inject malicious SQL commands into the background database, and can obtain the database on a website with security vulnerabilities by inputting malicious SQL statements in the Web form. Summary of the Invention

[0003] To solve the technical problems existing in the prior art, the present invention provides a method for detecting SQL injection attacks based on the iPGRU model, which can efficiently and accurately detect SQL injection attacks and perform excellently in terms of accuracy, precision, recall rate and false alarm rate.

[0004] To achieve the above object, the technical solution adopted by the present invention is a method for detecting SQL injection attacks based on the iPGRU model, which is operated according to the following steps. a. Construct a SQL injection attack data set, and perform decoding, deleting redundant columns and normalization processing on the data set to obtain a target data set; b. Then, use the SSA-Feature feature extraction method to extract features from the target data set to comprehensively capture the features of SQL injection attack samples; The SSA-Feature feature extraction method is as follows: The first step is to extract statistical features from SQL requests using the TF-IDF algorithm, quantify the importance of keywords by calculating the term frequency-inverse document frequency, highlight high-frequency abnormal words, and perform vectorization operations on the text with the help of TfidfVectorizer to extract highly discriminative features; The second step is to model from the semantic level using the Word2Vec model, learn the distributed representation of words through the Skip-gram architecture, capture context-sensitive word vectors, and obtain a Word2Vec feature vector that can reflect the semantic information of the text by taking the average value of its word vectors; The third step is to compress the Word2Vec feature vector through an Autoencoder encoder to extract a more concise and efficient latent representation; In the fourth step, the features processed in the second and third steps are concatenated, and then a temporal modeling is carried out on the concatenated features through the Bi-GRU module. At the same time, using the incremental learning mechanism of the iPNN module, when new data emerges continuously, the class probability density estimation can be dynamically updated without full retraining of the entire model.

[0005] Preferably, the Bi-GRU module consists of two GRU units. One processes the sequence backward from the starting position of the sequence, and the other processes it forward from the end of the sequence. Finally, the hidden states in the two directions are concatenated to comprehensively capture the context information in the input sequence. In the forward GRU, given an input sequence , the update process of the hidden state ht at each time step is determined by the following formula: , , , , where rt is the reset gate, ut is the update gate, is the candidate hidden state, ht is the current hidden state. The reverse GRU has a similar structure to the forward GRU, but it starts calculating from the last moment of the sequence. By processing the sequence backward, it captures future context information. In the Bi-GRU model, the forward and reverse hidden states are concatenated into a new hidden state at each time step: .

[0006] Preferably, the iPNN module includes an input layer, a pattern layer, a summation layer, and an output layer. The input layer is used to receive the input data vector x, representing the features of each sample. The pattern layer is responsible for calculating the probability density of each class. Usually, a kernel function is used to calculate the similarity. For the probability density estimation of a certain class Ck, it is calculated by the following formula: , where x is the new sample to be classified, xi is the feature vector of the i-th sample, Nk is the number of samples belonging to class Ck, and σ is the bandwidth parameter of the kernel function, usually using a Gaussian kernel. The summation layer is used to sum the probability densities of each class to obtain the probability output of each class. The output layer is used to finally output the class label according to the calculated probability. When new data samples arrive, the iPNN module adjusts the model by updating the probability density estimates of classes without retraining the entire network. Assuming there is already a training sample set Dold, and now a new sample xnew is received, the goal of the incremental learning mechanism is to update the probability density estimate of class Ck based on the new sample. The incremental update formula is as follows: , where, and are the weights of the old data and the new data, and are the probabilities of the old data and the new data for class Ck, respectively.

[0007] The probabilities of the old data and the new data for class Ck.

[0008] Compared with the prior art, the present invention has the following technical effects: The iPGRU model adopted by the present invention shows significant performance advantages in the SQL injection attack detection task. Especially in dealing with time series features and new attack patterns, the iPGRU model can more comprehensively capture the complex features of SQL injection attacks by combining the time series modeling ability of Bi - GRU and the incremental learning mechanism of iPNN, thus performing excellently in terms of accuracy, precision, recall rate and false alarm rate, and then achieving efficient and accurate detection of SQL injection attacks. Specific embodiments

[0009] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below in conjunction with examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0010] A method for detecting SQL injection attacks based on the iPGRU model, the operation steps are as follows: a. Construct a SQL injection attack data set, and perform decoding, deleting redundant columns and normalization processing on the data set to obtain a target data set; b. Then, use the SSA - Feature feature extraction method to extract features from the target data set to comprehensively capture the features of SQL injection attack samples.

[0011] For the SQL injection attack dataset, due to the possible encoding interference in SQL query statements, recursive decoding is performed to restore semantic features, and regular expressions are used to filter out non-ASCII characters and redundant spaces, retaining the context structure of key operators to ensure the semantic integrity of the attack statements. Subsequently, based on feature correlation analysis, redundant columns are removed, and only the core fields strongly related to SQL injection detection are retained to reduce the risk of the curse of dimensionality and improve computational efficiency. After the above data cleaning, the samples are labeled, with normal HTTP requests marked as 0 and SQL injection attack requests marked as 1, constructing the label system required for supervised learning.

[0012] To further improve data balance, a stratified sampling strategy is used to divide the dataset into a training set and a test set in a ratio of 8:2, ensuring that the distribution ratio of normal samples and attack samples remains consistent in the subsets. For the training set, one-hot encoding is performed on discrete features to generate binary feature vectors, avoiding the problem of pseudo-order introduced by numerical labels. For continuous features, min-max normalization is used to eliminate the dimensionality difference, and its formula is defined as follows: , where Xmin and Xmax are the minimum and maximum values of the feature column respectively, and all values are linearly mapped to the [0, 1] interval after normalization. The test set uses the same encoding and normalization parameters as the training set to avoid the risk of data leakage.

[0013] Through the above preprocessing process, the standardization and consistency of the data can be significantly improved, providing a reliable basis for subsequent feature modeling and algorithm optimization, and ensuring the scientificity and repeatability of experimental results.

[0014] Feature extraction: Feature extraction is one of the core tasks, aiming to extract highly distinguishable features from SQL injection attack samples and transform these features into a format suitable for model training. A SSA-Feature feature extraction method is proposed. The feature extraction process is as follows: First, the TF-IDF algorithm is used to extract statistical features from SQL requests. By calculating the term frequency-inverse document frequency to quantify the importance of keywords, high-frequency abnormal words are highlighted. The TfidfVectorizer is used to perform vectorization operations on the text, which can keenly highlight those words that appear frequently in the current document but are relatively rare in the entire document collection, and then extract highly discriminative features. TF-IDF features provide global importance information for the subsequent model from the unique perspective of statistical frequency.

[0015] Next, the Word2Vec model is used to model at the semantic level. Through the Skip-gram architecture, the distributed representation of words is learned to capture context-sensitive word vectors. Word2Vec divides the text into a list of sentences in detail and then trains the model to accurately capture the semantic associations between words, thereby generating word vectors. For each text sample, by taking the average of its word vectors, a Word2Vec feature vector that can reflect the semantic information of the text is obtained. This makes texts with similar semantics closer in position in the vector space, providing rich and accurate semantic feature support for the subsequent model.

[0016] To reduce the redundancy of high-dimensional semantic features, a dimensionality reduction module based on an autoencoder is designed. The autoencoder consists of two main parts: an encoder and a decoder. The encoder compresses the 300-dimensional Word2Vec features into a 256-dimensional latent space. This compression operation not only significantly reduces the dimension of the data, greatly reducing the computational complexity, but also greatly helps to improve the generalization ability of the model and effectively avoid the problem of overfitting. The decoder reconstructs the original features by minimizing the mean square error. After training, the low-dimensional features output by the encoder are concatenated with the TF-IDF features, and dimension alignment and fusion are performed through a fully connected layer.

[0017] After the above feature compression is completed, the Bi-GRU module conducts temporal modeling on the compressed Word2Vec features. The SQL request text essentially has obvious sequential characteristics, and the order of its elements contains rich semantic logic. Bi-GRU adopts a bidirectional learning method, considering both forward and backward context information: forward learning gradually captures the association between each element and subsequent elements from the beginning of the text, and backward learning analyzes the potential connections between elements from the end of the text forward. Such bidirectional modeling ensures that the model can comprehensively and deeply extract the deep semantic features in the text, significantly improving the accuracy of capturing temporal information and maintaining robustness even in the face of complex deformed or disguised attack methods.

[0018] At the same time, the iPNN module plays a key role in dynamic adaptation and optimization in the entire model. Considering the actual need for the continuous evolution of SQL injection attack methods and the emergence of new variants in the actual network environment, iPNN uses an incremental learning mechanism to dynamically update the category probability density estimation when new data emerges, without the need for full retraining of the entire model. Specifically, when new SQL request data appears, iPNN can quickly identify the differences between the new attack patterns and the existing categories, and timely adjust the probability distributions of relevant categories to ensure the sensitivity and adaptability of the model to new attacks, thereby greatly improving the real-time performance and overall detection performance of the model.

[0019] During the model training process, the design of the loss function is equally crucial. This paper adopts a hybrid strategy that combines traditional cross-entropy loss with a dedicated loss for incremental learning: the cross-entropy loss is used to measure the difference between the model's prediction results and the true labels, guiding the update of overall parameters; while the incremental learning loss focuses on constraining the change in the model's performance when new data arrives, ensuring that the model can adjust parameters in a timely manner during the continuous reception of new samples and always maintain a sensitive response to new attack patterns. In addition, since the number of normal requests in the actual dataset is much larger than that of attack requests, by reasonably designing the loss function, the data imbalance problem can be effectively alleviated, enabling the model to pay more attention to the attack samples of the minority class, thereby improving the detection effect as a whole.

[0020] Among them, the Bi-GRU module is a variant of the recurrent neural network based on GRU, which mainly captures more context information by simultaneously utilizing the forward and backward information of the input sequence. Traditional GRU can only rely on past information when processing sequence data, while Bi-GRU makes the output at each time step combine the forward and backward contexts by using forward and backward GRUs in parallel, thus enhancing the model's modeling ability. The Bi-GRU structure consists of two GRU units, one processes the sequence backward from the starting position of the sequence, and the other processes the sequence forward from the end of the sequence. Finally, the hidden states in both directions are concatenated to comprehensively capture the context information in the input sequence.

[0021] In the forward GRU, given an input sequence The update process of the hidden state ht at each moment is determined by the following formula: , where rt is the reset gate, ut is the update gate, is the candidate hidden state, ht is the current hidden state. The backward GRU has a similar structure to the forward GRU, but it starts calculating from the last moment of the sequence. By processing the sequence backward, it captures future context information. In the Bi-GRU model, the forward and backward hidden states are concatenated into a new hidden state at each time step: .

[0022] This approach enables the Bi-GRU to utilize both forward and backward context information simultaneously, significantly enhancing the model's capabilities. The advantage of Bi-GRU lies in its ability to capture the semantic information of sequences more comprehensively than traditional unidirectional GRUs, making it particularly suitable for text and time series data processing. Additionally, compared to LSTMs, GRUs themselves have fewer parameters and a more concise structure, making Bi-GRU computationally more efficient while providing similar performance to LSTMs.

[0023] The iPNN module is a variant based on the classical probabilistic neural network, designed specifically for processing streaming data such as time series data and real-time data streams. The core advantage of iPNN is its ability to perform incremental updates on new data in real-time without retraining the entire model, effectively addressing the real-time changes in data. This makes iPNN particularly suitable for dynamic data environments and enables continuous optimization and adaptation while continuously receiving new data.

[0024] iPNN combines probabilistic inference methods and incremental learning mechanisms, achieving continuous learning and dynamic adaptation of the model by performing incremental updates on each incoming new data. Different from traditional PNN models, iPNN does not need to retrain the entire model every time new data is received. Instead, it gradually adjusts the existing model, enabling the model to be optimized and updated according to new data. Therefore, iPNN can efficiently process real-time data streams without worrying about waste of computing resources and long training times.

[0025] The core idea of the incremental learning mechanism is that when new data samples arrive, iPNN adjusts the model by updating the probability density estimates of classes without retraining the entire network. Suppose there is already a training sample set Dold, and now a new sample xnew is received. The goal of the incremental learning mechanism is to update the probability density estimate of class Ck according to the new sample. The incremental update formula is as follows: , where, and are the weights of the old data and new data, and are the probabilities of the old data and new data for class Ck respectively. Through this incremental update mechanism, iPNN can quickly adapt to the changes in new data without losing the memory of old data, ensuring the continuous effectiveness of the model in a real-time changing environment.

[0026] The iPNN module mainly consists of the following layers: (1) Input layer: This layer accepts the input data vector x, representing the features of each sample, (2) Pattern layer: This layer is responsible for calculating the probability density of each category. Usually, a kernel function (such as Gaussian kernel) is used to calculate the similarity. The probability density estimation for a certain category Ck can be calculated by the following formula: , where x is the new sample to be classified. xi is the feature vector of the i-th sample. Nk is the number of samples belonging to category Ck. σ is the bandwidth parameter of the kernel function, and usually the Gaussian kernel is used.

[0027] (3) Summation layer: Sum the probability densities of each class to obtain the probability output of each class.

[0028] (4) Output layer: Finally, output the class label according to the calculated probability.

[0029] In this way, the iPGRU model based on SSA-Feature feature extraction constructs a detection framework that can not only adapt to the changing attack environment but also efficiently capture the features of complex SQL injection attacks by fusing statistical, semantic, and latent structure features, and combining the incremental learning advantage of iPNN and the bidirectional time series modeling ability of Bi-GRU.

[0030] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the scope of the present invention.

Claims

1. A SQL injection attack detection method based on the iPGRU model, characterized by: Follow the steps below, a. Construct a SQL injection attack dataset, decode it, delete redundant columns, and normalize it to obtain the target dataset; b. Then, the SSA-Feature feature extraction method is used to extract features from the target data set to fully capture the features of SQL injection attack samples; The SSA-Feature feature extraction method is: In the first step, we use the TF-IDF algorithm to extract statistical features from SQL requests, quantify the importance of keywords by calculating the word frequency-inverse document frequency, highlight high-frequency abnormal words, and use TfidfVectorizer to vectorize the text and extract highly discriminative features. The second step is to use the Word2Vec model to build a semantic model. The Skip-gram architecture is used to learn the distributed representation of vocabulary, capture context-sensitive word vectors, and average the word vectors to obtain the Word2Vec feature vector that can reflect the semantic information of the text. The third step is to compress the Word2Vec feature vector through the Autoencoder encoder to extract a more concise and efficient potential representation; In the fourth step, the features processed in the second and third steps are concatenated, and then the Bi-GRU module is used to perform time series modeling on the concatenated features. At the same time, the incremental learning mechanism of the iPNN module is used to dynamically update the category probability density estimate when new data continues to emerge, without the need to retrain the entire model.

2. According to claim 1, a SQL injection attack detection method based on the iPGRU model is characterized in that: The Bi-GRU module consists of two GRU units, one processes the sequence backward from the beginning of the sequence, and the other processes it forward from the end of the sequence, and finally concatenates the hidden states in the two directions, thereby fully capturing the contextual information in the input sequence; In the forward GRU, given an input sequence , the update process of the hidden state ht at each moment is determined by the following formula: , Among them, rt is the reset gate, ut is the update gate, is the candidate hidden state, ht is the current hidden state, and the reverse GRU is similar to the forward GRU structure, but it is calculated from the last moment of the sequence, and captures future context information by reverse processing of the sequence. In the Bi-GRU model, the forward and backward hidden states are concatenated into a new hidden state at each time step: .

3. The SQL injection attack detection method based on the iPGRU model according to claim 1, characterized in that: The iPNN module includes an input layer, a pattern layer, a summation layer, and an output layer. The input layer is used to accept the input data vector x, which represents the characteristics of each sample; The pattern layer is responsible for calculating the probability density of each category. The kernel function is usually used to calculate the similarity. The probability density estimation of a certain category Ck is calculated by the following formula: , Among them, x is the new sample to be classified, xi is the feature vector of the i-th sample, Nk is the number of samples belonging to category Ck, σ is the bandwidth parameter of the kernel function, usually a Gaussian kernel is used; The summation layer is used to sum the probability density of each class to obtain the probability output of each category; The output layer is used to finally output the category label according to the calculated probability; When new data samples arrive, the iPNN module adjusts the model by updating the probability density estimate of the category without retraining the entire network. Assuming that there is already a training sample set Dold, and now a new sample xnew is received, the goal of the incremental learning mechanism is to update the probability density estimate of the category Ck according to the new sample. The incremental update formula is as follows: , in, and is the weight of old data and new data, and are the probabilities of the old data and the new data for category Ck respectively.