SQL (Structured Query Language) injection detection method and device based on SQL context semantic analysis
Through the detection method based on SQL context semantic analysis, using improved model generation and feature extraction technology, the problems of sample data requirements, interpretability and sample imbalance in SQL injection detection are solved, and efficient and robust SQL injection detection is achieved.
Patent Information
- Application Number
- CN202411905202.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces the problems of massive sample data requirements, insufficient model interpretability and imbalance of samples in different categories in SQL injection detection, making it difficult to effectively detect and defend against complex SQL injection attacks.
Using a detection method based on SQL context semantic analysis, SQL injection statements are generated through the improved augmented attack tree model, combined with the improved TextCNN and LSTM models for feature extraction and pattern recognition, and an injection detection model is built to realize SQL injection detection.
This method can systematically cover various SQL injection attack types, improve detection accuracy and robustness, enhance the recognition ability of complex injection patterns, and improve the generalization ability of the model and anti-attack performance.
Smart Images

Figure CN120068064A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of database SQL injection detection and defense, and particularly to a SQL injection detection method and device based on SQL context semantic analysis. Background Art
[0002] SQL (Structured Query Language) injection attack is one of the most serious database attack means in current network security. SQL Injection is an attack method that embeds user data into database instructions through the external interface of the database system, aiming to invade the database and even the operating system. Such attacks are achieved by directly or indirectly inserting malicious code into user input linked to SQL commands. For example, an attacker may prematurely terminate a text string and append a new SQL command, or use the comment symbol "--" to abort the original SQL code, causing subsequent instructions to be ignored. The core of SQL injection attack lies in exploiting SQL syntax vulnerabilities to manipulate data and perform unauthorized queries by embedding SQL statements.
[0003] Compared with traditional vulnerabilities, the risk of SQL injection is more serious. It can bypass the firewall and directly access the database, and may even obtain the access right to the database server. In a specific environment, the risk of SQL injection is higher than other vulnerability types. Its characteristics mainly include the following points: ① Strong concealment: SQL injection uses Web vulnerabilities to attack application systems and is not easily detected. Since network firewalls are usually completely open to HTTP / HTTPS, coupled with the diversification of Web attack types, traditional signature-based intrusion detection systems (IDS) are difficult to cope with. ② Fast attack speed: A single SQL injection attack may be completed within seconds to minutes, sufficient to achieve data theft, Trojan implantation, or even obtain control of the entire database or server, making it almost impossible for manual response. ③ Great harmfulness: Many enterprises and institutions (such as banks, telecommunications, governments, e-commerce platforms) store confidential data in the background database. Once an attack occurs, it may lead to data leakage and tampering, thereby causing serious property and reputation losses to enterprises or individuals. In addition, the attack and tampering of government websites may cause social instability and even be exploited by external forces. ④ Serious tangible and intangible losses: SQL injection attacks pose a great threat to various enterprises and government agencies. Especially when such incidents occur in listed companies, it may trigger stock price fluctuations and reputation damage, resulting in inestimable economic and social impacts.
[0004] Although the industry currently mainly addresses the SQL injection risk by optimizing code quality or strengthening protection, these methods do not fundamentally solve the problem. For example, existing program-based SQL injection vulnerability detection methods intercept and integrate HTTP packets before external data enters the database. Although they can mitigate the SQL injection risk to a certain extent, they often have a negative impact on external access efficiency.
[0005] In the research on preventing SQL injection, the main means currently applied include parameterized queries, input validation and filtering, using ORM frameworks, and the principle of least privilege. Although these methods can improve the security of the database to a certain extent, they also have some drawbacks: ① Parameterized queries are a very effective defense method that prevents SQL injection through pre-compiled SQL statements and parameterized inputs. However, this method may require developers to have a relatively high technical level and use parameterized queries for every database operation, which may increase the complexity of development and maintenance in actual operations. ② Input validation and filtering can prevent inputs that do not conform to the expected format from entering the system, but this method may miss certain attacks due to improper rule settings or affect the user experience due to overly strict rules. ③ ORM frameworks reduce the risk of SQL injection by automatically handling parameterized queries, but they may be attacked due to vulnerabilities in the framework itself or improper configuration. At the same time, developers may not have a good understanding of the underlying SQL operations, making it impossible to effectively defend against SQL injection in some cases. ④ The principle of least privilege can reduce the damage after an attack is successful, but implementing it may increase the complexity of system management and maintenance.
[0006] The main difficulties in SQL injection detection lie in the diversity and concealment of attack methods. Attackers can construct various complex SQL statements, attempt to bypass the input validation of the application program, and inject malicious code into the database. These attack methods include, but are not limited to, numeric, character, and search-based injections, as well as injection attacks submitted through different methods such as GET, POST, Cookie, and HTTP request headers. In addition, attackers may also use attack methods with different execution effects such as error-based injection, blind injection, union query injection, and heap query injection, making detection even more difficult. Deep learning models can automatically extract features and establish a discrimination model by learning a large number of normal and malicious SQL statements, thereby improving the accuracy and robustness of detection.
[0007] When using deep learning for SQL injection detection, its feasibility is mainly reflected in its ability to process and learn a large amount of data, automatically extract features, and identify complex attack patterns. Deep learning models, especially convolutional neural networks and recurrent neural networks, have achieved remarkable results in the field of natural language processing. They can capture local correlations and long-term dependencies in SQL statements. In addition, by using pre-trained language models and combining them with deep learning models, the detection ability of the model for SQL injection attacks can be further improved. However, the advancement of deep learning models in SQL injection detection also faces challenges, including the need for a large amount of sample data, insufficient interpretability of the model, and imbalance of different categories of samples. Summary of the Invention
[0008] In order to solve the technical problems existing in the prior art that the advancement of deep learning models in SQL injection detection faces challenges, including the need for a large amount of sample data, insufficient interpretability of the model, and imbalance of different categories of samples, the embodiments of the present invention provide a SQL injection detection method and device based on SQL context semantic analysis. The technical solutions are as follows:
[0009] On the one hand, a SQL injection detection method based on SQL context semantic analysis is provided. This method is implemented by a SQL injection detection device, and the method includes:
[0010] S1. Obtain SQL injection attack samples and normal query samples, use an improved augmented attack tree SQL injection attack model to process the SQL injection attack samples to generate SQL injection statements, and construct a SQL data set according to the SQL injection statements and normal query samples.
[0011] S2. Use an improved TextCNN model to extract features from the SQL injection statements.
[0012] S3. Train and optimize an improved LSTM model according to the SQL data set and the extracted features to obtain an injection detection model.
[0013] S4. Obtain the input data to be detected, and detect the input data according to the injection detection model to obtain a SQL injection detection result.
[0014] Optionally, using an improved augmented attack tree SQL injection attack model in S1 to process the SQL injection attack samples to generate SQL injection statements includes:
[0015] S11. Use an attack tree model to formally describe the attack path from the attack starting point to the target achievement.
[0016] S12. For the formalized attack path, define the characteristics of SQL injection attacks using a formal language and generate SQL injection statements.
[0017] Optionally, defining the characteristics of SQL injection attacks using a formal language and generating SQL injection statements in S12 includes:
[0018] According to the characteristics of SQL injection attacks, construct a set of attack payloads, define logical connectives to combine the attack payloads, and generate SQL injection statements.
[0019] Among them, defining logical connectives includes: defining || as the OR operation of attack payloads, as shown in the following formula (1):
[0020] S(1)‖S(2) = {x|x ∈ S(1) ∨ x ∈ S(2)} (1)
[0021] Among them, S(1)‖S(2) means either one of the two attack payloads S(1) and S(2) is taken.
[0022] Define && as the AND operation of attack payloads, as shown in the following formula (2):
[0023] S(1)&&S(2) = {x,y|x ∈ S(1) ∧ y ∈ S(2)} (2)
[0024] Among them, S(1)&&S(2) means that the two attack payloads S(1) and S(2) need to be used simultaneously.
[0025] Define * as the composite operation between attack payloads, as shown in the following formula (3):
[0026] S(1)*S(2) = {x|x ∈ S(1) ∧ S(2)} (3)
[0027] Among them, S(1)*S(2) means using S(1) to process the attack payloads of S(2).
[0028] Optionally, using the improved TextCNN model to extract features from SQL injection statements in S2 includes:
[0029] S21. Segment the SQL injection statements, extract local context information based on a sliding window, and use the local context information as the input of the TextCNN model.
[0030] S22. According to the input, input layer, one-dimensional convolutional layer, and max pooling layer of the TextCNN model, extract features from the SQL injection statements.
[0031] Among them, the input layer adopts a two-channel form.
[0032] Optionally, the improved LSTM model in S3 includes: a dual-channel feature extraction module, a dual-channel feature fusion module, an input gate, a forget gate, a cell state, an output gate, and a classification module.
[0033] Among them, the dual-channel feature extraction module includes: a character-level channel and a statement-level channel.
[0034] The character-level channel is used to extract the temporal features of the special character sequence features in the SQL injection statement.
[0035] The statement-level channel is used to model the context dependency of keywords in the SQL injection statement.
[0036] On the other hand, a SQL injection detection device based on SQL context semantic analysis is provided. This device is applied to the SQL injection detection method based on SQL context semantic analysis. The device includes:
[0037] An acquisition module, configured to acquire SQL injection attack samples and normal query samples, process the SQL injection attack samples using an improved augmented attack tree SQL injection attack model to generate SQL injection statements, and construct a SQL data set according to the SQL injection statements and normal query samples.
[0038] A feature extraction module, configured to extract features from the SQL injection statements using an improved TextCNN model.
[0039] A training module, configured to train and optimize the improved LSTM model according to the SQL data set and the extracted features to obtain an injection detection model.
[0040] An output module, configured to acquire the input data to be detected, and detect the input data according to the injection detection model to obtain a SQL injection detection result.
[0041] Optionally, the acquisition module is further configured to:
[0042] S11. Formally describe the attack path from the attack starting point to the target achievement using an attack tree model.
[0043] S12. Define the features of the SQL injection attack using a formal language for the formally described attack path to generate SQL injection statements.
[0044] Optionally, the acquisition module is further configured to:
[0045] Construct an attack payload set according to the features of the SQL injection attack, define logical connectives to combine the attack payloads, and generate SQL injection statements.
[0046] Among them, logical connectives are defined, including: defining || as the OR operation of attack payloads, as shown in the following formula (1):
[0047] S(1)‖S(2) = {x|x∈S(1)∨x∈S(2)} (1)
[0048] Among them, S(1)‖S(2) means either one of the two attack payloads S(1) and S(2) is selected.
[0049] Defining && as the AND operation of attack payloads, as shown in the following formula (2):
[0050] S(1)&&S(2) = {x,y|x∈S(1)∧y∈S(2)} (2)
[0051] Among them, S(1)&&S(2) means that the two attack payloads S(1) and S(2) need to be used simultaneously.
[0052] Defining * as the composite operation between attack payloads, as shown in the following formula (3):
[0053] S(1)*S(2) = {x|x∈S(1)∧S(2)} (3)
[0054] Among them, S(1)*S(2) means using S(1) to process the attack payload of S(2).
[0055] Optionally, the feature extraction module is further used for:
[0056] S21. Tokenize the SQL injection statement, extract local context information based on a sliding window, and use the local context information as the input of the TextCNN model.
[0057] S22. Extract features from the SQL injection statement according to the input, input layer, one-dimensional convolutional layer, and max pooling layer of the TextCNN model.
[0058] Among them, the input layer adopts a two-channel form.
[0059] Optionally, the improved LSTM model includes: a two-channel feature extraction module, a two-channel feature fusion module, an input gate, a forget gate, a cell state, an output gate, and a classification module.
[0060] Among them, the two-channel feature extraction module includes: a character-level channel and a statement-level channel.
[0061] The character-level channel is used to extract the temporal features of the special character sequence features in the SQL injection statement.
[0062] The statement-level channel is used to model the context dependency relationship of keywords in the SQL injection statement.
[0063] On the other hand, a SQL injection detection device is provided, which includes: a processor; a memory storing computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the SQL injection detection methods based on SQL context semantic analysis as described above is implemented.
[0064] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the SQL injection detection methods based on SQL context semantic analysis as described above.
[0065] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0066] In the present invention, the collection of the SQL dataset is the basis for SQL injection detection. However, traditional data collection methods often have the problem of incomplete coverage and are difficult to include all types of SQL injection attacks. The method of generating SQL injection statements using a specific algorithm can systematically cover various attack types, including numeric, character, and search-type injections, as well as injection attacks submitted in different ways, such as GET, POST, Cookie, and HTTP request header injections. The advanced feature of this method is that it can simulate more comprehensive and diverse attack scenarios, thus providing a higher-quality dataset for subsequent model training.
[0067] Improve the application of the TextCNN model in SQL injection detection. Compared with traditional methods, the improved TextCNN introduces an attention mechanism on the basis of the original TextCNN, enabling the model to better capture the semantic dependency relationships and potential patterns in SQL statements. Specifically, the SQL statement is tokenized before the embedding layer, and local context information is extracted based on a sliding window and used as the input of the TextCNN. This method can not only retain the overall semantics of the statement but also enhance the model's ability to identify SQL injection features, especially when dealing with complex injection patterns. Compared with the traditional TextCNN method that only relies on local pattern detection, the innovation lies in introducing the attention mechanism, combining deep context information with convolutional features, making the model more sensitive to the hidden patterns of SQL injection.
[0068] Compared with introducing a dual-channel feature extraction architecture from traditional methods, the improved LSTM captures abnormal features at two levels: the character level and the statement level. The character-level channel uses LSTM to extract the temporal features of special character sequences (such as --, / *...* / , 'OR 1=1) in SQL statements, effectively dealing with non-standard syntax and hidden character injection patterns; the statement-level channel models the context dependencies of keywords (such as SELECT, UNION, WHERE) in SQL statements to parse complex semantic and structural patterns. After feature extraction, the attention mechanism is combined to fuse the dual-channel features, focusing on high-risk areas (such as string concatenation and conditional statements), and the fused features are input into the classification layer for discrimination. Incorporating prior knowledge in the field such as SQL keyword marking and abnormal character frequencies into the feature layer and working together with deep features can handle complex and variant attack scenarios. Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0070] Figure 1 is a flowchart of a SQL injection detection method based on SQL context semantic analysis provided by an embodiment of the present invention;
[0071] Figure 2 is a flowchart of dataset construction provided by an embodiment of the present invention;
[0072] Figure 3 is a diagram of a SQL injection attack model of an augmented attack tree provided by an embodiment of the present invention;
[0073] Figure 4 is a flowchart of text convolution feature extraction provided by an embodiment of the present invention;
[0074] Figure 5 is a flowchart of identifying sequence anomalies based on a long short-term memory network provided by an embodiment of the present invention;
[0075] Figure 6 is a block diagram of a SQL injection detection device based on SQL context semantic analysis provided by an embodiment of the present invention;
[0076] Figure 7 is a schematic structural diagram of a SQL injection detection device provided by an embodiment of the present invention. Detailed Embodiments
[0077] The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0078] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0079] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0080] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0081] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0082] The embodiments of the present invention provide a SQL injection detection method based on SQL context semantic analysis. This method can be implemented by a SQL injection detection device, and the SQL injection detection device can be a terminal or a server. As Figure 1 shown in the flowchart of the SQL injection detection method based on SQL context semantic analysis, the processing flow of this method can include the following steps:
[0083] S1. Obtain SQL injection attack samples and normal query samples, use an improved augmented attack tree SQL injection attack model to process the SQL injection attack samples, generate SQL injection statements, and construct a SQL data set according to the SQL injection statements and normal query samples.
[0084] In a feasible implementation, as Figure 2 shown, collect network SQL injection statements. To prevent disadvantages such as incompleteness and failure to cover all types of SQL injection attacks, use a specific algorithm to generate SQL injection statements.
[0085] Specifically, SQL injection attack examples are obtained through the open-source dataset SQLiV3. There are 24,500 normal samples and 25,527 attack samples in the training set; 10,000 normal samples and 10,000 injection attack samples in the validation set; 4,000 normal samples and 4,000 injection attack samples in the test set.
[0086] Optionally, in the SQL injection attack model using the improved augmented attack tree in S1, the SQL injection attack samples are processed to generate SQL injection statements, which may include the following steps S11 - S12:
[0087] S11. Use the attack tree model to formally describe the attack path from the attack starting point to the goal achievement.
[0088] S12. For the formally described attack path, use a formal language to define the characteristics of the SQL injection attack and generate SQL injection statements.
[0089] In a feasible implementation, the SQL injection statements are modeled and analyzed to obtain a formal description of the attack statements, and then the final samples are generated through instantiation. The model used in the present invention is an improved type of the SQL injection attack model of the augmented attack tree as Figure 3 shown.
[0090] Specifically, first, the attack tree model is used to systematically describe and analyze potential attack paths. The attack tree model formally describes all possible paths from the attack starting point to the goal achievement through a tree structure, which helps to comprehensively understand the strategies and steps that an attacker may take. In the research, a formal language can be used, that is, in the attack language of SQL injection, when what reflection characteristics appear can it be determined that it has a SQL injection security vulnerability. The instantiated data will provide guidance for the training of the SQL injection detection machine learning model.
[0091] Next, the present invention uses a formal language to precisely define the characteristics of the SQL injection attack, including various attack payload sets, and uses logical connectives to combine these payloads to form a complete attack statement. This step is crucial because it ensures that the diversity and complexity of the SQL injection attack can be captured. By instantiating these definitions and combinations, a large number of SQL injection samples can be generated, and these samples will be used to train the machine learning model to enable it to identify and defend against SQL injection attacks.
[0092] The test cases for SQL injection are composed of some basic attack payloads combined in a certain rule. Therefore, a formal description of the attack payloads will be given here. Among them, S(X) represents a set of a certain type of attack input. In particular, when there are multiple examples in the example, they are separated by the symbol " / ". After defining the attack payloads, in order to combine the attack payloads into attack inputs according to a certain rule, it is also necessary to define the operators between the attack payloads to describe the rules and relationships between the attack payloads when implementing a certain type of SQL injection attack.
[0093] Furthermore, define the following operators: Define || as the OR operation of the attack payloads, as shown in the following formula (1):
[0094] S(1)‖S(2)={x|x∈S(1)∨x∈S(2)} (1)
[0095] Among them, S(1)‖S(2) means that either S(1) or S(2) of the two attack payloads can be selected.
[0096] Define && as the AND operation of the attack payloads, as shown in the following formula (2):
[0097] S(1)&&S(2)={x,y|x∈S(1)∧y∈S(2)} (2)
[0098] Among them, S(1)&&S(2) means that the two attack payloads of S(1) and S(2) need to be used simultaneously.
[0099] Define * as the composite operation between the attack payloads, as shown in the following formula (3):
[0100] S(1)*S(2)={x|x∈S(1)∧S(2)} (3)
[0101] Among them, S(1)*S(2) means to process the attack payload of S(2) with S(1) to generate a new or composite attack payload. The operation order is from right to left, that is, S(1)*S(2)*S(3)=S(1)*(S(2)*S(3)). Next, it is also necessary to define the priority of the budget. Among various operators, the priority of parentheses is the highest, and the operation level of * is higher than that of || and &&.
[0102] To ensure the accuracy and robustness of the model, when constructing the dataset in the present invention, it not only includes a wide range of attack types but also covers normal query statements, so that the model can distinguish between normal and malicious SQL statements. This balanced dataset design helps improve the generalization ability of the model, enabling it to maintain a high recognition rate even when facing unknown attacks. In addition, the present invention pays attention to the diversity of the dataset, including SQL injection attack samples under different database systems and different programming language environments, to ensure that the model can perform well in various actual application scenarios. By such a method, the present invention can construct a comprehensive, accurate and highly practical SQL injection detection dataset, providing a solid foundation for the training and deployment of machine learning models.
[0103] In the embodiment of the present invention, the collection of the SQL dataset is the basis of SQL injection detection, but the traditional data collection methods often have the problem of incomplete coverage and are difficult to include all types of SQL injection attacks. The method of generating SQL injection statements using a specific algorithm can systematically cover various attack types, including numeric, character, and search-type injections, as well as injection attacks submitted in different ways, such as GET, POST, Cookie, and HTTP request header injections. The advanced nature of this method lies in its ability to simulate more comprehensive and diverse attack scenarios, thus providing a higher-quality dataset for subsequent model training.
[0104] S2. Use the improved TextCNN model to extract features from SQL injection statements.
[0105] In a feasible implementation manner, as Figure 4 shown, use the improved TextCNN (Text Convolutional Neural Network, a convolutional neural network for processing text data) model to extract features from SQL statements. The improved TextCNN model mainly uses a one-dimensional convolutional layer and a max-pooling layer. By capturing local features, combining and screening the features, semantic information at different abstraction levels can be obtained.
[0106] Optionally, the above step S2 may include the following steps S21 - S22:
[0107] S21. Segment the SQL injection statements and extract local context information based on a sliding window, and use the local context information as the input of the TextCNN model.
[0108] In a feasible implementation, pre - input introduces an attention mechanism, enabling the model to better capture semantic dependency relationships and potential patterns in SQL statements. Specifically, the SQL statement is tokenized before the embedding layer, and local context information is extracted based on a sliding window and used as the input to TextCNN.
[0109] The improved TextCNN of the present invention introduces an attention mechanism on the basis of the original TextCNN, enabling the model to better capture semantic dependency relationships and potential patterns in SQL statements. Specifically, the SQL statement is tokenized before the embedding layer, and local context information is extracted based on a sliding window and used as the input to TextCNN. This method can not only retain the overall semantics of the statement but also enhance the model's ability to identify SQL injection features, especially when dealing with complex injection patterns. Compared with the traditional TextCNN method that only relies on local pattern detection, the innovation lies in introducing an attention mechanism to combine deep context information with convolutional features, making the model more sensitive to the hidden patterns of SQL injection.
[0110] S22. Extract features from the SQL injection statement according to the input, input layer, one - dimensional convolutional layer, and max - pooling layer of the TextCNN model.
[0111] Among them, the input layer adopts a two - channel form.
[0112] In a feasible implementation, the first layer is the input layer. The input layer is an n*k matrix, where n is the number of words in a sentence and k is the dimension of the word vector corresponding to each word. That is to say, each row of the input layer is a k - dimensional word vector corresponding to a word. Additionally, here a padding operation is performed on the original sentence to make the vector lengths consistent. The input layer of the present invention adopts a two - channel form, that is, there are two n*k input matrices. One is expressed by pre - trained word embeddings and does not change during the training process; the other is also initialized in the same way but will be used as a parameter and change during the training process of the network.
[0113] The text of the input layer is a word matrix composed of word vectors, and the width of the convolutional kernel is the same as the width of the word matrix, and this width is the word vector size, and the convolutional kernel only moves in the height direction. Therefore, the position where the convolutional kernel slides through each time is a complete word, and it will not perform convolution on a part of the text of several words. The rows of the word matrix represent discrete symbols (i.e., words), which ensures the rationality of word as the smallest granularity in language.
[0114] On the matrix of n*k as the input, use a kernelw∈R hk and a window (window)x i:i+h-1Perform a convolution operation to generate a feature C i , that is: C i = f(w * x i:i+h-1 + b), where x i:i+h-1 represents a window of size h * k composed of the i-th row to the (i + h - 1)-th row of the input matrix, and is composed of x i , x i+1 , …, x i:i+h-1 concatenated. h represents the number of words in the window, w is a weight matrix of dimension h * k (so the number of parameters to be learned by a filter is h * k), b is a bias parameter, and f is a non-linear function. w and x i:i+h-1 perform a dot product operation (dot product, the result of element-wise multiplication is then summed). The filter is applied to a sentence, moving one step downwards at a time (i = 1…n - h + 1). For example, on h, after the convolution operation, C 1 is obtained, and on x 2i+h-1 , after the convolution operation, C 2 …, and then the concatenated c = [C 1 , c 2 , …, C n-h+1 is the feature map of the present invention. Each convolution operation is equivalent to an extraction of a feature vector. By defining different windows, different feature vectors can be extracted to form the output of the convolutional layer.
[0115] Finally, there is the pooling layer. The network uses 1-Max pooling, that is, selecting the largest feature from the feature vectors generated by each sliding window, and then concatenating these features to form a vector representation. K-Max pooling (selecting the largest K features in each feature vector) or average pooling (taking the average of each dimension in the feature vector) can also be selected. The effect achieved is to obtain a fixed-length vector representation for sentences of different lengths through pooling.
[0116] The improved TextCNN model is a powerful text classification tool that extracts features from SQL statements through one-dimensional convolutional layers and max-pooling layers. The core advantage of this model lies in its ability to capture local features, similar to finding n-gram patterns in text, that is, by using multiple convolutional kernels of different sizes to extract key information in sentences, thereby better capturing local correlations. When processing SQL statements, this local feature extraction is particularly important because the structure and keywords of SQL statements are crucial for understanding their intent and function. In the improved TextCNN model, the input is usually text data processed by word embeddings, forming a three-dimensional matrix that contains the sequence length of the text and the dimension of the word vectors. The convolutional layer performs convolutional operations on the input matrix using convolutional kernels of different sizes, extracting text features of different lengths, which can be regarded as local patterns or structures in the text. Subsequently, the max-pooling layer reduces the dimension of the feature maps output by the convolutional layer, extracting the most important features. This step helps the model focus on the most significant signals, thereby improving the accuracy of classification. Through this way of feature extraction and combination, the improved TextCNN model can obtain semantic information at different abstraction levels, which is very helpful for understanding and classifying SQL statements. For example, it can help distinguish query statements, update statements, insert statements, or delete statements, and even identify SQL injection attack attempts because these malicious behaviors often contain specific patterns and keywords. In this way, the improved TextCNN model can not only extract the features of SQL statements but also provide rich information for further analysis and processing.
[0117] In the embodiments of the present invention, the application of the improved TextCNN model in SQL injection detection, compared with traditional methods, the improved TextCNN introduces an attention mechanism on the basis of the original TextCNN, enabling the model to better capture semantic dependency relationships and potential patterns in SQL statements. Specifically, before the embedding layer, the SQL statement is tokenized, and local context information is extracted based on a sliding window and used as the input of TextCNN. This method can not only retain the overall semantics of the statement but also enhance the model's ability to identify SQL injection features, especially when dealing with complex injection patterns. Compared with the traditional TextCNN method that only relies on local pattern detection, the innovation lies in introducing an attention mechanism to combine deep context information with convolutional features, making the model more sensitive to hidden patterns of SQL injection.
[0118] S3. According to the SQL dataset and the extracted features, train and optimize the improved LSTM model to obtain an injection detection model.
[0119] In a feasible implementation, such as Figure 5As shown, by improving the LSTM (Long Short-Term Memory) model to train and optimize the extracted features, the improved LSTM model can effectively process sequence data and capture long-distance dependencies.
[0120] Optionally, the improved LSTM model in S3 includes: a dual-channel feature extraction module, a dual-channel feature fusion module, an input gate, a forget gate, a cell state, an output gate, and a classification module.
[0121] Among them, the dual-channel feature extraction module includes: a character-level channel and a statement-level channel.
[0122] The character-level channel is used to extract the temporal features of the special character sequence features in the SQL injection statement.
[0123] The statement-level channel is used to model the context dependency of the keywords in the SQL injection statement.
[0124] In a feasible implementation, in order to more comprehensively capture SQL anomaly features, the improved LSTM introduces a dual-channel feature extraction architecture to capture anomaly features from two levels: the character level and the statement level. The character-level channel uses LSTM to extract the temporal features of special character sequence features (such as --, / *...* / , 'OR 1=1) in the SQL statement, effectively dealing with non-standard syntax and hidden character injection patterns; the statement-level channel models the context dependency of keywords (such as SELECT, UNION, WHERE) in the SQL statement to parse complex semantic and structural patterns. After feature extraction, the attention mechanism is combined to fuse the dual-channel features, focusing on high-risk areas (such as string concatenation and conditional statements), and the fused features are input into the classification layer for discrimination. The innovation of the present invention is mainly reflected in integrating prior knowledge in fields such as SQL keyword marking and abnormal character frequency into the feature layer, working together with the deep features to enhance the robustness and interpretability of the model. This design enables the model to not only efficiently identify traditional SQL injection patterns but also cope with complex and variant attack scenarios, having significant practical application value.
[0125] Furthermore, the improved LSTM model is further trained and optimized for the abnormal features extracted from SQL statements. The improved LSTM model is a special type of Recurrent Neural Network (RNN) designed to address the vanishing gradient or exploding gradient problems encountered by traditional RNNs when dealing with long sequence data. By introducing gating mechanisms such as the input gate, forget gate, and output gate, the improved LSTM can selectively retain or forget information, effectively process sequence data, and capture long-range dependencies. In the context of SQL anomaly detection, the improved LSTM model can analyze the execution sequence of SQL statements and identify the subtle differences between normal operations and potential abnormal behaviors. Since SQL statements usually contain multiple steps and operations with complex dependencies between them, the long-range dependency capture ability of the improved LSTM model makes it an ideal choice for analyzing such data. Through training, the improved LSTM model can learn the patterns of normal SQL statements and identify abnormal features that deviate from these patterns, which may indicate potential security threats or errors. Additionally, the improved LSTM model continuously optimizes its parameters during training to improve the accuracy of abnormal feature recognition. This optimization not only involves the model's ability to process current data but also its generalization ability for unseen future data. In this way, the present invention can construct a powerful anomaly detection system that can monitor the execution of SQL statements in real time in practical applications, detect and respond to abnormal behaviors in a timely manner, thereby improving the security and stability of the database system.
[0126] Specifically, the long short-term memory network unit consists of four main parts: the input gate, forget gate, cell state, and output gate. Each part has its specific function and formula. ① Forget Gate: Determines which information should be forgotten from the cell state. ② Input Gate: Determines which new information will be stored in the cell state. ③ Cell State: Carries information about the observed input sequence. ④ Output Gate: Determines the output value based on the calculation of the cell state.
[0127] Suppose there is a time series data x t , where t is the time step. The improved LSTM cell calculates the following at each time step t:
[0128] The forget gate determines which information should be forgotten from the cell state. It uses the sigmoid activation function to output a value between 0 and 1, indicating the degree of retention or forgetting.
[0129] f t= σ(W f · [h t -1, x t + b f ) (4)
[0130] where f t is the output of the forget gate at time step t. σ is the sigmoid function that restricts the output range. W f is the weight matrix of the forget gate. h t -1 is the hidden state of the previous time step. x t is the input at the current time step. b f is the bias term of the forget gate.
[0131] The input gate consists of two parts: input modulation and candidate values for new cells. Input modulation determines the degree of writing new information, while candidate values for new cells are the candidate representations of new information.
[0132] i t = σ(W i · [h t -1, x t + b i ) (5)
[0133] C t = tanh(W C · [h t -1, x t + b C ) (6)
[0134] where i t is the output of the input gate at time step t. tanh is the hyperbolic tangent function. W i and W C are the weight matrices of the input gate and candidate values for new cells respectively. b i and b C are the bias terms of the input gate and candidate values for new cells respectively.
[0135] The cell state C t is the cell state at the current time step, which combines the cell state C t of the previous time step and the new information C t-1 .
[0136] C t = f t × C t-1 + i t × C t (7)
[0137] The output gate determines the calculation of the output value based on the cell state. It uses the sigmoid function to output a value between 0 and 1, indicating the degree of output.
[0138] o t = σ(W o · [h t -1,x t + b o ) (8)
[0139] h t = o t × tanh(C t ) (9)
[0140] Where: o t is the output of the forget gate at time step t. σ is the sigmoid function that restricts the output range. W o is the weight matrix of the forget gate. b o is the bias term of the forget gate. h t is the hidden state at time step t.
[0141] Furthermore, the present invention uses standard detection metrics F1, Precision, Recall. In classification tasks, Precision, Recall, and F1-score are the core metrics for evaluating model performance, each revealing the classification ability of the model from different perspectives. Precision reflects the accuracy of the model in predicting positive class samples and measures the effectiveness of the model in reducing false positives (false positive examples), so it is particularly crucial in scenarios that require high precision. Recall evaluates the model's ability to identify all actual positive class samples and reveals its effectiveness in avoiding false negatives (false negative examples), which is of great significance for tasks that require high coverage. F1-score, as the harmonic mean of Precision and Recall, balances these two metrics and provides a robust evaluation of the model's overall performance in the case of class imbalance. Generally speaking, these three metrics not only demonstrate the performance of the model in terms of accuracy and comprehensiveness, but also reflect its ability to achieve accurate identification and comprehensive coverage among different classes, which has profound guiding significance for the performance of classification models in practical applications.
[0142]
[0143] Where all detctions represents the number of all predictions that are unsafe, and all ground truths represents the number of all actual unsafe ones. Where TP, FP, FN:
[0144] True Positive (TP): IOU > IOU threshold (IOU threshold (usually taken as 0.5) the number of prediction results (each Ground Truth is only counted once).
[0145] False Positive(FP): The number of prediction results where IOU <= IOU threshold or the number of redundant detection boxes for the same GT detected.
[0146] False Negative(FN): The number of GTs not detected.
[0147] In the embodiments of the present invention, compared with introducing a dual-channel feature extraction architecture from traditional methods, the improved LSTM captures abnormal features at two levels: the character level and the statement level. The character-level channel uses LSTM to extract the temporal features of special character sequences (such as --, / *...* / , 'OR 1=1) in SQL statements, effectively dealing with non-standard syntax and hidden character injection patterns; the statement-level channel models the context dependencies of keywords (such as SELECT, UNION, WHERE) in SQL statements to parse complex semantic and structural patterns. After feature extraction, the attention mechanism is combined to fuse the dual-channel features, focusing on high-risk areas (such as string concatenation and conditional statements), and the fused features are input into the classification layer for discrimination. Incorporating prior knowledge in the field such as SQL keyword tagging and abnormal character frequencies into the feature layer and acting together with the deep features can handle complex and variant attack scenarios.
[0148] S4. Obtain the input data to be detected, and detect the input data according to the injection detection model to obtain the SQL injection detection result.
[0149] The advancement of deep learning models in SQL injection detection also faces challenges, including the need for a large amount of sample data, insufficient interpretability of the models, and imbalance of different categories of samples. To solve these problems, the present invention explores the combination of semantic feature extraction and ensemble learning methods to improve the generalization ability and anti-attack performance of the models. The present invention not only focuses on the statistical features of vocabulary but also on the context semantic of vocabulary, thus achieving better results than traditional detection methods in practice.
[0150] In the embodiments of the present invention, the collection of the SQL dataset is the basis for SQL injection detection, but traditional data collection methods often have the problem of incomplete coverage and are difficult to include all types of SQL injection attacks. The method of generating SQL injection statements using a specific algorithm can systematically cover various attack types, including numeric, character, and search-type injections, as well as injection attacks submitted in different ways, such as GET, POST, Cookie, and HTTP request header injections. The advantage of this method is that it can simulate more comprehensive and diverse attack scenarios, thus providing a higher-quality dataset for subsequent model training.
[0151] Improve the application of the TextCNN model in SQL injection detection. Compared with traditional methods, the improved TextCNN introduces an attention mechanism on the basis of the original TextCNN, enabling the model to better capture semantic dependency relationships and potential patterns in SQL statements. Specifically, the SQL statement is tokenized before the embedding layer, and local context information is extracted based on a sliding window and used as the input to TextCNN. This method can not only preserve the overall semantics of the statement but also enhance the model's ability to identify SQL injection features, especially when dealing with complex injection patterns. Compared with the traditional TextCNN method that only relies on local pattern detection, the innovation lies in introducing an attention mechanism to combine deep context information with convolutional features, making the model more sensitive to hidden patterns of SQL injection.
[0152] The improved LSTM introduces a dual-channel feature extraction architecture compared with traditional methods, capturing abnormal features from two levels: the character level and the statement level. The character-level channel uses LSTM to extract the temporal features of special character sequences in the SQL statement (such as --, / *...* / , 'OR 1=1), effectively dealing with non-standard syntax and hidden character injection patterns; the statement-level channel models the context dependency relationships of keywords in the SQL statement (such as SELECT, UNION, WHERE) through LSTM to parse complex semantic and structural patterns. After feature extraction, the attention mechanism is combined to fuse the dual-channel features, focusing on high-risk areas (such as string concatenation and conditional statements), and the fused features are input into the classification layer for discrimination. Incorporating prior knowledge in the field such as SQL keyword marking and abnormal character frequency into the feature layer and working together with deep features can handle complex and variant attack scenarios.
[0153] Figure 6 It is a block diagram of an SQL injection detection device based on SQL context semantic analysis shown according to an exemplary embodiment. This device is used for the SQL injection detection method based on SQL context semantic analysis. Refer to Figure 6 This device includes an acquisition module 310, a feature extraction module 320, a training module 330, and an output module 340. Among them:
[0154] The acquisition module 310 is used to acquire SQL injection attack samples and normal query samples, process the SQL injection attack samples using an improved augmented attack tree SQL injection attack model to generate SQL injection statements, and construct an SQL data set based on the SQL injection statements and normal query samples.
[0155] The feature extraction module 320 is used to extract features from the SQL injection statements using the improved TextCNN model.
[0156] A training module 330, configured to train and optimize an improved LSTM model according to an SQL data set and the extracted features, so as to obtain an injection detection model.
[0157] An output module 340, configured to obtain input data to be detected, and detect the input data according to the injection detection model, so as to obtain an SQL injection detection result.
[0158] In the embodiment of the present invention, the collection of the SQL data set is the basis of SQL injection detection. However, traditional data collection methods often have the problem of incomplete coverage and are difficult to cover all types of SQL injection attacks. The method of generating SQL injection statements by using a specific algorithm can systematically cover various attack types, including numeric, character, and search injection, as well as injection attacks submitted in different ways, such as GET, POST, Cookie, and HTTP request header injection. The advantage of this method is that it can simulate more comprehensive and diverse attack scenarios, thereby providing a higher-quality data set for subsequent model training.
[0159] Regarding the application of the improved TextCNN model in SQL injection detection, compared with traditional methods, the improved TextCNN introduces an attention mechanism on the basis of the original TextCNN, enabling the model to better capture semantic dependency relationships and potential patterns in SQL statements. Specifically, the SQL statement is tokenized before the embedding layer, and local context information is extracted based on a sliding window and used as the input of the TextCNN. This method can not only retain the overall semantics of the statement but also enhance the model's ability to recognize SQL injection features, especially when dealing with complex injection patterns. Compared with the traditional TextCNN method that only relies on local pattern detection, the innovation lies in introducing an attention mechanism to combine deep context information with convolutional features, making the model more sensitive to hidden patterns of SQL injection.
[0160] Compared with traditional methods, the improved LSTM introduces a dual-channel feature extraction architecture to capture abnormal features at the character level and statement level. The character-level channel uses LSTM to extract the temporal features of special character sequences (such as --, / *...* / , 'OR 1=1) in the SQL statement, effectively dealing with non-standard syntax and hidden character injection patterns; the statement-level channel models the context dependency relationships of keywords (such as SELECT, UNION, WHERE) in the SQL statement to parse complex semantic and structural patterns. After feature extraction, the attention mechanism is used to fuse the dual-channel features, focusing on high-risk areas (such as string concatenation and conditional statements), and the fused features are input into the classification layer for discrimination. Incorporating domain prior knowledge such as SQL keyword marking and abnormal character frequencies into the feature layer and jointly acting with the deep features can handle complex and variant attack scenarios.
[0161] Figure 7 This is a schematic structural diagram of an SQL injection detection device provided by an embodiment of the present invention. As shown in Figure 7 the figure, the SQL injection detection device may include the above-mentioned Figure 6 SQL injection detection device based on SQL context semantic analysis shown in the figure. Optionally, the SQL injection detection device 410 may include a first processor 2001.
[0162] Optionally, the SQL injection detection device 410 may further include a memory 2002 and a transceiver 2003.
[0163] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.
[0164] Next, in combination with Figure 7 each component of the SQL injection detection device 410 will be specifically introduced:
[0165] Among them, the first processor 2001 is the control center of the SQL injection detection device 410, which may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0166] Optionally, the first processor 2001 may execute various functions of the SQL injection detection device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0167] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 7 CPU0 and CPU1 shown in the figure.
[0168] In a specific implementation, as an embodiment, the SQL injection detection device 410 may also include multiple processors, such as Figure 7The first processor 2001 and the second processor 2004 shown in []. Each of these processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0169] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0170] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the SQL injection detection device 410 ( Figure 7 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0171] The transceiver 2003 is used to communicate with network devices or terminal devices.
[0172] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 7 not shown separately in []). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0173] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the SQL injection detection device 410 ( Figure 7 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0174] It should be noted that Figure 7 The structure of the SQL injection detection device 410 shown in Figure 7 does not limit the router. The actual knowledge structure recognition device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0175] In addition, the technical effects of the SQL injection detection device 410 can refer to the technical effects of the SQL injection detection method based on SQL context semantic analysis described in the above method embodiments, and will not be elaborated here.
[0176] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0177] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0178] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0179] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0180] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0181] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0182] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0183] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0184] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0185] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0186] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0187] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0188] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A SQL injection detection method based on SQL context semantic analysis, characterized in that: The method comprises: S1. Obtain SQL injection attack samples and normal query samples, use an improved SQL injection attack model of an augmented attack tree to process the SQL injection attack samples, generate SQL injection statements, and construct an SQL data set according to the SQL injection statements and normal query samples; S2. Using the improved TextCNN model, feature extraction is performed on the SQL injection statement; S3. According to the SQL data set and the extracted features, the improved LSTM model is trained and optimized to obtain an injection detection model; S4. Obtain input data to be detected, detect the input data according to the injection detection model, and obtain a SQL injection detection result.
2. The SQL injection detection method based on SQL context semantic analysis according to claim 1 is characterized in that: The improved augmented attack tree SQL injection attack model in S1 processes the SQL injection attack sample to generate an SQL injection statement, including: S11. Use the attack tree model to formally describe the attack path from the attack starting point to the target achievement; S12. For the formally described attack path, use a formal language to define the characteristics of the SQL injection attack and generate SQL injection statements.
3. The SQL injection detection method based on SQL context semantic analysis according to claim 2 is characterized in that: The step S12 uses a formal language to define the features of the SQL injection attack and generates an SQL injection statement, including: According to the characteristics of SQL injection attacks, an attack payload set is constructed, logical connectives are defined to combine the attack payloads, and SQL injection statements are generated; The defining of the logical connective includes: defining || as an OR operation of the attack payload, as shown in the following formula (1): S(1)‖S(2)={x|x∈S(1)∨x∈S(2)} (1) Where S(1)‖S(2) represents any one of the two attack payloads S(1) and S(2); Define && as the AND operation of the attack payload, as shown in the following formula (2): S(1)&&S(2)={x,y|x∈S(1)∧y∈S(2)} (2) Where S(1)&&S(2) means that both attack payloads S(1) and S(2) need to be used simultaneously; Define * as the composite operation between attack payloads, as shown in the following formula (3): S(1)*S(2)={x|x∈S(1)∧S(2)} (3) Here, S(1)*S(2) means using S(1) to process the attack payload of S(2).
4. The SQL injection detection method based on SQL context semantic analysis according to claim 1, characterized in that: The improved TextCNN model is used in S2 to extract features of the SQL injection statement, including: S21, segmenting the SQL injection statement, extracting local context information based on a sliding window, and using the local context information as input of a TextCNN model; S22, performing feature extraction on the SQL injection statement according to the input, input layer, one-dimensional convolution layer and maximum pooling layer of the TextCNN model; Wherein, the input layer adopts a dual-channel form.
5. The SQL injection detection method based on SQL context semantic analysis according to claim 1 is characterized in that: The improved LSTM model in S3 includes: a dual-channel feature extraction module, a dual-channel feature fusion module, an input gate, a forget gate, a unit state, an output gate and a classification module; Wherein, the dual-channel feature extraction module includes: a character-level channel and a sentence-level channel; The character-level channel is used to extract the temporal features of special character sequence features in the SQL injection statement; The statement-level channel is used to model the contextual dependencies of keywords in SQL injection statements.
6. A SQL injection detection device based on SQL context semantic analysis, the SQL injection detection device based on SQL context semantic analysis is used to implement the SQL injection detection method based on SQL context semantic analysis as claimed in any one of claims 1 to 5, characterized in that: The device comprises: An acquisition module is used to acquire SQL injection attack samples and normal query samples, process the SQL injection attack samples using an improved augmented attack tree SQL injection attack model, generate SQL injection statements, and construct an SQL data set according to the SQL injection statements and normal query samples; A feature extraction module, used to extract features of the SQL injection statement using an improved TextCNN model; A training module, used to train and optimize the improved LSTM model according to the SQL data set and the extracted features to obtain an injection detection model; The output module is used to obtain input data to be detected, detect the input data according to the injection detection model, and obtain SQL injection detection results.
7. The SQL injection detection device based on SQL context semantic analysis according to claim 6, characterized in that: The improved SQL injection attack model using the augmented attack tree is used to process the SQL injection attack sample and generate an SQL injection statement, including: S11. Use the attack tree model to formally describe the attack path from the attack starting point to the target achievement; S12. For the formally described attack path, use a formal language to define the characteristics of the SQL injection attack and generate SQL injection statements.
8. The SQL injection detection device based on SQL context semantic analysis according to claim 7, characterized in that: The method of using a formal language to define the characteristics of an SQL injection attack and generating an SQL injection statement includes: According to the characteristics of SQL injection attacks, an attack payload set is constructed, logical connectives are defined to combine the attack payloads, and SQL injection statements are generated; The defining of the logical connective includes: defining || as an OR operation of the attack payload, as shown in the following formula (1): S(1)‖S(2)={x|x∈S(1)∨x∈S(2)} (1) Among them, S(1)‖S(2) means that either of the two attack payloads S(1) and S(2) is selected; Define && as the AND operation of the attack payload, as shown in the following formula (2): S(1)&&S(2)={x,y|x∈S(1)∧y∈S(2)} (2) Among them, S(1)&&S(2) means that both S(1) and S(2) attack payloads need to be used at the same time; Define * as the composite operation between attack payloads, as shown in the following formula (3): S(1)*S(2)={x|x∈S(1)∧S(2)} (3) Here, S(1)*S(2) means using S(1) to process the attack payload of S(2).
9. A SQL injection detection device, characterized in that: The SQL injection detection device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 5.