SQL (Structured Query Language) injection feature extraction method and system and medium
By using custom feature groups and standardization, the problem of traditional feature extraction methods affecting the accuracy of machine learning models is solved, achieving more efficient SQL injection detection.
Patent Information
- Application Number
- CN202510869546.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, traditional feature extraction methods are affected by factors such as the web application's running server, operating environment, database, and input data format, leading to inaccurate learning by machine learning models and resulting in low accuracy in SQL injection detection.
A custom feature group is used, including basic scale features, syntax construction features, syntax interpretation features, complexity features, comprehensive description features, and character separation features. By initializing the original load feature vector, the feature group is defined and a feature matrix is constructed. After standardization, SQL injection features are extracted.
It improves the recognition accuracy of machine learning models, reduces false positive rates, enhances the model's ability to identify abnormal parts, and ensures accurate description of SQL injection characteristics in a low-dimensional context.
Smart Images

Figure CN120995061A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital information transmission technology, and in particular to a method, system and medium for extracting SQL injection features. Background Technology
[0002] With the increasing popularity of computers and the internet, and the rapid development of Web technologies, various Web applications have become deeply integrated into people's daily learning, work, and life. However, the convenience and openness of Web applications have also made their security issues increasingly prominent. Various forms of attacks targeting Web applications seriously threaten people's privacy, property, and other security. Among these, SQL injection attacks, characterized by rapid mutation, numerous types, and stealth, are one of the most important methods hackers use to attack Web applications and a major channel for Web information leakage. Therefore, SQL injection detection has become a key focus of current Web application security research. Depending on the implementation technology, SQL injection attack detection mainly includes pattern matching-based SQL injection detection and machine learning-based SQL injection detection techniques.
[0003] Early researchers typically relied on pattern matching techniques to detect SQL injection attacks. The basic principle of this technique is to represent known SQL injection patterns as specific rules and then match data that conforms to these rules in network traffic. For example, a legitimate SQL Master File (SQLMF) is generated from the web application, and the input SQL query is then dynamically matched against the SQLMF. However, dynamic matching incurs significant performance overhead when detecting a large number of SQL query requests in real time, thus affecting the response speed and throughput of SQL injection detection. To address this, some researchers have proposed a hybrid SQL injection detection technique combining static and dynamic analysis. While the static analysis phase matches strings, the dynamic analysis phase performs tracing to identify SQL injection attacks. Alternatively, the Aho-Corasick algorithm can be used to perform static pattern matching on user-generated SQL queries, storing the abnormal keyword set from the static pattern matches, and then updating the static pattern table with newly generated abnormal patterns through dynamic pattern matching. However, attackers can bypass the static phase detection by modifying the structure or format of the attack payload, making it impossible for the dynamic analysis phase to capture these mutated attacks. Other researchers have utilized syntax pattern matching to extract and customize web access log fields from web applications, designing syntax pattern recognizers to detect SQL injection attacks. Alternatively, an abstract syntax tree library of SQL injection feature patterns can be generated based on SQL syntax parsing and SQL semantic parsing, establishing an SQL injection feature library based on SQL syntax elements and fields. However, syntax pattern matching methods cannot detect SQL injection attacks that mix normal query statements with malicious statements. Due to the limitations of pattern matching technology, more and more researchers are using machine learning algorithms to identify SQL injection attack behavior in web applications, detecting SQL injection attacks by learning normal and abnormal behavior patterns from large amounts of data. This technology can detect unknown attacks and has a certain degree of adaptability. In recent years, machine learning has been widely recognized in the detection and prediction of SQL injection attacks, but most of these studies use text classification methods to extract features from SQL injection payloads. For example, using words such as "OR=", "AND=", "HAVING=", "WAITFOR DELAY", and "GROUP BY" from the built-in keyword set to convert SQL statements into feature vectors. The CBOW natural language processing model is used to train word vectors to extract SQL injection features. For words that appear less frequently in SQL injection, the CBOW model generates sparse word vectors, resulting in inaccurate representations of some keywords. Based on semantic features, the Word2vec algorithm is used to convert SQL data into word vectors in the payload, extracting SQL injection payload feature information. A TF-IDF-based SQL injection feature extraction method is also employed.The feature values of SQL statements in the statement whitelist and statement blacklist are calculated using the weighted TF-IDF algorithm, resulting in whitelist statement feature vectors and blacklist statement feature vectors. The N-GRAM method is used to convert samples into feature vectors. The vectors extracted by N-GRAM are optimized using the TF-IDF algorithm. In the preprocessing stage, N-GRAM technology is used to select feature words, and then TF-IDF technology is used to perform SQL statement text vectorization to extract SQL injection features. A BOW-based SQL injection feature extraction method is also proposed. The bag-of-words model is used to extract text features from the preprocessed SQL injection sample data. Then, the preprocessed SQL injection sample data is matched with a one-hot encoded general rule base to obtain the rule code for each sample as a rule feature. The text features and rule features are then fused to obtain a feature dataset. A Feature Ratio Method (FRM) SQL injection feature extraction method is also proposed. This method is not based on words, but rather extracts feature points of individual bit characters. By combining static and dynamic features of SQL statements and through reasonable feature selection and feature vector synthesis, features that can identify the self-generated structure and execution result of SQL statements are obtained. This paper utilizes an improved TextCNN model to extract features from SQL injection statements. Convolutional neural networks (CNNs) capture local patterns within SQL statements, improving the extraction of key information. Feature vectors for SQL statements are extracted using the collapse probability of wavefunction vector substate transitions. Initial features of the SQL query and execution plan are generated using a BERT model. The query is then input into a self-attention-based encoder to obtain semantic and SQL injection features. The execution plan is then input into a relation-aware attention-based encoder to obtain execution logic features. Finally, the two sets of features are fused using a cross-modal encoder to obtain the SQL feature representation. The SQL statements are segmented to obtain corresponding word vectors. These word vectors are then averaged to obtain the sentence vector for each SQL statement, initializing a category vector representing injection and non-injection. Next, the sentence word vectors are multiplied by the corresponding category representation vector to obtain the category representation word vector. Finally, the obtained category representation word vectors are fed into the TextCNN model to obtain the classification semantic vector. Finally, the text frequency index (TVI) and the super Term Vector algorithm are combined to transform the text data of injection attack statements in the dataset into corresponding numerical features. However, SQL injection features differ from ordinary text features. Existing feature extraction methods, such as word segmentation, suffer from problems such as inaccurate word segmentation due to segmentation errors, low precision, and excessively large word feature vector dimensions.Therefore, this invention proposes a new SQL injection feature extraction method and standardizes it according to the SQL injection feature dimension, so that the feature vector can correctly describe the features of SQL injection attacks in a low dimension, thereby reducing the precision loss and low accuracy caused by indistinct features and word segmentation errors. Summary of the Invention
[0004] This application provides a method, system, and medium for extracting SQL injection features. It addresses the problems of traditional feature extraction methods being affected by factors such as the web application's running server, operating environment, database, and input data format, leading to inaccurate machine learning model learning and low SQL injection detection accuracy. By extracting features from a custom feature group, it helps the machine learning model train and identify more accurately from different dimensions, ensuring both model performance and recognition capability.
[0005] This application provides a method for extracting SQL injection features, including:
[0006] Step S1: Initialize the original load feature vector by inputting and traversing the original load dataset;
[0007] Step S2: Read the raw payload and define the feature group as the SQL injection feature vector;
[0008] Step S3: Extract the basic scale features, grammatical construction features, grammatical interpretation features, complexity features, comprehensive description features, and character separation features from the feature group in sequence.
[0009] Step S4: Construct the feature matrix and output the standardized feature matrix.
[0010] Preferably, the
[0011] The basic scale characteristics are: query length and number of words;
[0012] The grammatical construction features are: the number of single quotes, the number of double quotes, the number of left parentheses, and the number of right parentheses;
[0013] The grammatical characteristics are: the number of special symbols and the number of annotation symbols;
[0014] The complexity characteristics are: the number of logical operators and the number of arithmetic operators;
[0015] The comprehensive description features include: the number of hexadecimal numbers, the number of letters, the number of numbers, the number of SQL keywords, and the number of SQL functions;
[0016] The character segmentation feature is the number of whitespace characters.
[0017] Preferably, the grammatical construction features also introduce a single quote parity flag, a double quote parity flag, a left bracket parity flag, and a right bracket parity flag.
[0018] Preferably, the feature matrix expression in step S4 is:
[0019]
[0020] Where: q i , i∈Z + For the original load; Φ ss () represents the original load (q) i The SQL injection feature extraction function set; τ is the function set for extracting Φ. ss The result of () is converted into an SQL injection feature matrix; m, n∈Z + For standardized calculations.
[0021] Preferably, the u m Let u be the sample mean of the m-th column. m The expression is:
[0022]
[0023] Where x mn This represents the value in the m-th column and n-th row of the matrix.
[0024] Preferably, the s m s is the sample standard deviation of the m-th column. m The expression is:
[0025]
[0026] This application also proposes an SQL injection feature extraction system to implement the above feature extraction method, including:
[0027] The feature vector input module is used to initialize the original load feature vector, input and traverse the original load dataset;
[0028] The feature group definition module is used to define feature groups as SQL injection feature vectors;
[0029] The feature extraction module is used to extract basic scale features, syntax construction features, syntax interpretation features, complexity features, comprehensive description features, and character separation features from the feature group in sequence.
[0030] The feature matrix output module is used to construct the feature matrix and output the standardized feature matrix.
[0031] This application also proposes a computer-readable storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the above-described feature extraction method.
[0032] The technical solutions provided in this application embodiment have the following technical effects or advantages:
[0033] 1. Due to the use of custom feature groups, which include six feature modules—basic scale features, syntax construction features, syntax interpretation features, complexity features, comprehensive description features, and character separation features—a total of 20 features are employed. These 20 features together form a complete SQL injection feature vector. These 20 features, defined above, help the machine learning model to train and identify SQL injection more accurately from different dimensions. Adding too many redundant features will lead to a decrease in model performance, while reducing relevant features will lead to a decrease in the model's recognition ability.
[0034] 2. This application proposes introducing parity features for single quotes, double quotes, left parentheses, and right parentheses into the syntax construction feature set. This solves the technical problem that existing detection methods would directly classify a user's incorrect input of a quote or parenthesis as SQL injection, leading to an increased false positive rate. By introducing parity features, the false positive rate is reduced.
[0035] 3. In the original payload, this application constructs a feature vector that fuses numerical and Boolean values. The numerical part mainly reflects the intensity or scale signal of the feature, while the Boolean part reflects whether it exists. By fusing these two types of feature values, the machine learning model can learn both the linear or nonlinear effects caused by the difference in magnitude and capture key content, thereby achieving information complementarity and improving the model's discriminative ability.
[0036] 4. This application standardizes the feature matrix, resolving issues such as large differences in dimensions and inconsistent distribution among SQL injection features. The original feature data is linearly transformed according to its sample mean and standard deviation to eliminate the influence of different feature scales and enhance the model's ability to identify outliers. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the feature extraction method in Embodiment 1 of this application;
[0038] Figure 2 This is a flowchart of the feature extraction method in Embodiment 1 of this application. Detailed Implementation
[0039] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0040] Example 1
[0041] This application proposes a method for extracting SQL injection features. In machine learning, the raw SQL payload cannot be directly used for model training; feature extraction is required to generate a feature vector, which can then be used for machine learning. The SQL injection features of the raw payload in this invention refer to the following: number of single quotes (sq), number of double quotes (dq), number of left brackets (lparen), number of right brackets (rparen), number of special characters (puncts), number of comments (comments), number of logical operators (logic), number of arithmetic operators (arith), number of hexadecimal numbers (hexnum), number of letters (alpha), number of numbers (digit), number of SQL keywords (sqlkw), number of SQL functions (sqlfunc), raw payload length (qlen), number of words (wcount), number of whitespace characters (spaces), parity flags for single quotes (sq_mismatch), parity flags for double quotes (dq_mismatch), parity flags for left brackets (lparen_mismatch), and parity flags for right brackets (rparen_mismatch).
[0042] Among them, sq, dq, lparen, rparen, puncts, comments, logic, arith, hexnum, alpha, digit, sqlkw, and sqlfunc primarily describe character or word-level features in the original payload. qlen, wcount, spaces, sq_mismatch, dq_mismatch, lparen_mismatch, and rparen_mismatch focus on structural features such as symbol combinations, matching patterns, and overall payload length and word count. In the original payload, for the same request within the same function, there are significant differences between normal and malicious payloads in several aspects of SQL injection characteristics. First, malicious original payloads typically add extra SQL keywords and special characters, causing their length (qlen) to exceed that of normal original payloads. Second, the word count (wcount) in malicious original payloads is also significantly higher than in normal original payloads, which generally only contain fixed parameter words. The number of single quotes (sq) is also an important distinguishing feature. Normal raw payloads contain few or no single quotes, while malicious raw payloads often include single quotes to disrupt the original quote closure structure and perform malicious splicing. The usage of double quotes follows a similar pattern to single quotes; therefore, the number of double quotes (dq) is also an important characteristic distinguishing malicious raw payloads from normal ones. When disrupting the original closure structure, parentheses may also exist, so the number of left parentheses (lparen) and right parentheses (rparen) is another distinguishing feature. Normal raw payloads similarly contain only a few left and right parentheses or none at all. However, the number of single and double quotes and the number of left and right parentheses are insufficient to determine malicious intent. Therefore, the parity flags for single quotes (sq_mismatch) and double quotes (dq_mismatch) are introduced to further assist in identifying malicious intent. This is mainly because sometimes users might accidentally press the wrong key or for other reasons, resulting in quotes. However, in malicious attacks, the number of quotes is often odd. Therefore, this flag can be used to assist in the judgment. In addition, there should be parity flags for left and right parentheses, which also play a role in determining whether it is malicious. Normal requests generally do not contain parentheses, or if they do, they are in an even number. Therefore, the parity flags for left and right parentheses (lparen_mismatch) and right parentheses (rparen_mismatch) can be combined for a comprehensive judgment. For example, if an attacker constructs a payload like ") = (", even though the number is even, when broken down into left and right parentheses for analysis, it can be identified that the attacker is constructing this type of payload to close the original parentheses.Regarding the number of whitespace characters, normal raw payload parameters and values are usually placed close together without using whitespace characters, while malicious payloads require whitespace characters for word segmentation, thus causing semantic changes. Therefore, the number of whitespace characters is also a characteristic that distinguishes malicious from normal raw payloads in SQL injection. The number of special characters (punctuation) also shows a significant difference. Malicious raw payloads often use special characters (such as !, #, $, %, &, *, +, , -, ., / , :, ;, <, =, >, ?, @, [, \, ], ^, _, `, {, |,}, ~). Normal requests do not contain a large number of special characters. When a large number of special characters are present, it indicates that an attacker may be trying to use these special characters to interfere with detection and bypass it. The presence of comments often indicates code injection or bypass detection behavior. For example, malicious raw payloads often contain multiple lines of comment symbols ( / ** / ). Logical operators (logic) and arithmetic operators (arithm) frequently appear in malicious raw payloads. The former is used for logical judgments to construct malicious queries, while the latter is often used to bypass detection rules, such as using the plus sign (+) to represent a space. Regarding the number of hexadecimal numbers (hexnum), letters (alpha), and numbers (digit), malicious raw payloads often contain a large number of hexadecimal codes, letters, and numbers to obfuscate content and bypass detection, while normal raw payloads contain fewer or no of these. Finally, the number of SQL keywords (sqlkw) and SQL functions (sqlfunc) can also clearly distinguish normal and malicious raw payloads. Normal raw payloads generally do not contain SQL keywords. If they do contain SQL keywords, they are only a small number, and they will not contain a large number of SQL functions, because attacks often use multiple SQL functions to obtain information or achieve bypass effects. For example, attackers may use error reporting functions to cause information leakage. Therefore, malicious raw payloads, in addition to containing a large number of SQL keywords such as "union" and "select", will also contain a large number of SQL functions, such as updatexml() and json_arry().
[0043] In summary, through in-depth analysis and extraction of the above-mentioned SQL injection features, this invention constructs a feature set that effectively distinguishes between normal and malicious original payloads, as shown in Table 1.
[0044] Table 1. SQL Injection Characteristics of the Raw Load
[0045]
[0046]
[0047]
[0048] Of the 20 features listed above, query length (qlen) and word count (wcount) reflect the basic scale of the original payload, thus reflecting the basic scale feature. The number of whitespace characters (spaces) and word count (wcount) typically show an approximately linear relationship, indicating that the separation between words in a normal original payload is relatively uniform, thus reflecting the character separation feature. The counts of single quotes (sq), double quotes (dq), left parentheses (lparen), and right parentheses (rparen), as well as the parity flags for single quotes (sq_mismatch), double quotes (dq_mismatch), left parentheses (lparen_mismatch), and right parentheses (rparen_mismatch), reveal the completeness of the syntax construction in the original payload. Outliers may... The frequency of special symbols (puncts) and comments indicates that attackers may use non-standard delimiters or comments to disrupt syntax parsing, thus reflecting syntax interpretation characteristics. The frequency of logical and arithmetic operators (arithm) indicates the complexity of conditional expressions and numerical operations, thus reflecting complexity characteristics. In addition, the statistics of hexadecimal numbers (hexnum), letters (alpha), and numbers (digit), the matching of SQL keywords (sqlkw), and SQL functions (sqlfunc) together constitute a comprehensive description of the potential SQL injection characteristics in the original payload, thus providing multi-dimensional quantitative evidence for distinguishing between normal queries and malicious SQL injection, thus reflecting comprehensive descriptive characteristics.
[0049] By extracting 20 features from Table 1 from the original payload, a set of numerical values that express the SQL injection characteristics of the original payload can be obtained.
[0050] This invention is a method for feature extraction of the original payload. It mainly extracts the SQL injection features of the original payload, constructs the feature value vector of the original payload, and transforms it into a numerical matrix vector so that the machine learning model can recognize it. The specific expression is shown in formula (1):
[0051]
[0052] in:
[0053] q i , i∈Z + : This is the original load.
[0054] Φ ss (): represents the original load (q) iThis is a set of SQL injection feature extraction functions, through which the SQL injection feature values of each original payload are obtained.
[0055] τ: Φ ss The result of () is converted into an SQL injection feature matrix.
[0056] m, n∈Z + : For standardized calculations, where x mn To perform τ(Φ) ss (q i After transformation, the value in the m-th column and n-th row of the matrix; u m Let m be the sample mean of the m-th column, which is calculated as shown in formula (2):
[0057]
[0058] s m is the sample standard deviation of the m-th column, which is calculated as shown in formula (3):
[0059]
[0060] After the original payload is processed by the method of this invention, a feature matrix expressing the SQL injection characteristics of the original payload will be obtained. According to Table 1 and formulas (1), (2) and (3), as follows Figure 1 As shown, the specific steps of the SQL injection feature extraction method proposed in this invention are as follows:
[0061] Step S1: Construct the feature vector of the original load and convert it into a numerical matrix vector. Initialize the original load feature vector by inputting and traversing the original load dataset.
[0062] Step S2: Read the original payload and define the feature group as the SQL injection feature vector.
[0063] Step S3, as follows Figure 2 As shown, the basic scale features, grammatical construction features, grammatical interpretation features, complexity features, comprehensive description features, and character separation features are extracted sequentially from the feature group.
[0064] Basic scale characteristics include: query length (qlen) and word count (wcount).
[0065] The grammatical construction features include: the number of single quotes (sq), the number of double quotes (dq), the number of left parentheses (lparen), and the number of right parentheses (rparen); it also introduces a parity flag for single quotes (sq_mismatch), a parity flag for double quotes (dq_mismatch), a parity flag for left parentheses (lparen_mismatch), and a parity flag for right parentheses.
[0066] Grammatical interpretation features include the number of puncts and comments.
[0067] Complexity characteristics include the number of logical operators and the number of arithmetic operators.
[0068] The comprehensive description features include: number of hexadecimal values (hexnum), number of letters (alpha), number of numbers (digit), number of SQL keywords (sqlkw), and number of SQL functions (sqlfunc).
[0069] Character segmentation feature: number of whitespace characters.
[0070] Step S4: Construct the feature matrix and output the standardized feature matrix.
[0071] Example 2
[0072] This application proposes an SQL injection feature extraction system, including:
[0073] The feature vector input module is used to construct the feature value vector of the original load and convert it into a numerical matrix vector. It initializes the original load feature vector by inputting and iterating through the original load dataset.
[0074] The feature group definition module is used to define feature groups as SQL injection feature vectors.
[0075] The feature extraction module is used to extract basic scale features, grammatical construction features, grammatical interpretation features, complexity features, comprehensive description features, and character separation features from the feature group in sequence.
[0076] The feature matrix output module is used to construct the feature matrix and output the standardized feature matrix.
[0077] In this application, for the original payload, 20 features are defined to form a complete SQL injection feature vector. These 20 features, from different dimensions, help the machine learning model to train and identify SQL injection more accurately. Adding too many redundant features will lead to a decrease in model performance, while reducing relevant features will decrease the model's recognition ability. For example, if the feature `puncts` is removed, it will be difficult to identify SQL injection based solely on features like `sqlkw` and `sqlfunc`, because a large number of SQL keywords and functions are insufficient to prove SQL injection. However, the simultaneous use of a large number of special characters can give the model stronger confidence in determining whether it is SQL injection. This is because advanced SQL injection often uses a large number of SQL keywords and functions to execute more sophisticated threats, and evades detection by padding or relying on a large number of special characters. Furthermore, the absence of the features `sqlkw` or `sqlfunc` will reduce the model's recognition ability. If an attacker wants to execute more advanced or harmful SQL injection, they must use SQL functions, such as writing to files. Relying solely on SQL keywords will make it difficult for the model to better identify SQL injection. Similarly, without SQL keywords, relying solely on SQL functions makes it difficult to detect SQL injection. In addition, features like `sq`, `dq`, `lparen`, `rparen`, `sq_mismatch`, `dq_mismatch`, `lparen_mismatch`, and `rparen_mismatch` can easily lead to false positives. If a user incorrectly enters quotation marks or parentheses, existing detection methods will directly classify it as SQL injection, leading to a higher false positive rate. However, if we use sq_mismatch and dq_mismatch to assist in detection, for example, if a user accidentally enters a single quote, sq_mismatch will be marked as 1, which does pose a certain risk of SQL injection. But in our invented method, other features are also used as the basis for judgment. For example, the model will observe features such as sqlkw, sqlfunc, and puncts. If they all have high signals, then it can be judged as an SQL injection. But if only the single quote-related features are triggered, while other features are not triggered, then it can be determined that it is not an SQL injection. Therefore, these 20 features are complementary and the most important in the identification of SQL injection. None of them can be omitted.
[0078] This application introduces parity characteristics for single quotes, double quotes, left parentheses, and right parentheses. In the original payload, a feature vector fused with numerical and Boolean values is constructed. The numerical part primarily reflects the strength or scale of the feature signal, while the Boolean part reflects its presence or absence. By fusing these two types of feature values, the machine learning model can learn both the linear or nonlinear effects caused by differences in magnitude and capture key content, thereby achieving information complementarity and improving the model's discriminative ability.
[0079] This application targets the XGBoost model, combining the feature vectors and labels of the original payloads in the dataset into a feature matrix, and then standardizing the feature matrix. In SQL injection detection, there are often problems such as large differences in the dimensions and uneven distribution among SQL injection features. For example, the orders of magnitude of features `qlen` and `comments` are not the same. If directly used for model training, the model will focus on `qlen` and ignore `comments`. However, after standardization, the values of `qlen` and `comments` will be on the same order of magnitude. This standardized value represents the degree to which the feature deviates from the norm, allowing the model to focus on the extent of the feature's deviation from the norm, rather than just its magnitude. To address the problem of dimensional differences among SQL injection features, this invention further scales the feature values in the SQL injection feature matrix, linearly transforming the original feature data according to its sample mean and sample standard deviation. This eliminates the influence of scale differences between different features and enhances the model's ability to distinguish anomalies.
[0080] Example 3
[0081] This application also provides a computer storage medium including instructions, characterized in that, when the instructions are executed on a computer, they cause the computer to perform the method in Embodiment 1 above.
[0082] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0083] The aforementioned computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0084] The embodiments described herein are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made to the structure, shape, and principle of the present invention should be covered within the scope of protection of the present invention. Although preferred embodiments of the present invention have been described, those skilled in the art, once they understand the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention also intends to include these modifications and variations.
Claims
1. A method for extracting SQL injection features, characterized in that, include: Step S1: Initialize the original load feature vector by inputting and traversing the original load dataset; Step S2: Read the raw payload and define the feature group as the SQL injection feature vector; Step S3: Extract the basic scale features, grammatical construction features, grammatical interpretation features, complexity features, comprehensive description features, and character separation features from the feature group in sequence. Step S4: Construct the feature matrix and output the standardized feature matrix.
2. The feature extraction method as described in claim 1, characterized in that, The The basic scale characteristics are: query length and number of words; The grammatical construction features are: the number of single quotes, the number of double quotes, the number of left parentheses, and the number of right parentheses; The grammatical characteristics are: the number of special symbols and the number of annotation symbols; The complexity characteristics are: the number of logical operators and the number of arithmetic operators; The comprehensive description features include: the number of hexadecimal numbers, the number of letters, the number of numbers, the number of SQL keywords, and the number of SQL functions; The character segmentation feature is the number of whitespace characters.
3. The feature extraction method as described in claim 2, characterized in that, The grammatical construction features also introduce parity flags for single quotes, double quotes, left parentheses, and right parentheses.
4. The feature extraction method as described in claim 1, characterized in that, The feature matrix expression in step S4 is: Where: q i , i∈Z + For the original load; Φ ss () represents the original load (q) i The SQL injection feature extraction function set; τ is the function set for extracting Φ. ss The result of () is converted into an SQL injection feature matrix; m, n∈Z + For standardized calculations; i represents the sequence number of the original load data; q i This represents the i-th raw load data.
5. The feature extraction method as described in claim 4, characterized in that, The u m Let u be the sample mean of the m-th column. m The expression is: Where x mn This represents the value in the m-th column and n-th row of the matrix.
6. The feature extraction method according to claim 4, characterized in that, The s m s is the sample standard deviation of the m-th column. m The expression is:
7. A SQL injection feature extraction system, characterized in that, To implement the feature extraction method according to any one of claims 1-6, comprising: The feature vector input module is used to initialize the original load feature vector, input and traverse the original load dataset; The feature group definition module is used to define feature groups as SQL injection feature vectors; The feature extraction module is used to extract basic scale features, syntax construction features, syntax interpretation features, complexity features, comprehensive description features, and character separation features from the feature group in sequence. The feature matrix output module is used to construct the feature matrix and output the standardized feature matrix.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when loaded and executed by a processor, implement the feature extraction method as described in any one of claims 1-6.