Abnormal access detection method and system based on data analysis
By processing SQL log data and using the OCSVM model for anomaly detection, the data scarcity problem of network information security isolation devices in detailed security event analysis and abnormal access detection is solved, and efficient and accurate user behavior anomaly detection is achieved.
Patent Information
- Application Number
- CN202510659338.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
Existing network information security isolation devices make it difficult to collect and analyze security incidents in detail, cannot effectively assess the security status of boundaries, and have difficulty tracing malicious attacks. In addition, abnormal access detection in confidential units faces the problem of data scarcity.
SQL log data is processed through log collection, feature selection, lexical analysis and syntax analysis, and trained using a one-class support vector machine model (OCSVM) to generate anomaly detection classifiers, achieving high-precision description of user behavior and anomaly detection.
It achieves high-precision description of user behavior, reduces data redundancy, simplifies the complexity of model training, improves the efficiency and accuracy of anomaly detection, and can issue alarms in time to prevent potential security risks.
Smart Images

Figure CN120671123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to network data communications, and in particular to a method and system for detecting abnormal access based on data analysis. Background Art
[0002] In recent years, with the rapid development of network attack technology, defense-related confidential units, confidentiality departments, key industries and corporate units are faced with a large number of high-risk operations such as abnormal access, abnormal operations, batch downloads, account theft, etc. in their daily operations, posing a serious threat to system security.
[0003] Currently, key organizations typically employ dual-network isolation as the primary security protection system. Information security network isolation devices are deployed at the network boundary to achieve strong logical isolation, effectively ensuring the secure and reliable operation of network services. However, with the development of networks, the exchange of information between internal and external networks has become increasingly frequent, increasing the potential for network attacks. As the second line of defense for information networks, the internal and external network boundary carries a large amount of information exchange between critical external services and internal databases.
[0004] However, current network information security isolation devices are only capable of online, real-time SQL statement filtering, lacking the capacity for more detailed security event collection and analysis, making it difficult to assess perimeter security status and trace malicious attacks. Therefore, while maintaining high-intensity security protection capabilities, it is urgent to establish online database monitoring capabilities to further enhance the overall defense capabilities of the power information network. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the present invention provides the following technical solutions:
[0006] A method and system for detecting abnormal access based on data analysis, the method comprising the following steps:
[0007] S1, log collection, collects and centrally stores the original SQL logs generated by the database management system;
[0008] S2, feature selection, extracts key information that can represent user behavior while reducing data redundancy;
[0009] S3, through lexical analysis and syntax analysis, converts raw SQL log data into structured datasets that can be used by machine learning algorithms;
[0010] S4, trains the dataset in S3 through a one-class support vector machine model, namely OCSVM, to generate a classifier for anomaly detection;
[0011] S5, model evaluation and validation, divides the data set in S3 into a training set and a test set. The training set is used to train the one-class support vector machine model, and the test set is used to evaluate the performance of the trained one-class support vector machine model.
[0012] A method and system for detecting abnormal access based on data analysis, the system includes: a SQL log preprocessing module, a machine learning training module, and a model evaluation and verification module;
[0013] The SQL log preprocessing module is responsible for processing the original audit SQL log data and generating input data for OCSVM training and recognition. The SQL log preprocessing module processing flow includes: log collection, feature selection, and SQL lexical analysis, syntax analysis and numerical processing;
[0014] The machine learning training module includes: an OCSVM model and a detection unit; the machine learning training module uses the data obtained by the SQL log preprocessing module as input samples for OCSVM training. After OCSVM training, a classifier is obtained for use by the detection unit to detect SQL log data, thereby realizing abnormal detection of user access behavior;
[0015] The model evaluation and verification module is used for training the OCSVM model and evaluating the performance of the OCSVM model.
[0016] Furthermore, in S2, the selected key information includes user name, operation time, client IP address, and operation type.
[0017] Furthermore, in S3, the input SQL statement is decomposed into a series of lexical units through lexical analysis to generate a lexical unit sequence; then, through grammatical analysis, the lexical unit sequence is read and combined into a statement or expression with a grammatical structure to construct a grammar tree.
[0018] Furthermore, in S4, the data set is: The one-class support vector machine model aims to find a hyperplane that is as far away from the origin as possible and separates most or all normal data points. The optimization objective of the one-class support vector machine model can be expressed as:
[0019]
[0020] Among them, ν is a hyperparameter that controls the ratio of support vectors, ξ i is a slack variable used to deal with inseparable situations; the objective function can be viewed as minimizing the misclassified support vector while maximizing the interval between the hyperplane and the origin; through the Lagrange multiplier method, the dual problem of a class of support vector machine models can be obtained.
[0021] Furthermore, by the Lagrange multiplier method, for each training sample x i Introducing the Lagrange multiplier α i and β i , we can get the following dual problem:
[0022]
[0023] in,
[0024] K(x i ,x j )=φ(x i ) T φ(x j )
[0025] is the kernel function, which allows computing inner products in high-dimensional feature spaces;
[0026] After training, the decision function of a class of support vector machine models can be expressed as:
[0027]
[0028] Among them, α i is a non-zero Lagrange multiplier corresponding to the support vector, K(x i ,x) is the value of the kernel function, and ρ is the intercept obtained through training; therefore, when a new sample point is mapped to the feature space, it can be classified according to its position relative to the hyperplane to determine whether it is an abnormal sample.
[0029] Furthermore, in S5, the indicators of model evaluation include accuracy, drawing ROC curves and calculating AUC values, evaluating the classification ability of the model under different thresholds, and using the K-fold cross-validation method to divide the training set multiple times, train multiple models and average the evaluation results to verify the stability and generalization ability of the model.
[0030] Compared with the existing technology, the technical solution of this application has the following beneficial effects:
[0031] 1. The present invention uses the SQL log preprocessing module to select key information that can represent user behavior, reduce data redundancy, and use lexical analysis and syntax analysis to analyze log data with efficient data preprocessing methods to achieve high-precision description of user behavior. It also selects features from the original log, reduces data redundancy, and ensures efficient training.
[0032] 2. In view of the limited training data available in key confidential units, the present invention classifies user behavior based on a first-class support vector machine (OCSVM). Only normal behavior data is needed to establish a detection model, and no abnormal data is required, which greatly simplifies the complexity of model training. Compared with traditional methods, OCSVM has distinct advantages in processing single-class data and anomaly detection. Secondly, the present invention adopts efficient data preprocessing means to parse log data through lexical analysis and grammatical analysis to achieve high-precision description of user behavior; feature selection is performed on the original log to extract sensitive information such as user name, operation time, client IP, etc., to reduce data redundancy and ensure high efficiency of model training and recognition. The anomaly detection method proposed in the present invention can efficiently analyze abnormal user behavior information in the database, which helps to issue alarms in time and take necessary countermeasures to prevent potential security risks or data leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the technical solution flow of the present invention;
[0034] Figure 2 Schematic diagram of the system structure of the present invention;
[0035] Figure 3 This is a flow chart of the lexical analysis steps of the present invention;
[0036] Figure 4 This is a flow chart of the syntax analysis steps of the present invention;
[0037] Figure 5 Schematic diagram of actual detection results of the present invention. DETAILED DESCRIPTION
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] Example 1:
[0040] See also Figure 1-5 , a method and system for detecting abnormal access based on data analysis, the method comprising the following steps:
[0041] S1, log collection, collects and centrally stores the original SQL logs generated by the database management system. These logs contain various user operation information, including query, insert, update and delete;
[0042] S2, feature selection. Each raw SQL log is very complex and redundant and cannot be used directly for training. For example, the main information of user operations on the database includes: the username of the database login, the start time of the user logging into the database, the time when the user performs the operation, the client IP address and port number, the IP address and port number of the database server, and the SQL statement submitted by the user to the database. This information describes the access behavior of database users from different aspects. However, there is a lot of redundancy in these raw data. In order to improve the efficiency of OCSVM training and recognition, it is necessary to perform accurate feature selection to extract key features from the raw logs to describe user behavior. The goal of feature selection is to extract key information that can represent user behavior while reducing data redundancy. Common features include: username, operation time, client IP address, and operation type (such as SELECT, INSERT, UPDATE, DELETE).
[0043] S3, by using the LEX lexical analyzer and the YACC parser for lexical analysis and grammatical analysis, converts the original SQL log data into a structured dataset that can be used by machine learning algorithms;
[0044] The purpose of lexical analysis is to decompose the input SQL statement into a series of lexical units (Tokens). Each lexical unit represents a basic element in the SQL statement, such as a keyword, identifier, operator, etc.
[0045] The syntax analysis phase combines the sequence of lexical units into a statement or expression with a grammatical structure based on the lexical analysis. The purpose is to generate a syntax tree or abstract syntax tree (AST) based on the sequence of lexical units. This tree structure represents the grammatical structure of the SQL statement. The key processing steps of lexical analysis and syntax analysis are as follows: Figure 3 、 Figure 4 As shown;
[0046] S4, trains the dataset in S3 through a one-class support vector machine model, namely OCSVM, to generate a classifier for anomaly detection;
[0047] The One-Class Support Vector Machine (OCSVM) model is a variant of the SVM and is used for anomaly detection and outlier detection problems. Unlike traditional SVMs, which can only handle binary classification problems, the OCSVM aims to learn a hyperplane that describes the characteristics of positive samples by using only positive samples and keep negative samples as far away from the hyperplane as possible.
[0048] In OCSVM, training samples only include positive samples. The goal is to find an optimal hyperplane so that positive samples are as far above the hyperplane as possible and negative samples are as far below the hyperplane as possible. Specifically, the OCSVM principle is described in the following formula: The dataset after the SQL log data preprocessing module is: OCSVM aims to find a hyperplane that is as far away from the origin as possible and separates most (or all) normal data points; therefore, its optimization objective can be expressed as:
[0049]
[0050] Among them, ν is a hyperparameter that controls the ratio of support vectors, ξ i is a slack variable used to deal with inseparable cases; this objective function can be viewed as minimizing the misclassified support vector while maximizing the interval between the hyperplane and the origin; through the Lagrange multiplier method, the dual problem of OCSVM can be obtained, that is, for each training sample x i Introducing the Lagrange multiplier α i and β i , we can get the following dual problem:
[0051]
[0052]
[0053] in,
[0054] K(x i ,x j )=φ(x i ) T φ(x j )
[0055] is a kernel function that allows the calculation of inner products in high-dimensional feature space. After training, the decision function of OCSVM can be expressed as:
[0056]
[0057] Among them, α i is a non-zero Lagrange multiplier corresponding to the support vector, K(x i ,x) is the value of the kernel function, ρ is the intercept obtained through training;
[0058] Therefore, when a new sample point is mapped to the feature space, it can be classified according to its position relative to the hyperplane to determine whether it is an abnormal sample;
[0059] Database user behavior anomaly detection based on one-class support vector machine (OCSVM) has the following significant advantages and characteristics:
[0060] (1) Good single-class data processing effect. OCSVM is a variant of support vector machine (SVM). Its characteristic is that it only requires normal class data as training samples, and does not require abnormal class data. This is particularly important in practical applications because in the abnormal access detection of defense confidential units, confidentiality departments, key industries and enterprises, abnormal data is often difficult to obtain or very rare, while normal data is relatively easy to obtain. Therefore, using OCSVM can simplify the model training process and avoid the difficulty of requiring a large amount of labeled abnormal data.
[0061] (2) Efficient model training. Compared with traditional methods that require both normal and abnormal data for training, OCSVM's training process is simpler and more efficient. It can quickly build a model for identifying normal behavior and determine whether unknown data is abnormal by capturing the characteristics and distribution of normal data.
[0062] (3) Strong generalization ability. OCSVM is based on the theory of support vector machines and has excellent generalization ability. This means that even when faced with a small amount of normal data and a variety of abnormal situations, the model can effectively identify and classify unknown data, thereby improving the system's adaptability and response speed to unknown threats.
[0063] S5, model evaluation and validation, divides the data set in S3 into a training set and a test set. The training set is used to train the one-class support vector machine model, and the test set is used to evaluate the performance of the trained one-class support vector machine model.
[0064] After completing SQL log data collection, feature selection, and data conversion, the processed data set is divided into a training set and a test set. The training set is used for model training, and the test set is used to evaluate the performance of the model. The ratio is set to 8:2.
[0065] The performance of the trained OCSVM model was evaluated using the test set. Evaluation metrics included accuracy, ROC curve plotting, and AUC calculation to assess the model's classification capabilities at different thresholds. To verify the model's stability and generalization capabilities, the training set was divided multiple times using the K-fold cross-validation method. Multiple models were trained and the evaluation results were averaged. The overall process is as follows:
[0066] #Predict test set data
[0067] #Calculate evaluation metrics
[0068] #Calculate ROC curve and AUC
[0069] #K-fold cross validation evaluation model
[0070] After optimization, multiple samples are tested in the detection unit and the following results are obtained: Figure 5 As shown in the accuracy rate, the detection rate of unauthorized operations in the test samples is 82.4%, the detection rate of illegal user operations is 95.6%, the detection rate of access to sensitive database resources is 98.1%, and the detection rate of normal behavior is 88.1%. It can be seen that the technical solution of the present application has a very high detection rate, especially the detection rate for illegal user operations and access to sensitive database resources.
[0071] Example 2:
[0072] See also Figure 1-5 ,According to embodiment 1, a method and system for detecting abnormal access based on data analysis,,the system includes: a SQL log preprocessing module, a machine learning training module,,a model evaluation and verification module;
[0073] The SQL log preprocessing module is the foundation of the entire user behavior anomaly detection system. It is responsible for processing raw audit SQL log data and generating input data for OCSVM training and recognition. The SQL log preprocessing module processing flow includes: log collection, feature selection, and SQL lexical analysis, syntax analysis, and numerical processing.
[0074] The machine learning training module is a key part of the entire abnormal access detection system, including: OCSVM model and detection unit. The machine learning training module uses the data obtained by the SQL log preprocessing module as the input sample for OCSVM training. After OCSVM training, the classifier is obtained and used by the detection unit to detect SQL log data, thereby realizing abnormal user access behavior detection.
[0075] The model evaluation and verification module is an indispensable part of the user behavior anomaly detection system based on machine learning. This module ensures that the established anomaly detection model can run robustly and effectively detect abnormal behavior in actual scenarios through multi-stage data processing, training and evaluation steps.
[0076] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for detecting abnormal access based on data analysis, characterized in that: The following steps are involved: S1, log collection, collects and centrally stores the original SQL logs generated by the database management system; S2, feature selection, extracts key information that can represent user behavior while reducing data redundancy; S3, through lexical analysis and syntax analysis, converts raw SQL log data into structured datasets that can be used by machine learning algorithms; S4, trains the dataset in S3 through a one-class support vector machine model, namely OCSVM, to generate a classifier for anomaly detection; S5, model evaluation and verification, divides the data set in S3 into a training set and a test set. The training set is used to train the one-class support vector machine model, and the test set is used to evaluate the performance of the trained one-class support vector machine model.
2. The abnormal access detection method based on data analysis according to claim 1, characterized in that: In S2, the selected key information includes user name, operation time, client IP address, and operation type.
3. The abnormal access detection method based on data analysis according to claim 1 is characterized in that: In S3, the input SQL statement is decomposed into a series of lexical units through lexical analysis to generate a lexical unit sequence; then, through syntax analysis, the lexical unit sequence is read and combined into statements or expressions with grammatical structure to construct a syntax tree.
4. The abnormal access detection method based on data analysis according to claim 1, characterized in that: The dataset is: The one-class support vector machine model aims to find a hyperplane that is far away from the origin and separates most or all normal data points. The optimization objective of the one-class support vector machine model can be expressed as: Among them, ν is a hyperparameter that controls the ratio of support vectors, ξ i is a slack variable used to deal with inseparable situations; the objective function can be viewed as minimizing the misclassified support vector while maximizing the interval between the hyperplane and the origin; through the Lagrange multiplier method, the dual problem of a class of support vector machine models can be obtained.
5. The abnormal access detection method based on data analysis according to claim 4 is characterized in that: By Lagrange multiplier method, for each training sample x i Introducing the Lagrange multiplier α i and β i , we can get the following dual problem: in, K(x i ,x j )=φ(x i ) T φ(x j ) is the kernel function, which allows computing inner products in high-dimensional feature spaces; After training, the decision function of a class of support vector machine models can be expressed as: Among them, α i is a non-zero Lagrange multiplier corresponding to the support vector, K(x i ,x) is the value of the kernel function, and ρ is the intercept obtained through training; therefore, when a new sample point is mapped to the feature space, it can be classified according to its position relative to the hyperplane to determine whether it is an abnormal sample.
6. The abnormal access detection method based on data analysis according to claim 1, characterized in that: In S5, the model evaluation indicators include accuracy, drawing ROC curves and calculating AUC values to evaluate the classification ability of the model at different thresholds, and using the K-fold cross-validation method to divide the training set multiple times, train multiple models and average the evaluation results to verify the stability and generalization ability of the model.
7. A system comprising the abnormal access detection method based on data analysis according to any one of claims 1 to 6, characterized in that: include: SQL log preprocessing module, machine learning training module, model evaluation and verification module; The SQL log preprocessing module is responsible for processing the original audit SQL log data and generating input data for OCSVM training and recognition. The SQL log preprocessing module processing flow includes: log collection, feature selection, and SQL lexical analysis, syntax analysis and numerical processing; The machine learning training module includes: an OCSVM model and a detection unit; the machine learning training module uses the data obtained by the SQL log preprocessing module as input samples for OCSVM training. After OCSVM training, a classifier is obtained for use by the detection unit to detect SQL log data, thereby realizing abnormal detection of user access behavior; The model evaluation and verification module is used for training the OCSVM model and evaluating the performance of the OCSVM model.
Citation Information
Cited By
Charging selection probability prediction method and system for expressway EV users
CN121365997A
A charging selection probability prediction method and system for highway ev users
CN121365997B