Database user abnormal behavior detection method

By adopting DTMC-based user behavior abnormality detection methods in the database system, the problem that the prior art is difficult to effectively detect and prevent abnormal behavior of database users is solved, and fast and accurate abnormality detection is achieved, which improves detection efficiency and accuracy.

CN120046143APending Publication Date: 2025-05-27SHANGHAI YOUJIA NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150189.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and prevent abnormal behaviors of database users, especially to effectively stop internal attacks and abuse of permissions, resulting in data leakage, tampering or corruption.

Method used

A database user abnormal behavior detection method is adopted, including data preprocessing module, data training module and exception detection module. The data preprocessing module extracts meaningful features by cleaning and denoising the original data. The data training module uses discrete time Markov chain (DTMC) to construct a user behavior feature model. The abnormality detection module monitors user behavior in real time, and judges whether the behavior is abnormal by comparing the deviation between the detection sequence and historical data.

Benefits of technology

By building a DTMC-based user behavior abnormality detection system, it can quickly and accurately analyze user behavior, reduce computing complexity, improve detection efficiency, and directly understand whether the user has performed unusual operations, which has higher accuracy and targeting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046143A_ABST
    Figure CN120046143A_ABST
Patent Text Reader

Abstract

The invention discloses a database user abnormal behavior detection method, and belongs to the technical field of network information security, the database user abnormal behavior detection method comprises three subsystems of a data preprocessing module, a data training module and an abnormal detection module, the data preprocessing module is responsible for identifying and processing original data, cleaning and removing repeated data, the data training module is responsible for receiving data transmitted by the data preprocessing module, and the anomaly detection module analyzes and judges database user behavior data obtained in real time by using a trained model. The DTMC-based user behavior anomaly detection system is constructed by applying the discrete time Markov chain to database system anomaly detection, so that when the system analyzes user behaviors, only the current state and the transition probability need to be considered, all previous historical states do not need to be considered, the analysis process is greatly simplified, and the analysis efficiency is improved. The calculation complexity is reduced, and the detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network information security, and particularly relates to a method for detecting abnormal behaviors of database users. Background Art

[0002] In today's digital age, data has become the core asset of enterprises and organizations. A large amount of key information is stored in databases, covering user profiles, financial data, business records, etc. This data is crucial for the operation, decision-making, and competitiveness of enterprises. With the continuous increase in the value of data, it has become the main target of various attackers. Abnormal behaviors of database users may lead to data leakage, tampering, or destruction, causing huge economic losses and reputational damage to enterprises. Therefore, effective means are needed to detect and prevent such behaviors.

[0003] In the overall security of information systems, databases are often the most attractive targets for attackers. The fundamental purpose of many network attacks is to obtain important information stored in databases. Traditional database security protection methods have improved the security of database systems to a certain extent, but most of them are passive security technologies that focus on prevention and are unable to effectively stop intrusion behaviors, especially powerless against internal attacks such as abuse of database user permissions. According to statistics, 80% of attacks on databases come from within, and internal attacks are the main threats to database systems. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method for detecting abnormal behaviors of database users.

[0005] The technical solution adopted to solve the above technical problem is: A method for detecting abnormal behaviors of database users, including three major subsystems: a data preprocessing module, a data training module, and an anomaly detection module. The data preprocessing module is responsible for identifying and processing the cleaning of raw data to remove duplicate data. The data training module is responsible for receiving the data transmitted by the data preprocessing module and extracting and selecting features meaningful for anomaly behavior detection from the preprocessed data. The anomaly detection module uses the trained model to analyze and judge the real-time database user behavior data, and monitors every operation behavior of the user in real time to determine whether it conforms to the normal behavior pattern.

[0006] Further, the data preprocessing module mainly includes a data acquisition subunit and a preprocessing subunit;

[0007] For the data acquisition subunit, SQL statements are used as the data source for anomaly detection, and it is necessary to collect the user SQL operation requests in the database log file;

[0008] For the preprocessing subunit, methods of data cleaning, data denoising, and data transformation are adopted; data cleaning mainly performs duplicate removal, null value removal, and anomaly removal on the data to ensure the accuracy and integrity of the data; data denoising mainly filters out the noise in the data to improve the quality of the data; data transformation mainly transforms and standardizes the original data for subsequent data analysis and modeling. At the same time, the preprocessing subunit processes all the collected data into a sequence according to the requirements of the normal sequence library, extracts the keywords in the SQL request, and combines them into a training sequence.

[0009] Furthermore, the data training module includes a sequence library establishment subunit and a behavior feature extraction subunit based on DTMC:

[0010] For the sequence library establishment subunit; for the preprocessed sequence, W different lengths of short sequences are defined, and corresponding occurrence frequency weights are set, x = (x 1 , x 2 , …, x n ) is a sequence, is called a stream of, where is a short sequence of length k intercepted from x. Since the length of the stream is determined by k, is also called the stream generated by x with length k;

[0011] l(i) represents the length of the i-th short sequence, and e(i) represents the occurrence frequency weight of the i-th short sequence, where i = 1, 2, …, W, and l(1) < l(2) < … < l(W);

[0012] Obtain the preprocessed data and generate W different lengths of short sequences. The finally collected and preprocessed data will form a sequence, represented by s = (s 1 , s 2 , …, s r ), where r represents the number of SQL query fields in the sequence, and s j' (1 ≤ j' ≤ r) represents the j'-th field in the sequence s. Generate W different lengths of streams 1 of the sequence s = (s 2 , …, s r ). is a stream of length l(i), and there is

[0013]

[0014] (1 ≤ i ≤ W, 1 ≤ j' ≤ r - l(i) + 1)

[0015] Extract different short sequences in each stream and calculate their occurrence frequencies. In the data stream the different types of short sequences can be represented as m(i) is the number of different types of short sequences in the data stream Calculate the occurrence frequency of each short sequence according to the following steps;

[0016] is a short sequence of length l(i), represents the number of times it appears in the data stream Then the occurrence frequency in is

[0017]

[0018] Merge all the different short sequences in all data streams into the normal sequence library and calculate their weighted occurrence frequencies in the normal sequence library;

[0019] Merge all the different short sequences into the normal sequence library, then the normal sequence library LGS can be represented as:

[0020]

[0021] The calculation method of the weighted occurrence frequency of each short sequence in LGS is calculated as shown above.

[0022] For the behavior feature extraction subunit based on DTMC, construct a discrete-time Markov chain using the states of short sequences of different lengths as the behavior features of the user. The number of states N needs to be set in advance. It is necessary to obtain the behavior patterns of each command in the training data s=(s 1 , s 2 , …, s r ), and then generate the Markov chain states of each behavior pattern according to the N-1 groups in LGS, thereby obtaining a set of Markov chain states, and finally establish the transition probability matrix P of the Markov chain.

[0023] Furthermore, the anomaly detection module compares the detection sequence with the historical data to determine whether they belong to the same user. The deviation of the behavior features of the detected user from the historical data is an important indicator of intrusion. The collected data is preprocessed to form a sequence, c=(c 1 , c 2 , …, c t ) represents the detection sequence after preprocessing, where c j is the j-th command in the detection sequence, and t represents the number of commands in the detection sequence;

[0024] According to the behavior pattern extraction algorithm and the Markov chain state generation algorithm, taking the normal sequence library T and the detection sequence c = (c 1 , c 2 , …, c t ) as inputs, the corresponding state sequence q = (q 1 , q 2 , …, q M' ) can be obtained. The short sequences intercepted from c = (c 1 , c 2 , …, c t ) may not be included in the LGS. Therefore, in the state sequence q = (q 1 , q 2 , …, q M' ), the state value may be N;

[0025] Obtain the classification value corresponding to the state sequence. In the state sequence q = (q 1 , q 2 , …, q M' ), for each state q i (1 ≤ i ≤ M' - 1), the transition probability from q i to q i+1 is

[0026]

[0027] The classification value transferred from

[0028]

[0029] In the formula: is the classification value transferred from q i to q i+1 ; δ is the probability threshold value that needs to be set in advance. When , it represents a normal transition; otherwise, it indicates an abnormal transition. After calculating the classification values of all states, a sequence of classification values

[0030] At the in the classification value sequence , its decision value is defined as:

[0031]

[0032] In the formula: D(n) is the decision value corresponding to the classification value , n is incremented by 1 (w ≤ n ≤ M' - 1), and w is the window size;

[0033] The threshold λ for determining whether a user's behavior is abnormal needs to be set in advance according to system requirements. Whether the current behavior of the user being detected is abnormal is determined by D(n) and λ. If D(n) ≥ λ, it is considered that the current behavior of the user being detected is normal. If D(n) < λ, it is considered that the current behavior of the user being detected is abnormal;

[0034] The current behavior corresponds to w classification values and w + 1 states (q n-w+1 , q n-w+2 , …, q n , q n+1 ) are related. The threshold λ is a parameter with high sensitivity. The higher the threshold, the higher the success rate of anomaly detection, and at the same time, the higher the false alarm rate. The lower the threshold, the lower the success rate of anomaly detection, and the false alarm rate also decreases.

[0035] The beneficial effects of the present invention are as follows: By applying the discrete-time Markov chain to the anomaly detection of database systems, the present invention constructs a user behavior anomaly detection system based on DTMC, analyzes the SQL statements submitted by users as user behavior characteristics, and uses DTMC to extract the behavior characteristics of normal users and behaviors to be detected respectively, and compares the two. If the deviation degree between the two exceeds the threshold, the behavior is determined to be abnormal. In this way, when the system analyzes user behavior, it only needs to consider the current state and transition probability, without considering all previous historical states, greatly simplifying the analysis process, reducing the computational complexity, improving the detection efficiency, being able to quickly process and analyze a large amount of user behavior data, and directly understanding whether the user has performed unusual operations, such as attempting to access sensitive data, performing large-scale data deletion, etc. Compared with other indirect behavior characteristics, it has higher accuracy and pertinence. Brief Description of the Drawings

[0036] Figure 1 is a schematic diagram of the modules of the present invention. Detailed Embodiment

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0038] Such as Figure 1As shown in the figure, a method for detecting abnormal behaviors of database users in this embodiment includes three major subsystems: a data preprocessing module, a data training module, and an anomaly detection module. The data preprocessing module is responsible for identifying and processing the cleaning of raw data to remove duplicate data. The data training module is responsible for receiving the data transmitted from the data preprocessing module and extracting and selecting features meaningful for abnormal behavior detection from the preprocessed data. The anomaly detection module uses the trained model to analyze and judge the database user behavior data obtained in real time, and monitors every operation behavior of the user in real time to determine whether it conforms to the normal behavior pattern.

[0039] The data preprocessing module mainly includes a data collection subunit and a preprocessing subunit;

[0040] For the data collection subunit, SQL statements are used as the data source for anomaly detection, and it is necessary to collect the user SQL operation requests in the database log file;

[0041] For the preprocessing subunit, methods of data cleaning, data denoising, and data transformation are adopted; data cleaning mainly performs deduplication, removal of blanks, and anomaly processing on the data to ensure the accuracy and integrity of the data; data denoising mainly filters the noise in the data to improve the data quality; data transformation mainly transforms and standardizes the raw data for subsequent data analysis and modeling. At the same time, the preprocessing subunit processes all the collected data into a sequence according to the requirements of the normal sequence library, extracts the keywords in the SQL request, and combines them into a training sequence.

[0042] The data training module includes a sequence library establishment subunit and a behavior feature extraction subunit based on DTMC:

[0043] For the sequence library establishment subunit; for the preprocessed sequence, W different lengths of short sequences are defined, and corresponding occurrence frequency weights are set. x=(x 1 , x 2 , …, x n ) is a sequence, is called a stream of, where is a short sequence of length k intercepted from x. Since the length of the stream is determined by k, is also called the stream generated by x with length k;

[0044] l(i) represents the length of the i-th short sequence, and e(i) represents the occurrence frequency weight of the i-th short sequence, where i = 1, 2, …, W, and l(1) < l(2) < … < l(W);

[0045] Obtain preprocessed data and generate W short sequences of different lengths. The data after collection and preprocessing will ultimately form a sequence, represented by s=(s 1 , s 2 , …, s r ). Among them, r represents the number of SQL query fields in the sequence, and s j' (1≤j'≤r) represents the j'-th field in the sequence s. Generate W different-length streams of the sequence s=(s 1 , s 2 , …, s r ). The stream is a stream of length l(i), and there is

[0046]

[0047] (1≤i≤W, 1≤j'≤r - l(i)+1)

[0048] Extract different short sequences in each stream and calculate their occurrence frequencies. In the data stream , the different types of short sequences can be represented as m(i) is the number of different types of short sequences in the data stream . Calculate the occurrence frequency of each short sequence according to the following steps;

[0049] is a short sequence of length l(i), represents 's occurrence times in the data stream . Then the occurrence frequency of is

[0050]

[0051] Merge all different short sequences in all data streams into the normal sequence library and calculate their weighted occurrence frequencies in the normal sequence library;

[0052] Merge all different short sequences into the normal sequence library. Then the normal sequence library LGS can be represented as:

[0053]

[0054] The calculation method of the weighted occurrence frequencies of various short sequences in LGS is calculated as shown above.

[0055] For the behavior feature extraction subunit based on DTMC, construct a discrete-time Markov chain using the states of short sequences of different lengths as the user's behavior feature. The number of states N needs to be set in advance. It is necessary to obtain the training data s=(s 1 , s 2 , …, sr ) The behavior pattern of each command in it is then used to generate the Markov chain state of each behavior pattern according to N - 1 groups in the LGS, thus obtaining a set of Markov chain states. Finally, the transition probability matrix P of the Markov chain is established.

[0056] The anomaly detection module compares the detection sequence with the historical data to determine whether they belong to the same user. The deviation of the behavior characteristics of the detected user from the historical data is an important indicator of intrusion. The collected data is pre - processed to form a sequence, c=(c 1 , c 2 , …, c t ) represents the detection sequence after pre - processing, where c j is the j - th command in the detection sequence, and t represents the number of commands in the detection sequence;

[0057] According to the behavior pattern extraction algorithm and the Markov chain state generation algorithm, taking the normal sequence library T and the detection sequence c=(c 1 , c 2 , …, c t ) as the input, the corresponding state sequence q=(q 1 , q 2 , …, q M' ) can be obtained. The short sequence intercepted from c=(c 1 , c 2 , …, c t ) may not be included in the LGS. Therefore, in the state sequence q=(q 1 , q 2 , …, q M' ), the state value may be N;

[0058] Obtain the classification value corresponding to the state sequence. In the state sequence q=(q 1 , q 2 , …, q M' ), for each state q i (1≤i≤M' - 1), the transition probability from q i to q i+1 is,

[0059]

[0060] The classification value defined as the transition from to is:

[0061]

[0062] In the formula: is the classification value of the transition from q i to q i+1 ; δ is the probability critical value that needs to be set in advance. When When it is [a certain condition], it represents a normal transition; otherwise, it indicates that an abnormal transition has occurred. After calculating the classification values of all states, a sequence of classification values is obtained.

[0063] In the sequence of classification values at the position of , its decision value is defined as:

[0064]

[0065] In the formula, D(n) is the classification value corresponding decision value, n is incremented by 1 successively (w ≤ n ≤ M' - 1), and w is the window size;

[0066] The threshold λ for determining whether the user's behavior is abnormal needs to be set in advance according to the system requirements. Whether the current behavior of the currently detected user is abnormal is determined by D(n) and λ. If D(n) ≥ λ, it is considered that the current behavior of the detected user is normal. If D(n) < λ, it is considered that the current behavior of the detected user is abnormal;

[0067] The current behavior corresponds to w classification values related to w + 1 states (q n-w+1 , q h-w+2 , …, q n , q n+1 ). The threshold λ is a parameter with high sensitivity. The higher the threshold, the higher the success rate of abnormal detection, and at the same time, the higher the false alarm rate. The lower the threshold, the lower the success rate of abnormal detection, and the false alarm rate also decreases.

[0068] The above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention.

Claims

1. A method for detecting abnormal behavior of database users, characterized by: It includes three subsystems: data preprocessing module, data training module and anomaly detection module. The data preprocessing module is responsible for identifying and processing the cleaning of raw data to remove duplicate data. The data training module is responsible for receiving data transmitted by the data preprocessing module and extracting and selecting features that are meaningful for abnormal behavior detection from the preprocessed data. The anomaly detection module uses the trained model to analyze and judge the database user behavior data obtained in real time, monitors each user's operation behavior in real time, and determines whether it conforms to the normal behavior pattern.

2. A method for detecting abnormal behavior of database users according to claim 1, characterized in that: The data preprocessing module mainly includes a data acquisition subunit and a preprocessing subunit; For the data collection subunit, SQL statements are used as the data source for anomaly detection, and user SQL operation requests in the database log file need to be collected; For the preprocessing subunit, data cleaning, data denoising and data conversion methods are used; Data cleaning mainly involves removing duplicates, empty spaces, and anomalies from data to ensure data accuracy and integrity; data denoising mainly involves filtering out noise in data to improve data quality; Data conversion mainly involves converting and standardizing the original data for subsequent data analysis and modeling. At the same time, the preprocessing subunit processes all collected data into a sequence according to the requirements of the normal sequence library, extracts the keywords in the SQL request, and combines them into a training sequence.

3. A method for detecting abnormal behavior of database users according to claim 1, characterized in that: The data training module includes a sequence library building subunit and a DTMC-based behavior feature extraction subunit: Subunits are established for the sequence library; for the preprocessed sequences, W short sequences of different lengths are defined and the corresponding frequency weights are set, x = (x1, x2, ..., x n ), is a sequence, is called a stream, where is a short sequence of length k intercepted by x. The length of is determined by k, It is also called generating a stream of length k from x; l(i) represents the length of the i-th short sequence, e(i) represents the frequency weight of the i-th short sequence, where i = 1, 2, ..., W, l(1) <l(2)<…<l(W); The preprocessed data is obtained and W short sequences of different lengths are generated. The collected and preprocessed data will eventually form a sequence, which is composed of s = (s1, s2, ..., s r ), where r represents the number of SQL query fields in the sequence, and s j’ (1≤j'≤r) represents the j'th field in the sequence s, generating the sequence s = (s1, s2, ..., s r ) of W different lengths is a flow of length l(i) and has, (1≤i≤W,1≤j'≤rl(i)+1) Extract different short sequences in each stream and calculate their frequency of occurrence. In the above example, different types of short sequences can be represented as m(i) is the data stream The number of different short sequence types in , calculate the frequency of occurrence of each short sequence according to the following steps; is a short sequence of length l(i), express In the data flow The number of times it appears in (1≤i≤W), then exist The frequency of occurrence is: Merge all different short sequences in all data streams into a normal sequence library, and calculate their weighted occurrence frequencies in the normal sequence library; All different short sequences are merged into the normal sequence library, and the normal sequence library LGS can be expressed as: The weighted frequency calculation methods of various short sequences in LGS are calculated as shown above.

4. A method for detecting abnormal behavior of database users according to claim 3, characterized in that: For the behavior feature extraction subunit based on DTMC, a discrete-time Markov chain is constructed using the states of short sequences of different lengths as the user's behavior features. The number of states N needs to be set in advance, and the training data s = (s1, s2, ..., s r ), and then generate the Markov chain state of each behavior pattern according to the N-1 groups in LGS, thereby obtaining a Markov chain state set, and finally establishing the Markov chain transition probability matrix P.

5. A method for detecting abnormal behavior of database users according to claim 1, characterized in that: The anomaly detection module compares the detection sequence with the historical data to determine whether the two belong to the same user. The deviation between the behavior characteristics of the detected user and the historical data is an important indicator of intrusion. The collected data is preprocessed to form a sequence, c = (c1, c2, ..., c t ) represents the detection sequence after preprocessing, where c j is the jth command in the detection sequence, and t represents the number of commands in the detection sequence; According to the behavior pattern extraction algorithm and the Markov chain state generation algorithm, the normal sequence library T and the detection sequence c = (c1, c2, ..., c t ) as input, we can get the state sequence q=(q1, q2, ..., q M’ ), by c=(c1,c2,…,c t ) may not be included in LGS, so in the state sequence q = (q1, q2, ..., q M’ ), the state value may be N; The corresponding state sequence obtains the classification value. In the state sequence q=(q1,q2,…,q M’ ), for each state q i (1≤i≤M'-1), by q i and q i+1 The transition probability is, The classification value transferred from to is defined as: Where: For i Transfer to q i+1 The classification value; δ is the probability critical value, which needs to be set in advance. , it represents a normal transition; otherwise, it means an abnormal transition has occurred. After calculating the classification values ​​of all states, a sequence of classification values ​​is obtained. In the categorical value sequence In The judgment value is defined as: Where: D(n) is the classification value The corresponding judgment value, n, is incremented by 1 (w≤n≤M'-1), where w is the window size; The threshold λ for determining whether the user behavior is abnormal needs to be set in advance according to system requirements. Whether the current behavior of the detected user is abnormal is determined by D(n) and λ. If D(n)≥λ, the current behavior of the detected user is considered normal. If D(n)<λ, the current behavior of the detected user is considered abnormal. The current behavior corresponds to w classification values With w+1 states (q n-w+1 ,q n-w+2 ,…,q n ,q n+1 ), the threshold λ is a very sensitive parameter. The higher the threshold, the higher the anomaly detection success rate, and the higher the false alarm rate. The lower the threshold, the lower the anomaly detection success rate, and the lower the false alarm rate.