AI-based database operating system risk prevention and control system and method
Through the AI-based database operating system risk prevention and control system, combined with natural language processing and deep learning technology, dynamically identifying and responding to database operation risks, the limitations of traditional defense mechanisms and the shortcomings of backup strategies are solved, and efficient risk prevention and control and data management are achieved.
Patent Information
- Application Number
- CN202510610981.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
AI Technical Summary
When the existing technology faces complex and changing database operation scenarios, it is difficult to fully identify and prevent and control internal threats and advanced persistent threats. Traditional defense mechanisms lack dynamic adaptability, insufficient accuracy of risk analysis, and backup and recovery strategies cannot meet real-time requirements, resulting in data security and business continuity being threatened.
The AI-based database operating system risk prevention and control system is adopted, including data processing module, risk identification module, risk response module and user analysis module. Natural language processing, machine learning and deep learning technology are used to achieve comprehensive risk identification and automated response through grammatical semantic analysis, risk feature construction and dynamic threshold adjustment.
It significantly improves the accuracy and comprehensiveness of risk identification, dynamically adjusts risk thresholds, realizes automated risk response, enhances active defense capabilities, and improves data management efficiency.
Smart Images

Figure CN120256258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically to a risk prevention and control system and method for a database operating system based on AI. Background Art
[0002] With the rapid development of information technology, database operations have become increasingly frequent and complex, covering various operations such as data addition, deletion, modification, query, table structure modification, and permission adjustment. These operations have a significant impact on data security and business continuity. However, existing technologies have the following deficiencies when dealing with complex and changing attack environments and high-frequency update scenarios:
[0003] Limitations of traditional security protection means: Existing database security protection mainly relies on static defense mechanisms such as firewalls, access control lists (ACLs), and rule-based detection systems. Although these mechanisms can prevent external attacks to a certain extent, their defense capabilities against internal threats and complex attack behaviors are limited. For example, incorrect operations or malicious behaviors by internal personnel may lead to data leakage, data corruption, or system crashes, and traditional defense mechanisms are difficult to effectively identify and prevent such behaviors.
[0004] Challenges of backup and recovery: Although conventional backup methods ensure data security to a certain extent, in high-frequency update scenarios, traditional backup strategies cannot meet the requirements of real-time or near-real-time backup. There may be a lag in backup data, resulting in the inability to quickly recover to the latest state in case of a failure. In addition, traditional backup and recovery processes often consume a large amount of storage and computing resources, especially when dealing with large-scale data operations, which is likely to cause resource waste and affect the overall performance of the system.
[0005] Deficiencies of database firewalls: There are obvious deficiencies in the implementation of existing database firewalls. For example, although the whitelist mechanism can prevent the execution of unauthorized SQL statements, it cannot guarantee the recording of all SQL statements involved in complex business systems within a certain period of time. For complex queries containing a large number of attribute conditions, it is difficult for the whitelist to enumerate all cases. Direct interception may lead to business interruption, while non-interception poses a security risk. In addition, traditional database firewalls are insufficient in defending against advanced persistent threats (APTs) and zero-day attacks.
[0006] Limitations of risk analysis: In the comprehensive risk analysis of user operation behaviors, existing technologies do not fully consider risk factors such as the user's permission level, operation time, and location. This results in insufficient accuracy and comprehensiveness of risk analysis and makes it difficult to effectively identify potential high-risk operations. For example, some operations may be regarded as normal behaviors during normal working hours, but may carry higher risks when performed outside of working hours or by high-privilege users.
[0007] Lack of dynamic adaptability: Most existing risk prevention and control systems are based on static rules and preset thresholds, lacking dynamic adaptability. When facing constantly changing attack patterns and business requirements, these systems are difficult to adjust risk thresholds and response strategies in real time, resulting in high false alarm rates and missed alarm rates.
[0008] How to comprehensively identify and prevent database operation risks and improve data management efficiency is a technical problem that needs to be solved. Summary of the Invention
[0009] The technical task of the present invention is to address the above deficiencies and provide an AI-based risk prevention and control system and method for database operating systems to solve the technical problems of how to comprehensively identify and prevent database operation risks and improve data management efficiency.
[0010] In a first aspect, an AI-based risk prevention and control system for a database operating system of the present invention includes a data processing module, a risk identification module, a risk response module, and a user analysis module;
[0011] The data processing module is used to obtain historical operation logs from a database management system and perform data preprocessing on the historical operation logs to obtain processed operation logs, where the historical operation logs include SQL statements, execution times, user IDs, and operation results;
[0012] The risk identification module is used to perform syntactic and semantic analysis on the SQL statements in the processed operation logs based on natural language processing technology, construct risk characteristics of the SQL statements, and use the risk characteristics as inputs to perform risk prediction through a trained risk identification model, and output a risk level;
[0013] The risk response module is used to perform risk warnings based on predefined risk strategies and the risk levels output by the risk identification module, classify and prioritize the risk warnings based on a Bayesian network, and push the risk warnings;
[0014] The user analysis module is used to analyze users based on the processed operation logs and the risk levels output by the risk identification module, and real-time monitor user behaviors based on user portraits and anomaly detection algorithms, and form and push alarm information based on the detection results.
[0015] Preferably, the data preprocessing module is used to perform the following data preprocessing on the historical operation logs:
[0016] Perform data cleaning on the historical operation logs, remove invalid and incorrect data, and fill in missing values;
[0017] Perform data format conversion on the cleaned historical operation logs to obtain normalized historical operation logs;
[0018] Extract features from the normalized historical operation logs, and extract the feature vectors reflecting operation risks in the historical operation logs. Among them, the feature vectors reflecting operation risks include operation types, operation frequencies, the number of tables involved, and operation time distributions. The operation types include addition, deletion, modification, and query.
[0019] Preferably, the risk identification module is used to perform the following operations:
[0020] Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis;
[0021] Combine the results of lexical analysis to perform syntax analysis on the SQL statements in the processed operation logs, construct a syntax tree based on the results of syntax analysis, establish operation modes of multi-table joins and nested queries through the syntax tree, and obtain syntax structure features;
[0022] Perform semantic analysis on the processed operation logs based on the results of lexical analysis, syntax analysis, and business logic rules, construct risk features of SQL statements based on the results of semantic analysis and syntax structure features, and perform vectorization operations on the risk features of SQL statements to obtain vectorized risk features;
[0023] Use the vectorized risk features as input, perform risk prediction through the trained risk identification model, and output risk level labels. Among them, the risk identification model is a deep learning model constructed based on a long short-term memory network or a Transformer architecture.
[0024] Preferably, for the risk identification model, the trained risk identification model is obtained through the following operations: Use the vectorized risk features corresponding to the historical operation logs as samples, perform model training, model testing, and model verification on the constructed risk identification model based on the samples, and obtain the finally trained risk identification model.
[0025] Among them, when performing model training, the random forest or gradient boosting tree method in supervised learning algorithms is selected for model training. When performing model testing and model verification, the model parameters are optimized through cross-validation. When performing model verification, the trained risk identification model is verified with accuracy, recall rate, and F1 score as performance indicators.
[0026] Preferably, the user analysis module is used to perform the following operations:
[0027] Analyze users through the K-means clustering algorithm based on the processed operation logs and the risk levels output by the risk identification module, and generate multi-dimensional user portrait features including common operation types, operation time distributions, operation frequencies, and preferences for accessing sensitive data.
[0028] Based on user portrait features, analyze abnormal operations that deviate from normal behavior patterns through the Isolation Forest algorithm, and introduce an encoder to monitor user behavior in real time and identify users' abnormal operations to obtain user risk behaviors;
[0029] Based on user risk behaviors, give early warnings, generate and push warning messages, where the warning messages include the types of abnormal operations, the scope of involved data, and the risk levels.
[0030] Second Party, a risk prevention and control method for a database operating system based on AI, performs risk prevention and control based on a risk prevention and control system for a database operating system based on AI as described in any item of the first aspect. The method includes the following steps:
[0031] Data processing: Obtain historical operation logs from the database management system, and perform data preprocessing on the historical operation logs to obtain processed operation logs. Among them, the historical operation logs include SQL statements, execution times, user IDs, and operation results;
[0032] Risk identification: Based on natural language processing technology, perform syntactic and semantic analysis on the SQL statements in the processed operation logs, construct risk features of the SQL statements, and use the risk features as input to perform risk prediction through a trained risk identification model, and output the risk levels;
[0033] Risk response: Based on predefined risk strategies and the risk levels obtained from risk identification, give risk warnings, classify and prioritize the risk warnings based on a Bayesian network, and push the risk warnings;
[0034] User analysis: Analyze users based on the processed operation logs and the risk levels obtained from risk identification, and based on user portraits and anomaly detection algorithms, monitor user behavior in real time, form warning messages based on the detection results, and push the warning messages.
[0035] Preferably, when performing data preprocessing, perform the following data preprocessing on the historical operation logs:
[0036] Perform data cleaning on the historical operation logs, remove invalid and error data, and fill in missing values;
[0037] Perform data format conversion on the cleaned historical operation logs to obtain normalized historical operation logs;
[0038] Perform feature extraction on the normalized historical operation logs, extract feature vectors reflecting operation risks in the historical operation logs. Among them, the feature vectors reflecting operation risks include operation types, operation frequencies, the number of involved tables, and operation time distributions. The operation types include addition, deletion, modification, and search.
[0039] Preferably, risk identification includes the following operations:
[0040] Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis;
[0041] Perform syntax analysis on the SQL statements in the processed operation logs in combination with the results of lexical analysis, construct a syntax tree based on the results of syntax analysis, establish an operation mode of multi-table connection and nested query through the syntax tree, and obtain syntax structure features;
[0042] Perform semantic analysis on the processed operation logs based on the results of lexical analysis, syntax analysis, and business logic rules, construct risk features of the SQL statements based on the results of semantic analysis and syntax structure features, and perform vectorization operations on the risk features of the SQL statements to obtain vectorized risk features;
[0043] Use the vectorized risk features as input, perform risk prediction through the trained risk identification model, and output risk level labels. Among them, the risk identification model is a deep learning model constructed based on the long short-term memory network or Transformer architecture.
[0044] Preferably, for the risk identification model, the trained risk identification model is obtained through the following operations: Use the vectorized risk features corresponding to the historical operation logs as samples, perform model training, model testing, and model verification on the constructed risk identification model based on the samples to obtain the finally trained risk identification model.
[0045] Among them, during model training, the random forest or gradient boosting tree method in the supervised learning algorithm is selected for model training. During model testing and model verification, the model parameters are optimized through cross-validation. During model verification, the trained risk identification model is verified using accuracy, recall rate, and F1 score as performance indicators.
[0046] Preferably, user analysis includes the following operations:
[0047] Based on the processed operation logs and the risk levels obtained from risk identification, analyze users through the K-means clustering algorithm, and generate multi-dimensional user portrait features including common operation types, operation time distribution, operation frequency, and preferences for accessing sensitive data;
[0048] Based on the user portrait features, analyze abnormal operations that deviate from the normal behavior pattern through the isolation forest algorithm, and introduce an encoder to monitor user behavior in real time and identify the abnormal operations of users to obtain user risk behaviors;
[0049] Based on the user's risk behavior, alarms and early warnings are generated and alarm information is pushed. The alarm information includes the type of abnormal operation, the scope of involved data, and the risk level.
[0050] The risk prevention and control system and method for database operating system based on AI of the present invention have the following advantages:
[0051] 1. Comprehensively identify high-risk operations: By combining natural language processing, machine learning, and deep learning technologies, various database operation risks can be comprehensively identified, significantly improving the accuracy and comprehensiveness of risk identification;
[0052] 2. Dynamically adjust the risk threshold: Introduce online learning and reinforcement learning algorithms to dynamically adjust the risk threshold according to real-time data and system status, improving the adaptability and flexibility of the system;
[0053] 3. Automatically respond to risks: Based on alarms, automatic risk handling is realized, reducing manual intervention and improving operation and maintenance efficiency.
[0054] 4. User behavior analysis and early warning: Through user profiling and anomaly detection algorithms, user behavior is monitored in real time, and potential risks are early warned, enhancing the active defense ability. Brief Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] The present invention will be further described below in conjunction with the drawings.
[0057] Figure 1 It is a flowchart of a risk prevention and control method for a database operating system based on AI in Embodiment 2. Detailed Embodiments
[0058] The following will further illustrate the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the specific embodiments cited are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0059] The embodiments of the present invention provide a risk prevention and control system and method for a database operating system based on AI, which are used to solve the technical problems of how to comprehensively identify and prevent database operation risks and improve data management efficiency.
[0060] Embodiment 1:
[0061] The present invention discloses an AI-based database operating system risk prevention and control system, comprising a data processing module, a risk identification module, a risk response module and a user analysis module.
[0062] The data processing module is used to obtain historical operation logs from the database management system and perform data preprocessing on the historical operation logs to obtain processed operation logs, wherein the historical operation logs include SQL statements, execution time, user ID and operation results.
[0063] As a specific implementation of the data preprocessing module, the module is used to perform the following data preprocessing on the historical operation logs:
[0064] (1) Clean the historical operation logs, remove invalid and erroneous data, and fill in missing values;
[0065] (2) converting the data format of the cleaned historical operation log to obtain a normalized historical operation log;
[0066] (3) Perform feature extraction on the normalized historical operation logs to extract feature vectors reflecting the operation risks in the historical operation logs. The feature vectors reflecting the operation risks include the operation type, operation frequency, number of tables involved, and operation time distribution. The operation types include adding, deleting, modifying, and searching.
[0067] In this embodiment, data collection, data cleaning and normalization, and feature extraction operations are performed through this module. Data collection: comprehensively collect historical operation logs from the database management system, and record key information such as SQL statements, execution time, user ID, operation results, etc. in detail to accumulate original data materials for subsequent analysis; data cleaning and normalization: eliminate invalid and erroneous parts of the collected data, fill in missing values, ensure data consistency and integrity, and uniformly convert data formats to meet the standard requirements of subsequent processing; feature extraction: for the cleaned and normalized data, extract feature vectors that can reflect operational risks, covering dimensions such as operation type (such as addition, deletion, modification, and query), operation frequency, number of tables involved, and operation time distribution. These feature vectors will serve as key inputs to the subsequent risk assessment model to accurately depict the risk profile of the operational behavior.
[0068] The risk identification module is used to perform grammatical and semantic analysis on the SQL statements in the processed operation logs based on natural speech processing technology, construct risk features of the SQL statements, and use the risk features as input to perform risk prediction through the trained risk identification model to output the risk level.
[0069] As a specific implementation of the risk identification module, the module is used to perform the following operations:
[0070] (1) Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis;
[0071] (2) Perform syntax analysis on the SQL statements in the processed operation logs in combination with the results of lexical analysis, construct a syntax tree based on the results of syntax analysis, establish operation modes for multi-table joins and nested queries through the syntax tree, and obtain syntax structure features;
[0072] (3) Perform semantic analysis on the processed operation logs based on the results of lexical analysis, syntax analysis, and business logic rules, construct risk features of the SQL statements based on the results of semantic analysis and syntax structure features, and perform vectorization operations on the risk features of the SQL statements to obtain vectorized risk features;
[0073] (4) Use the vectorized risk features as input, perform risk prediction through the trained risk identification model, and output risk level labels. Among them, the risk identification model is a deep learning model constructed based on the long short-term memory network or Transformer architecture.
[0074] In this embodiment, this module performs operations such as SQL statement parsing and feature fusion, machine learning model training and evaluation, deep learning model training and optimization, and dynamic adjustment of risk thresholds.
[0075] SQL statement parsing and feature fusion involve lexical and syntax analysis as well as semantic analysis. During lexical and syntax analysis, natural language processing technology is used to deeply parse SQL statements. First, lexical analysis is performed to accurately identify basic elements such as keywords, table names, and column names in SQL statements; then syntax analysis is carried out to construct a syntax tree and analyze complex operation modes such as multi-table joins and nested queries. The syntax structure features extracted in this step can reflect the complexity and normativity of SQL statements, providing an important basis for evaluating potential operation risks. For example, multi-table join operations may cause performance risks due to the large amount of data involved, while nested queries may lead to data consistency risks due to complex logic. Semantic analysis: Based on lexical and syntax analysis, combined with business logic rules, semantic interpretation of SQL statements is carried out. Judge whether an UPDATE statement involves key business fields, whether a SELECT statement queries sensitive privacy data, etc., and accurately identify the compliance of operations with business rules. The results of semantic analysis will jointly form a comprehensive risk feature portrait of SQL statements with the syntax structure features extracted earlier, providing detailed risk information for subsequent model evaluation. For example, if an UPDATE statement modifies the key amount field of the core financial data table and there is no approval process associated during off-peak business hours, its risk score will increase significantly.
[0076] Machine learning model training and evaluation involves data preparation and feature engineering as well as model training and tuning. Data preparation and feature engineering: Based on the extracted feature vectors, the data is refined and various features are converted into a format acceptable to the machine learning model, such as one-hot encoding of operation types and normalization of operation frequencies, etc., to build a complete training data set and lay a solid foundation for model learning. Model training and tuning: Random forests, gradient boosting trees and other methods in supervised learning algorithms are used to train risk identification models. With feature vectors as input and the risk level of the corresponding operation (labeled as high risk, medium risk, and low risk) as output, the model learns the intrinsic mapping relationship between features and risks. Through cross-validation and other means, the model parameters are continuously optimized to improve the model's performance indicators such as accuracy, recall rate, and F1 score in risk prediction tasks, to ensure that the model can accurately locate high-risk operations while avoiding excessive false positives that interfere with normal business processes.
[0077] Deep learning model training and optimization involve data vectorization and model architecture construction, model training and evaluation improvement. Data vectorization and model architecture construction: With the help of word embedding technology (such as Word2Vec or BERT), SQL statements are converted into vector representations to fully capture the semantic and contextual information. Long short-term memory network (LSTM) or Transformer architecture is used to build deep learning models, and its powerful time series feature capture ability is used to mine the potential risk evolution pattern in SQL statement sequences. Model training and evaluation improvement: With the embedded vector sequence as input and the corresponding risk label as output, the model parameters are continuously adjusted through back propagation and gradient descent algorithms, the binary cross entropy loss function is optimized, and the performance of the model in risk identification tasks is improved. The model is rigorously evaluated using the validation set, and the optimal deep learning model version is selected based on indicators such as accuracy, recall rate and F1 score, so that it can accurately identify various risk operations in complex and changeable database operation scenarios, providing deep intelligent support for risk prevention and control.
[0078] Dynamically adjusting risk thresholds involves embedding online learning mechanisms and optimizing reinforcement learning mechanisms.
[0079] The embedding of the online learning mechanism involves the application of incremental learning algorithms and the associated adjustment of the system state. Application of incremental learning algorithms: Introduce incremental learning algorithms such as online stochastic gradient descent, enabling the model to accept the latest operation data in real time and dynamically update the model parameters. With the changes in business scenarios and the continuous accumulation of data, the model adjusts the benchmark of risk assessment in real time, always accurately matching the current risk situation. For example, during the business peak period, if a large number of high-frequency but low-risk query operations flood in, the model will moderately adjust the threshold according to the online learning mechanism to avoid excessive alarms caused by a fixed threshold and ensure the smoothness of the business. Associated adjustment of the system state: Comprehensively consider key indicators such as the load status and data update frequency during system operation, as well as the current business requirements, and flexibly and dynamically adjust the risk threshold. In high-load scenarios, to balance business continuity and risk prevention and control, appropriately increase the risk threshold to reduce the false alarm rate; while in the data sensitive period or the business off-peak period, lower the threshold to enhance the sensitivity of risk monitoring and promptly capture potential threats.
[0080] Building the reinforcement learning environment and designing the reward function: Carefully design the reinforcement learning environment, map the risk threshold adjustment strategy to the action space of the intelligent agent, and construct the reward function based on core indicators such as the false alarm rate and miss rate feedback by the system. When the system accurately identifies high-risk operations without false alarms, give the intelligent agent a positive reward; conversely, if false alarms or miss alarms occur, impose a negative penalty. The intelligent agent continuously learns and optimizes the threshold adjustment strategy through interaction and trial and error with the environment. After multiple iterations, a dynamic threshold adjustment scheme that can perfectly balance the risk identification accuracy and business smoothness is derived, enabling the system to have an adaptive risk prevention and control ability in complex and dynamic business scenarios.
[0081] The risk response module is used to perform risk alerts based on predefined risk strategies and the risk levels output by the risk identification module, classify and prioritize the risk alerts based on the Bayesian network, and push the risk alerts.
[0082] For the risk identification model, the trained risk identification model is obtained through the following operations: Use the vectorized risk features corresponding to the historical operation logs as samples, and perform model training, model testing, and model verification on the constructed risk identification model based on the samples to obtain the final trained risk identification model.
[0083] Among them, when training the model, select the random forest or gradient boosting tree method in the supervised learning algorithm for model training. When testing the model, optimize the model parameters through cross-validation. When validating the model, use accuracy, recall rate, and F1 score as performance indicators to validate the trained risk identification model.
[0084] This implementation realizes the construction and optimization of the AI decision-making engine through this module, involving the refinement and dynamic adjustment of the rule engine, the upgrade of the intelligent alarm system, and the feedback integration.
[0085] Refinement and dynamic adjustment of the rule engine: Based on the risk assessment results, carefully design a multi-level response strategy. For low-risk operations (risk score <0.3), only the operation log is recorded to ensure the normal progress of the business; for medium-risk operations (0.3≤risk score<0.7), the administrator is notified to intervene in the review, and detailed operation information is recorded for traceability; for high-risk operations (risk score ≥0.7), the operation execution is decisively and automatically interrupted, and complete information is recorded simultaneously and an emergency alarm is triggered. Machine learning algorithms such as Bayesian networks are introduced to continuously monitor and evaluate the effectiveness of response strategies. According to changes in business scenarios and historical response effect data, the rule engine is dynamically optimized and adjusted to ensure that it always meets the current business risk prevention and control needs, and achieves accurate and efficient risk response.
[0086] Intelligent alarm system upgrade and feedback integration: Use Bayesian networks to intelligently classify alarm information, and accurately prioritize it according to the urgency of risks and the potential impact range, and give priority to key high-risk alarms to help administrators focus on core risks. At the same time, establish an administrator feedback loop to collect administrators' opinions and evaluations on the alarm results. Based on this, the system deeply analyzes the deviations and deficiencies of the alarm strategy, optimizes the alarm model parameters and classification logic in a targeted manner, continuously improves the accuracy and practicality of alarms, and builds an efficient closed-loop risk alarm handling mechanism.
[0087] The user analysis module is used to analyze users based on the processed operation logs and the risk level output by the risk identification module, and to monitor user behavior in real time based on user portraits and anomaly detection algorithms, and to generate and push alarm information based on the detection results.
[0088] As a specific implementation of the user analysis module, this module is used to perform the following operations:
[0089] (1) Based on the processed operation logs and the risk level output by the risk identification module, the user is analyzed through the K-means clustering algorithm to generate multi-dimensional user profile features including common operation types, operation time distribution, operation frequency, and preference for sensitive data access;
[0090] (2) Based on user portrait features, the isolation forest algorithm is used to analyze abnormal operations that deviate from normal behavior patterns, and an encoder is introduced to monitor user behavior in real time and identify users' abnormal operations to obtain user risk behaviors;
[0091] (3) Generate and push warnings based on user risk factors. Warning information includes the type of abnormal operation, the scope of data involved, and the risk level.
[0092] In this embodiment, this module realizes user behavior analysis and risk warning, specifically involving user portrait construction and deepening, anomaly detection algorithm enhancement, and real-time monitoring and warning coordination.
[0093] User portrait construction and deepening: Comprehensively collect the user's historical operation logs. After data preprocessing, use data analysis methods such as the K-means clustering algorithm to deeply analyze the user's operation behavior patterns, and generate a user portrait containing multi-dimensional information such as common operation types, operation time distribution, operation frequency, and access preferences for sensitive data, laying a solid foundation for accurately identifying abnormal behaviors.
[0094] Anomaly detection algorithm enhancement: On the one hand, use the Isolation Forest algorithm to accurately locate abnormal operations that deviate significantly from the normal behavior pattern based on the user portrait features, and quickly capture potential risks; on the other hand, introduce an autoencoder to continuously monitor the subtle changes in the user behavior sequence, and rely on its efficient learning and reconstruction ability for the normal behavior pattern to sensitively identify even hidden and complex abnormal operations, realizing all-round monitoring of various risk behaviors.
[0095] Real-time monitoring and warning coordination: Rely on a high-performance stream processing framework (such as Apache Kafka, Apache Flink) to perform real-time millisecond-level processing on user operation data, seamlessly connect the user portrait and the anomaly detection model, and immediately evaluate the operation risk. Once an abnormal behavior is detected and triggers the warning threshold, the system immediately automatically activates the warning process, generates detailed warning information, including core contents such as abnormal operation types, data ranges involved, and risk levels, and quickly notifies the administrator through multiple channels (such as emails, text messages, in-site messages) to ensure that potential risks are detected and handled at the budding stage, comprehensively protecting the security of database operations.
[0096] The system of this embodiment comprehensively identifies and prevents high-risk risks in database operations through machine learning and natural language processing technologies. The system parses SQL statements through natural language processing technology to extract semantic information; combines machine learning and deep learning models for risk prediction and anomaly detection; and introduces online learning and reinforcement learning algorithms to dynamically adjust the risk threshold. In addition, the system also has an automated risk response ability, which can trigger predefined response strategies according to the risk assessment results and early warn of potential risks through user behavior analysis.
[0097] Embodiment 2:
[0098] A method for preventing and controlling database operation system risks based on AI according to the present invention performs risk prevention and control based on the system disclosed in Embodiment 1. This method includes data processing, risk identification, risk response, and user analysis.
[0099] Step S100 data processing: obtaining historical operation logs from the database management system, and performing data preprocessing on the historical operation logs to obtain processed operation logs, wherein the historical operation logs include SQL statements, execution time, user ID, and operation results.
[0100] As a specific implementation of step S100, this step performs data preprocessing on the historical operation log as follows:
[0101] (1) Clean the historical operation logs, remove invalid and erroneous data, and fill in missing values;
[0102] (2) converting the data format of the cleaned historical operation log to obtain a normalized historical operation log;
[0103] (3) Perform feature extraction on the normalized historical operation logs to extract feature vectors reflecting the operation risks in the historical operation logs. The feature vectors reflecting the operation risks include the operation type, operation frequency, number of tables involved, and operation time distribution. The operation types include adding, deleting, modifying, and searching.
[0104] In this embodiment, data collection, data cleaning and normalization, and feature extraction operations are performed through the steps: data collection: comprehensively collect historical operation logs from the database management system, and record key information such as SQL statements, execution time, user ID, operation results, etc. in detail to accumulate original data materials for subsequent analysis; data cleaning and normalization: eliminate invalid and erroneous parts of the collected data, fill in missing values, ensure data consistency and integrity, and uniformly convert data formats to meet the standard requirements of subsequent processing; feature extraction: for the cleaned and normalized data, extract feature vectors that can reflect operational risks, covering dimensions such as operation type (such as addition, deletion, modification, and query), operation frequency, number of tables involved, and operation time distribution. These feature vectors will serve as key inputs to the subsequent risk assessment model to accurately depict the risk profile of the operational behavior.
[0105] Step S200: Risk identification: Based on natural speech processing technology, grammatical and semantic analysis is performed on the SQL statements in the processed operation logs, risk features of the SQL statements are constructed, and risk features are used as input to perform risk prediction through the trained risk identification model, and the risk level is output.
[0106] As a specific implementation of risk identification, this step performs the following operations:
[0107] (1) Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis;
[0108] (2) Perform grammatical analysis on the SQL statements in the processed operation logs in combination with the lexical analysis results, build a syntax tree based on the syntax analysis results, establish multi-table connection and nested query operation modes through the syntax tree, and obtain grammatical structure features;
[0109] (3) Perform semantic analysis on the processed operation log based on the lexical analysis results, grammatical analysis results, and business logic rules, construct risk features of SQL statements based on the semantic analysis results and grammatical structure features, and perform vectorization operations on the risk features of SQL statements to obtain vectorized risk features;
[0110] (4) Using the vectorized risk features as input, risk prediction is performed through the trained risk identification model, and the risk level label is output. The risk identification model is a deep learning model built based on the long short-term memory network or the Transformer architecture.
[0111] This embodiment uses this step to perform operations such as SQL statement parsing and feature fusion, machine learning model training and evaluation, deep learning model training and optimization, and dynamic adjustment of risk thresholds.
[0112] SQL statement parsing and feature fusion involve lexical and grammatical analysis as well as semantic analysis. During lexical and grammatical analysis, natural language processing technology is used to deeply parse SQL statements. First, lexical analysis is performed to accurately identify basic elements such as keywords, table names, and column names in SQL statements; then grammatical analysis is performed to build a syntax tree and analyze complex operation modes such as multi-table joins and nested queries. The grammatical structure features extracted in this step can reflect the complexity and standardization of SQL statements, providing an important basis for assessing potential operational risks. For example, multi-table join operations may cause performance risks due to the large amount of data involved, and nested queries may cause data consistency risks due to complex logic. Semantic analysis: Based on lexical and grammatical analysis, SQL statements are semantically interpreted in combination with business logic rules. For example, it is determined whether the UPDATE statement involves key business fields, whether the SELECT statement queries sensitive privacy data, etc., to accurately identify the degree of fit between the operation and the business rules. The semantic analysis results will form a comprehensive risk feature portrait of the SQL statement together with the grammatical structure features extracted previously, providing detailed risk information for subsequent model evaluation. For example, if an UPDATE statement modifies a key amount field in a core financial data table, its risk score will increase significantly during non-peak business hours and without an approval process associated with it.
[0113] Machine learning model training and evaluation involves data preparation and feature engineering as well as model training and tuning. Data preparation and feature engineering: Based on the extracted feature vectors, the data is refined and various features are converted into a format acceptable to the machine learning model, such as one-hot encoding of operation types and normalization of operation frequencies, etc., to build a complete training data set and lay a solid foundation for model learning. Model training and tuning: Random forests, gradient boosting trees and other methods in supervised learning algorithms are used to train risk identification models. With feature vectors as input and the risk level of the corresponding operation (labeled as high risk, medium risk, and low risk) as output, the model learns the intrinsic mapping relationship between features and risks. Through cross-validation and other means, the model parameters are continuously optimized to improve the model's performance indicators such as accuracy, recall rate, and F1 score in risk prediction tasks, to ensure that the model can accurately locate high-risk operations while avoiding excessive false positives that interfere with normal business processes.
[0114] Deep learning model training and optimization involve data vectorization and model architecture construction, model training and evaluation improvement. Data vectorization and model architecture construction: With the help of word embedding technology (such as Word2Vec or BERT), SQL statements are converted into vector representations to fully capture the semantic and contextual information. Long short-term memory network (LSTM) or Transformer architecture is used to build deep learning models, and its powerful time series feature capture ability is used to mine the potential risk evolution pattern in SQL statement sequences. Model training and evaluation improvement: With the embedded vector sequence as input and the corresponding risk label as output, the model parameters are continuously adjusted through back propagation and gradient descent algorithms, the binary cross entropy loss function is optimized, and the performance of the model in risk identification tasks is improved. The model is rigorously evaluated using the validation set, and the optimal deep learning model version is selected based on indicators such as accuracy, recall rate and F1 score, so that it can accurately identify various risk operations in complex and changeable database operation scenarios, providing deep intelligent support for risk prevention and control.
[0115] Dynamically adjusting risk thresholds involves embedding online learning mechanisms and optimizing reinforcement learning mechanisms.
[0116] The embedding of the online learning mechanism involves the application of incremental learning algorithms and the associated adjustment of system states. Application of incremental learning algorithms: Introduce incremental learning algorithms such as online stochastic gradient descent, enabling the model to accept the latest operation data in real time and dynamically update model parameters. As the business scenario changes and data continues to accumulate, the model adjusts the benchmark for risk assessment in real time, always accurately matching the current risk situation. For example, during the business peak period, if a large number of high-frequency but low-risk query operations flood in, the model will moderately adjust the threshold according to the online learning mechanism to avoid over-alarming caused by a fixed threshold and ensure the smoothness of the business. Associated adjustment of system states: Comprehensively consider key indicators such as the load status and data update frequency during system operation, as well as the current business requirements, and flexibly and dynamically adjust the risk threshold. In high-load scenarios, to balance business continuity and risk prevention and control, appropriately increase the risk threshold to reduce the false alarm rate; while in the data sensitive period or the business off-peak period, lower the threshold to enhance the sensitivity of risk monitoring and promptly capture potential threats.
[0117] Construction of the reinforcement learning environment and design of the reward function: Carefully design the reinforcement learning environment, map the risk threshold adjustment strategy to the action space of the agent, and construct the reward function based on core indicators such as the false alarm rate and missed alarm rate fed back by the system. When the system accurately identifies high-risk operations without false alarms, give the agent a positive reward; conversely, if false alarms or missed alarms occur, impose a negative penalty. The agent continuously learns and optimizes the threshold adjustment strategy through interaction and trial and error with the environment. After multiple iterations, a dynamic threshold adjustment scheme that can perfectly balance the risk identification accuracy and business smoothness is derived, enabling the system to have an adaptive risk prevention and control ability in complex and dynamic business scenarios.
[0118] Step S300 Risk response: Based on the predefined risk strategy and the risk level obtained from risk identification, conduct risk alerts, classify and prioritize the risk alerts based on the Bayesian network, and push the risk alerts.
[0119] For the risk identification model, the trained risk identification model is obtained through the following operations: Use the vectorized risk features corresponding to the historical operation logs as samples, and based on the samples, conduct model training, model testing, and model verification on the constructed risk identification model to obtain the final trained risk identification model.
[0120] Among them, when training the model, select the random forest or gradient boosting tree method in the supervised learning algorithm for model training. When testing the model, optimize the model parameters through cross-validation. When validating the model, use accuracy, recall rate, and F1 score as performance indicators to validate the trained risk identification model.
[0121] This implementation realizes the construction and optimization of the AI decision-making engine through this step, involving the refinement and dynamic adjustment of the rule engine, the upgrade and feedback integration of the intelligent alarm system.
[0122] Refinement and dynamic adjustment of the rule engine: Based on the risk assessment results, carefully design a multi-level response strategy. For low-risk operations (risk score <0.3), only the operation log is recorded to ensure the normal progress of the business; for medium-risk operations (0.3≤risk score<0.7), the administrator is notified to intervene in the review, and detailed operation information is recorded for traceability; for high-risk operations (risk score ≥0.7), the operation execution is decisively and automatically interrupted, and complete information is recorded simultaneously and an emergency alarm is triggered. Machine learning algorithms such as Bayesian networks are introduced to continuously monitor and evaluate the effectiveness of response strategies. According to changes in business scenarios and historical response effect data, the rule engine is dynamically optimized and adjusted to ensure that it always meets the current business risk prevention and control needs, and achieves accurate and efficient risk response.
[0123] Intelligent alarm system upgrade and feedback integration: Use Bayesian networks to intelligently classify alarm information, and accurately prioritize it according to the urgency of risks and the potential impact range, and give priority to key high-risk alarms to help administrators focus on core risks. At the same time, establish an administrator feedback loop to collect administrators' opinions and evaluations on the alarm results. Based on this, the system deeply analyzes the deviations and deficiencies of the alarm strategy, optimizes the alarm model parameters and classification logic in a targeted manner, continuously improves the accuracy and practicality of alarms, and builds an efficient closed-loop risk alarm handling mechanism.
[0124] Step S400: User analysis: Analyze users based on the processed operation logs and the risk level obtained by risk identification, monitor user behavior in real time based on user portraits and anomaly detection algorithms, generate and push alarm information based on the detection results.
[0125] As a specific implementation of user analysis, this step performs the following operations:
[0126] (1) Based on the processed operation logs and the risk level output by the risk identification module, the user is analyzed through the K-means clustering algorithm to generate multi-dimensional user profile features including common operation types, operation time distribution, operation frequency, and preference for sensitive data access;
[0127] (2) Based on user portrait features, the isolation forest algorithm is used to analyze abnormal operations that deviate from normal behavior patterns, and an encoder is introduced to monitor user behavior in real time and identify users' abnormal operations to obtain user risk behaviors;
[0128] (3) Generate and push warnings based on user risk factors. Warning information includes the type of abnormal operation, the scope of data involved, and the risk level.
[0129] In this embodiment, this step realizes user behavior analysis and risk early warning, specifically involving user portrait construction and deepening, anomaly detection algorithm enhancement, and real-time monitoring and early warning coordination.
[0130] User portrait construction and deepening: Comprehensively collect the user's historical operation logs. After data preprocessing, use data analysis methods such as the K-means clustering algorithm to deeply analyze the user's operation behavior patterns, and generate a user portrait containing multi-dimensional information such as common operation types, operation time distribution, operation frequency, and access preferences for sensitive data, laying a solid foundation for accurately identifying abnormal behaviors.
[0131] Anomaly detection algorithm enhancement: On the one hand, use the Isolation Forest algorithm to accurately locate abnormal operations that deviate significantly from the normal behavior pattern based on the user portrait features, and quickly capture potential risks; on the other hand, introduce an autoencoder to continuously monitor the subtle changes in the user's behavior sequence, and rely on its efficient learning and reconstruction ability for the normal behavior pattern to keenly identify even hidden and complex abnormal operations, realizing all-round monitoring of various risk behaviors.
[0132] Real-time monitoring and early warning coordination: Rely on a high-performance stream processing framework (such as Apache Kafka, Apache Flink) to perform real-time millisecond-level processing on user operation data, seamlessly connect the user portrait and the anomaly detection model, and immediately evaluate the operation risk. Once an abnormal behavior is detected and triggers the early warning threshold, the system immediately automatically activates the early warning process, generates detailed warning information, including core contents such as abnormal operation types, data ranges involved, and risk levels, and quickly notifies the administrator through multiple channels (such as emails, text messages, in-site messages) to ensure that potential risks are detected and handled at the budding stage, comprehensively protecting the security of database operations.
[0133] The above has introduced in detail the risk prevention and control system and method of the AI-based database operating system provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An AI-based risk prevention and control system for database operating systems, characterized in that, It includes a data processing module, a risk identification module, a risk response module, and a user analysis module; The data processing module is used to obtain historical operation logs from a database management system and perform data preprocessing on the historical operation logs to obtain processed operation logs. Among them, the historical operation logs include SQL statements, execution times, user IDs, and operation results; The risk identification module is used to perform syntactic and semantic analysis on the SQL statements in the processed operation logs based on natural language processing technology, construct risk characteristics of the SQL statements, and use the risk characteristics as input to perform risk prediction through a trained risk identification model, and output a risk level; The risk response module is used to perform risk warnings based on predefined risk strategies and the risk levels output by the risk identification module, classify and prioritize the risk warnings based on a Bayesian network, and push the risk warnings; The user analysis module is used to analyze users based on the processed operation logs and the risk levels output by the risk identification module, and monitor user behaviors in real time based on user portraits and anomaly detection algorithms, and form warning information based on the detection results and push the warning information.
2. The risk prevention and control system for the AI-based database operating system according to claim 1, characterized in that, The data preprocessing module is used to perform the following data preprocessing on the historical operation logs: Perform data cleaning on the historical operation logs, remove invalid and error data, and fill in missing values; Perform data format conversion on the cleaned historical operation logs to obtain normalized historical operation logs; Perform feature extraction on the normalized historical operation logs, extract feature vectors reflecting operation risks in the historical operation logs. Among them, the feature vectors reflecting operation risks include operation types, operation frequencies, the number of tables involved, and operation time distributions. The operation types include addition, deletion, modification, and query.
3. The risk prevention and control system for the database operating system based on AI according to claim 1, characterized in that, The risk identification module is used to perform the following operations: Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis; Combine the lexical analysis results to perform syntactic analysis on the SQL statements in the processed operation logs, construct a syntax tree based on the syntactic analysis results, and establish operation modes of multi-table joins and nested queries through the syntax tree to obtain syntactic structure characteristics; Perform semantic analysis on the processed operation logs based on the lexical analysis results, syntactic analysis results, and business logic rules, construct risk characteristics of the SQL statements based on the semantic analysis results and syntactic structure characteristics, and perform vectorization operations on the risk characteristics of the SQL statements to obtain vectorized risk characteristics; Use the vectorized risk characteristics as input to perform risk prediction through a trained risk identification model, and output a risk level label. Among them, the risk identification model is a deep learning model constructed based on a long short-term memory network or a Transformer architecture.
4. The risk prevention and control system for an AI-based database operating system according to claim 3, wherein, For the risk identification model, the trained risk identification model is obtained through the following operations: Use the vectorized risk characteristics corresponding to the historical operation logs as samples, perform model training, model testing, and model verification on the constructed risk identification model based on the samples to obtain the final trained risk identification model. Among them, when training the model, the random forest or gradient boosting tree method in the supervised learning algorithm is selected for model training. When validating the model and testing the model, the model parameters are optimized through cross-validation. When validating the model, the accuracy, recall rate, and F1 score are used as performance indicators to validate the trained risk identification model.
5. The risk prevention and control system for the AI-based database operating system according to claim 1, wherein The user analysis module is used to perform the following operations: Based on the processed operation logs and the risk levels output by the risk identification module, analyze users through the K-means clustering algorithm to generate multi-dimensional user portrait features including common operation types, operation time distribution, operation frequency, and sensitive data access preferences; Based on the user portrait features, analyze abnormal operations deviating from the normal behavior pattern through the isolation forest algorithm, and introduce an encoder to monitor user behavior in real time and identify the abnormal operations of users to obtain user risk behaviors; Based on the user risk behaviors, give early warnings, generate and push warning messages, and the warning messages include abnormal operation types, data ranges involved, and risk levels.
6. A risk prevention and control method for a database operating system based on AI, characterized in that, Based on the risk prevention and control system of an AI-based database operating system described in any one of claims 1-5, perform risk prevention and control, and the method includes the following steps: Data processing: Obtain historical operation logs from the database management system, and perform data preprocessing on the historical operation logs to obtain processed operation logs. Among them, the historical operation logs include SQL statements, execution times, user IDs, and operation results; Risk identification: Based on natural language processing technology, perform syntax and semantic analysis on the SQL statements in the processed operation logs, construct risk features of the SQL statements, and use the risk features as input to perform risk prediction through the trained risk identification model, and output the risk level; Risk response: Perform risk warnings based on predefined risk strategies and the risk levels obtained from risk identification, classify and prioritize the risk warnings based on the Bayesian network, and push the risk warnings; User analysis: Analyze users based on the processed operation logs and the risk levels obtained from risk identification, and monitor user behavior in real time based on the user portrait and anomaly detection algorithm, and form and push warning messages based on the detection results.
7. The risk prevention and control method for an AI-based database operating system according to claim 6, characterized in that, When performing data preprocessing, perform the following data preprocessing on the historical operation logs: Perform data cleaning on the historical operation logs, remove invalid and error data, and fill in missing values; Perform data format conversion on the cleaned historical operation logs to obtain normalized historical operation logs; Perform feature extraction on the normalized historical operation logs, and extract feature vectors reflecting operation risks in the historical operation logs. Among them, the feature vectors reflecting operation risks include operation types, operation frequencies, the number of tables involved, and operation time distribution, and the operation types include addition, deletion, modification, and search.
8. The risk prevention and control method for an AI-based database operating system according to claim 6, characterized in that Risk identification includes the following operations: Perform lexical analysis on the SQL statements in the processed operation logs, and identify keywords, table names, and column names in the SQL statements through lexical analysis; Perform syntax analysis on the SQL statements in the processed operation log in combination with the lexical analysis results, construct a syntax tree based on the syntax analysis results, establish the operation modes of multi-table joins and nested queries through the syntax tree, and obtain the syntax structure features; Perform semantic analysis on the processed operation log based on the lexical analysis results, syntax analysis results, and business logic rules, construct the risk features of the SQL statements based on the semantic analysis results and syntax structure features, and perform vectorization operations on the risk features of the SQL statements to obtain the vectorized risk features; Use the vectorized risk features as input, perform risk prediction through the trained risk recognition model, and output the risk level label. Among them, the risk recognition model is a deep learning model constructed based on the long short-term memory network or Transformer architecture.
9. The risk prevention and control method for an AI-based database operating system according to claim 8, wherein For the risk recognition model, the trained risk recognition model is obtained through the following operations: Use the vectorized risk features corresponding to the historical operation log as samples, perform model training, model testing, and model verification on the constructed risk recognition model based on the samples to obtain the finally trained risk recognition model. Among them, when performing model training, the random forest or gradient boosting tree method in the supervised learning algorithm is selected for model training. When performing model testing and model verification, the model parameters are optimized through cross-validation. When performing model verification, the accuracy, recall rate, and F1 score are used as performance indicators to verify the trained risk recognition model.
10. The risk prevention and control method for an AI-based database operating system according to claim 6, wherein User analysis includes the following operations: Based on the processed operation log and the risk level obtained from risk recognition, analyze users through the K-means clustering algorithm to generate multi-dimensional user portrait features including common operation types, operation time distribution, operation frequency, and sensitive data access preferences; Based on the user portrait features, analyze abnormal operations that deviate from the normal behavior pattern through the isolation forest algorithm, and introduce an encoder to monitor user behavior in real time and identify the abnormal operations of users to obtain user risk behaviors; Perform alarm and early warning based on the user risk behaviors, generate and push alarm information, and the alarm information includes the abnormal operation type, the data range involved, and the risk level.
Citation Information
Cited By
Multi-terminal cooperative communication monitoring alarm system for 5G new call
CN120957172A
Intelligent agent cooperation system and method based on hierarchical context and global bridging scheduling
CN121116528A
Database operation log risk early warning method and system
CN121722650A
Safety management risk detection method and system for power monitoring system
CN122119123A
Intrusion detection method, electronic device, storage medium, and computer program product
CN122548736A