Log audit monitoring method and log audit monitoring system
By deploying agents in the IT audit system to collect raw log data and using the DeepLog model for data analysis and exception monitoring, the problems of data tampering, high labeling costs and security vulnerabilities in IT audits are solved, and efficient and accurate log audit monitoring is achieved.
Patent Information
- Application Number
- CN202510359956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-25
Smart Images

Figure CN120197216A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of log audit monitoring, and in particular to a log audit monitoring method and a log audit monitoring system. Background Art
[0002] IT audit refers to information system audit, which can be specifically referred to as the evaluation, supervision and standardization inspection of the internal information system, network security, production and operation, data flow, electronic signature, etc. of the audit department. IT audit mainly includes inspections of system development, system application, production and operation, security and risk management, etc. It is generally necessary to audit and monitor the logs generated in the above links, so as to effectively supervise and deal with abnormal situations found in log audit.
[0003] There are several main problems in performing IT audits using traditional technologies:
[0004] First, the risk of data tampering. In traditional technologies, logs generated in the business production system are exported from the system or manually sorted before audit monitoring. This process may lead to the risk of log data being tampered with, making it impossible to truly reflect the problems of log production through audit operations.
[0005] Second, it takes a lot of labor costs or algorithm models to label samples, which is labor-intensive and inaccurate. Traditional technologies generally use supervised algorithms to audit and monitor logs, which means that a lot of labeling work needs to be done in the early stages, and the labeling cost and difficulty are high. In traditional practices, due to the limited number and quality of manual and model-labeled samples, the training data is too small or the samples are highly homogeneous, resulting in overfitting in model training. In addition, this supervised model is difficult to identify abnormal situations outside the training samples and has poor generalization ability.
[0006] Third, the limitations of the audit monitoring function lead to security vulnerabilities. In the existing technology, the monitoring operation and maintenance platform and the company's organizational structure personnel information system are often independent of each other. This results in the operation and maintenance platform being unable to obtain these changes in a timely manner when personnel information changes (such as resignation, retirement, and job transfer), thus creating security risks. For example, the existing audit monitoring model is generally only responsible for monitoring abnormal operations at the operation and maintenance level, and cannot provide real-time insight into the log operations of users with unauthorized permissions. Summary of the invention
[0007] Based on the above problems, the present application provides a log audit monitoring method and a log audit monitoring system, the purpose of which is to reduce the risk of log data tampering and the cost of manual labeling, accurately and effectively realize log audit monitoring, and more comprehensively identify user authority security vulnerabilities in the audit process.
[0008] The embodiments of the present application disclose the following technical solutions:
[0009] In a first aspect of the present application, a log audit monitoring method is provided. This method is applied to a log audit monitoring system; the log audit monitoring system has a real-time embedded end-to-end architecture, and the log audit monitoring system at least includes: agent programs deployed on multiple key log production nodes, and a log audit monitoring model that has both abnormal behavior monitoring functions and personnel authorization monitoring functions for logs; the log audit monitoring model is a deep learning model based on the DeepLog architecture.
[0010] The method includes:
[0011] Collect the original log data generated at the key log production nodes through the agent programs;
[0012] Perform data parsing on the original log data through the log audit monitoring model to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the log key sequence is a sequence formed by arranging the log keys of each log in the original log data in the order of the log timestamps; the basic parameter value vector sequence is a sequence formed by arranging the basic parameter value vectors of each log in the original log data in the order of the log timestamps.
[0013] Match the operation user information of each log in the original log data with the existing personnel information dataset through the log audit monitoring model to obtain the personnel information parameters of each log in the original log data;
[0014] Fuse the basic parameter value vector sequence and the personnel information parameters of each log through the log audit monitoring model to obtain a new parameter value vector sequence;
[0015] Construct a log template sequence corresponding to the original log data through the log audit monitoring model based on the log key sequence and the new parameter value vector sequence;
[0016] Analyze the log template sequence through the log audit monitoring model to obtain a first abnormal monitoring result and a second abnormal monitoring result; the first abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal permissions.
[0017] In an alternative implementation, the existing personnel information dataset includes multiple personnel information entries; each personnel information entry includes employee information, user information, and permission information; the employee information at least includes: employee status, and the employee status is one of on-the-job, separated, transferred, or retired; the permission information includes: user role permissions and authorization period information;
[0018] The log audit monitoring model matches each log's operating user information in the original log data with the existing personnel information dataset to obtain the personnel information parameters of each log in the original log data, including:
[0019] The log audit monitoring model matches, based on each log's operating user information in the original log data, whether there is a corresponding personnel information entry in the existing personnel information dataset.
[0020] If it is determined through matching from the existing personnel information dataset that there is a corresponding personnel information entry for the user information, then some or all of the information in this personnel information entry is extracted as the personnel information parameters of the corresponding log in the original log data; among them, the partial information at least includes the employee status and the permission information;
[0021] If it is determined through matching from the existing personnel information dataset that there is no corresponding personnel information entry for the user information, then the corresponding personnel information parameters in the original log data are set to null.
[0022] In an alternative implementation, the log audit monitoring model fuses the obtained new parameter value vector sequence based on the basic parameter value vector sequence and the personnel information parameters of each log, including:
[0023] The log audit monitoring model splices the basic parameter value vectors in the basic parameter value vector sequence and the obtained personnel information parameters of each log one by one based on each log to obtain a new parameter value vector for each log in the original log data;
[0024] The obtained new parameter value vectors of each log are arranged in the order of the log timestamps to form a new parameter value vector sequence.
[0025] In an alternative implementation, the log audit monitoring model analyzes the log template sequence to obtain a first anomaly monitoring result and a second anomaly monitoring result, including:
[0026] The log audit monitoring model extracts a log key sequence w = {m to be analyzed from the log template sequence based on a preset log key window length h t-h,…,m t-2 ,m t-1}, based on the log key sequence w = {m t-h ,…,m t-2 ,m t-1} to analyze the next log key m t to check if there is an anomaly, and obtain the log key anomaly analysis result of the log corresponding to the log key m t ;
[0027] The log audit monitoring model extracts new parameter value vectors corresponding to each log key from the log template sequence. Based on the new parameter value vectors corresponding to each log key, it predicts whether there is an anomaly in the parameter value of the corresponding next log, and obtains the parameter value anomaly analysis result of each log; the parameter value anomaly analysis result is used to indicate whether there is an anomaly in the basic parameter value and whether there is an anomaly in the personnel permission information;
[0028] If the log key anomaly analysis result of the same log indicates an anomaly and / or the parameter value anomaly analysis result indicates an anomaly in the basic parameter value, then obtain a first anomaly monitoring result indicating that the generation of this log is triggered by an abnormal behavior pattern; if the log key anomaly analysis result of the same log indicates no anomaly, and the parameter value anomaly analysis result indicates no anomaly in the basic parameter value, then obtain a first anomaly monitoring result indicating that the generation of this log is triggered by a normal behavior pattern;
[0029] If the parameter value anomaly analysis result indicates an anomaly in the personnel permission information, then obtain a second anomaly monitoring result indicating that the generation of this log is triggered by a person with abnormal permissions; if the parameter value anomaly analysis result indicates no anomaly in the personnel permission information, then obtain a second anomaly monitoring result indicating that the generation of this log is triggered by a person with normal permissions.
[0030] In an alternative implementation, based on the log key sequence w = {m t-h ,…,m t-2 ,m t-1} to analyze the next log key m t to check if there is an anomaly, and obtain the log key anomaly analysis result of the log corresponding to the log key m t , including:
[0031] Based on the log key sequence w = {m t-h ,…,m t-2 ,m t-1} to predict the conditional probability distribution of the next log key m t , and the conditional probability distribution is used to characterize the possibilities of multiple candidate situations of the next log key m t ;
[0032] If the actual value of the next log key m t is within the top g candidate cases in the conditional probability distribution, it is determined that the log key anomaly analysis result of this log is normal; if the actual value of the next log key m t is not within the top g candidate cases in the conditional probability distribution, it is determined that the log key anomaly analysis result of this log is abnormal.
[0033] In an alternative implementation, predicting whether there is an anomaly in the parameter value of the corresponding next log based on the new parameter value vector corresponding to each log key, and obtaining the parameter value anomaly analysis result of each log, includes:
[0034] Predicting the first prediction distribution and the second prediction distribution of the corresponding next log based on the new parameter value vector corresponding to each log key; the first prediction distribution is the prediction distribution for the basic parameter value for anomaly behavior monitoring, and the second prediction distribution is the prediction distribution for the personnel information parameter for personnel authorization monitoring;
[0035] Calculating the first confidence interval of the basic parameter value of the next log according to the first prediction distribution. If the actual value of the basic parameter value of the next log is within the first confidence interval, the parameter value anomaly analysis result of the next log indicates that the basic parameter value is normal; if the actual value of the basic parameter value of the next log is not within the first confidence interval, the parameter value anomaly analysis result of the next log indicates that the basic parameter value is abnormal;
[0036] Calculating the second confidence interval of the personnel information parameter of the next log according to the second prediction distribution. If the actual value of the personnel information parameter of the next log is within the second confidence interval, the parameter value anomaly analysis result of the next log indicates that the generation of the next log is triggered by a person with normal permissions; if the actual value of the personnel information parameter of the next log is not within the second confidence interval, the parameter value anomaly analysis result of the next log indicates that the generation of the next log is triggered by a person with abnormal permissions.
[0037] In an alternative implementation, the log audit monitoring model is pre-trained in an unsupervised learning manner through a sample data set of normal log template sequences;
[0038] The sample data set of the normal log template sequences includes multiple normal log template sequence samples;
[0039] Each of the normal log template sequence samples includes multiple log template samples of historical normal logs with a timestamp sequence relationship;
[0040] The normal log template sequence sample is constructed based on the log key sequence sample of the multiple historical normal logs with timestamp sequence relationships and the new parameter value vector sequence sample;
[0041] The new parameter value vector sequence sample is obtained by fusing the basic parameter value vectors of the multiple historical normal logs with timestamp sequence relationships and the personnel information parameters of each historical normal log.
[0042] In an optional implementation manner, the log audit monitoring system is further configured with an information display and feedback mechanism;
[0043] After analyzing the log template sequence through the log audit monitoring model to obtain the first abnormal monitoring result and the second abnormal monitoring result, the method further includes:
[0044] If the first abnormal monitoring result indicates an abnormality and / or the second abnormal monitoring result indicates an abnormality, a risk warning message is generated and pushed to the audit end;
[0045] If the risk misjudgment information fed back by the audit end based on the risk warning message is received, the log audit monitoring model is continuously learned and optimized using the risk misjudgment information.
[0046] In an optional implementation manner, the log audit monitoring system further includes: a data integration module and a data storage module; a Kafka stream database is configured in the data storage module; wherein, the Kafka stream database is selected based on the real-time requirement of the log audit monitoring system;
[0047] After collecting the original log data generated at the key log production nodes through the proxy program, the method further includes:
[0048] The data integration module uses ETL technology to automatically extract, clean, filter, transform, and standardize the log data from the original log data in the audit scenario application of the log audit monitoring system, and load the processed data into the Kafka stream database; the Kafka stream database provides the log data to be analyzed processed by the ETL technology to the log audit monitoring model.
[0049] The second aspect of the present application provides a log audit monitoring system, which has a real-time embedded end-to-end architecture, and the log audit monitoring system at least includes: a data collection module and a model prediction module;
[0050] The data collection module is implemented by agent programs deployed on multiple key log production nodes;
[0051] The model prediction module is implemented by a log audit monitoring model that has both abnormal behavior monitoring function and personnel authorization monitoring function for logs; the log audit monitoring model is a deep learning model based on the DeepLog architecture;
[0052] The data collection module is used to collect the original log data generated at the key log production nodes through the agent programs;
[0053] The model prediction module is used to perform data parsing on the original log data through the log audit monitoring model to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the log key sequence is a sequence formed by arranging the log keys of each log in the original log data in the order of the log timestamps; the basic parameter value vector sequence is a sequence formed by arranging the basic parameter value vectors of each log in the original log data in the order of the log timestamps;
[0054] The model prediction module is also used to match the operation user information of each log in the original log data with the existing personnel information dataset by the log audit monitoring model to obtain the personnel information parameters of each log in the original log data;
[0055] The model prediction module is also used to fuse the basic parameter value vector sequence and the personnel information parameters of each log by the log audit monitoring model to obtain a new parameter value vector sequence;
[0056] The model prediction module is also used to construct a log template sequence corresponding to the original log data by the log audit monitoring model based on the log key sequence and the new parameter value vector sequence;
[0057] The model prediction module is also used to analyze the log template sequence through the log audit monitoring model to obtain a first abnormal monitoring result and a second abnormal monitoring result; the first abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal permissions.
[0058] In an optional implementation, the system is further configured with a display and feedback module, and the display and feedback module is used to implement an information display and feedback mechanism; the display and feedback module is connected to the model prediction module;
[0059] The display and feedback module is specifically used for:
[0060] If the first anomaly monitoring result indicates an anomaly and / or the second anomaly monitoring result indicates an anomaly, receive the risk warning information generated by the model prediction module, and push the risk warning information to the audit end for display;
[0061] If receiving the risk misjudgment information fed back by the audit end based on the risk warning information, feed the risk misjudgment information back to the model prediction module, so that the model prediction module can use the risk misjudgment information to continue learning and optimizing the log audit monitoring model.
[0062] In an alternative implementation, the system further includes: a data integration module and a data storage module; a Kafka stream database is configured in the data storage module; wherein, the Kafka stream database is selected based on the real-time requirements of the log audit monitoring system;
[0063] The data integration module is used to extract, clean, filter, transform, and standardize log data automatically from the original log data in the audit scenario application of the log audit monitoring system by using ETL technology, and load the processed data into the Kafka stream database;
[0064] The data storage module is used to provide the log data to be analyzed processed by the ETL technology to the log audit monitoring model through the Kafka stream database.
[0065] Compared with the prior art, the present application has the following beneficial effects:
[0066] The log audit monitoring method provided by the technical solution of the present application is applied to a log audit monitoring system. The system has a real-time embedded end-to-end architecture, which at least includes agent programs deployed at multiple key log production nodes, and a log audit monitoring model with both anomaly behavior monitoring function and personnel authorization monitoring function for logs; the model is a deep learning model based on the DeepLog architecture. Since this method uses agent programs deployed at key log production nodes to collect buried point data and directly processes the collected data through the model, the link of exporting log data or manually sorting data is avoided. The collected data is used as the source data for model processing, reducing the risk of log data being tampered with in the audit operation, and making the audit monitoring results more authentic and effective. Since the model is a deep learning model based on the DeepLog architecture, it is trained in an unsupervised learning manner and can identify anomalies without a large amount of annotation work, greatly reducing the manual annotation cost. At the same time, it also avoids the problems of model overfitting or poor generalization ability caused by sample annotation, and further leads to poor accuracy of audit monitoring.
[0067] In the method, the model is used to perform data parsing on the original log data to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the model matches the operator user information of each log in the original log data with the existing personnel information data set to obtain the personnel information parameters of each log in the original log data; the model fuses the basic parameter value vector sequence and the personnel information parameters of each log to obtain a new parameter value vector sequence; the model constructs a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence; the model analyzes the log template sequence to obtain a first anomaly monitoring result and a second anomaly monitoring result. Among them, the first anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal permissions. Since in the technical solution of this application, the personnel information parameters of the log are used to realize data enhancement on the basis of the original basic parameter values, and a new parameter value vector sequence is fused and constructed, so as to assist the log audit monitoring model to also have the function of personnel authorization monitoring on the basis of having the abnormal behavior monitoring function. It can be seen that the log audit monitoring model adopted in the technical solution of this application breaks the independent audit monitoring barrier between the monitoring and operation and maintenance platform and the personnel information system, can detect the log operations of users with permission violations, and can more effectively audit and monitor relevant security risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0069] Figure 1 It is an architecture diagram of a log audit monitoring system provided by an embodiment of the present application;
[0070] Figure 2 It is a schematic flowchart of a log audit monitoring method provided by an embodiment of the present application;
[0071] Figure 3 It is an architecture diagram of realizing dual anomaly monitoring through a log audit monitoring model provided by an embodiment of the present application;
[0072] Figure 4 It is a DeepLog architecture diagram;
[0073] Figure 5 It is an example flowchart of realizing dual anomaly monitoring through a log audit detection model provided by an embodiment of the present application;
[0074] Figure 6 It is a flowchart of another log audit monitoring method provided by an embodiment of the present application;
[0075] Figure 7 It is a schematic structural diagram of a log audit monitoring system provided by an embodiment of the present application;
[0076] Figure 8 It is a schematic structural diagram of another log audit monitoring system provided by an embodiment of the present application. Detailed implementation manners
[0077] Before formally introducing the specific implementation of the technical solution of the present application, first, three current mainstream log audit monitoring technologies will be introduced.
[0078] Technology 1: An IT audit method and device. In the method of Technology 1, start the audit workflow corresponding to the software development system architecture; obtain the corresponding data from the software development system architecture according to the audit workflow; automatically audit the corresponding data according to the audit workflow; and output an audit report. Technology 1 uses the audit workflow to automatically process and analyze audit data, thus eliminating the cumbersome processes of manual extraction, collation, querying and retrieving materials, and data extraction.
[0079] Technology 2: A method, system, device and medium for fusion risk analysis of multi-source security logs. The method of Technology 2 includes: obtaining security logs from multiple data sources, performing structured parsing, machine learning annotation and standardization processing on the security logs to generate standardized logs; cleaning, de-duplicating, filtering and preprocessing the standardized logs; storing the standardized logs in a distributed database, performing asset identification based on the standardized logs, and using a deep reinforcement learning model to perform risk identification according to the asset identification results, identifying risk logs and processing them; feeding back the processing results of the risk logs to the deep reinforcement learning model for continuous training and adjusting the judgment conditions of the risk logs. It can automatically process security logs from multiple sources, and through technologies such as machine learning, achieve log formatting, cleaning, asset identification and risk identification, thereby improving the automation degree and defense ability of the data security platform.
[0080] Technology Three: A log auditing and monitoring method and device. The method of Technology Three includes: collecting host logs and performing parsing and processing; using feature extraction to process log data; using the LSTM algorithm to train and establish a decision model and continuously optimize it; automatically auditing host logs, judging the risk level and processing it. By using a deep learning algorithm to establish a decision model, records that do not conform to the expected behavior are found and automated processing is performed. Since the Spell method is used to parse the logs and the LSTM algorithm is used to train and establish a decision model, focusing on the processing and use of log features, the heterogeneity of the logs is masked, which is applicable to the anomaly detection of multiple types of system logs.
[0081] The inventors analyzed and found that there are some drawbacks to the above three technologies, which are summarized as follows:
[0082] Drawback One: Existing technologies generally rely on the audit data submitted by the audited personnel and cannot obtain the most original and authentic data from the source systems of production operation and maintenance, resulting in the risk of data tampering in the audit data. Taking Technology One mentioned above as an example, in Technology One, through the workflow method, the materials submitted by each process node in the software development system architecture are directly used as input data. This data is input into the system after being sorted and processed by the audited personnel, which may deviate from the real original data and there is a risk of being tampered with.
[0083] Drawback Two: Existing technologies generally require a large number of training samples to be manually or model-labeled, increasing labor or computational costs, and the accuracy of sample labeling is not necessarily high, which has a negative impact on the accuracy of model prediction. In Technology Two, the sensitive types and sensitive levels of assets are labeled through a pre-trained neural network model, and the data is classified and graded. This processing method is generally affected by different data sets in practical applications. And it depends on manual experience and the accuracy of the model, resulting in low accuracy of the sample labels themselves. In addition, the natural characteristic of log data is that the positive and negative sample distributions are extremely unbalanced. If a supervised method is used for training, there are great challenges in the selection of sampling methods and the screening of negative samples. For example, in Technology Two, the sensitive types and sensitive levels of assets are labeled and processed. And this imbalance often causes the model to overly focus on the majority class samples during the training process, thus ignoring the minority class samples, which affects the generalization ability and accuracy of the model; it also leads to difficulties in model evaluation, increased consumption of computing resources and a series of other problems.
[0084] Currently, the models for risk detection of log data generally adopt traditional machine learning models and deep learning models, and the accuracy of risk detection for log data is relatively poor. In Technique 2, the graph database and the graph neural network model have certain limitations when processing sequential data such as log data. Graph algorithms store and process data with nodes and edges, and are good at processing data with complex relationships and multi-level associations, but they have insufficient ability to express and learn features of sequential data. Technique 3 uses LSTM for model training. Although the gated system can effectively capture long-term dependencies in sequential data, the LSTM model requires a large amount of training data to learn the patterns and rules of logs. If the amount of training data is small, the model cannot fully learn the features of the data, resulting in low accuracy. In addition, LSTM has high requirements for data quality. If there are noises, missing values or outliers in the input data, it will easily affect the accuracy of the model.
[0085] Disadvantage 3: Existing technologies generally only use host logs as input data, which leads to the risk monitoring model can only detect the behavioral risks existing in the log data itself, and cannot detect the potential association information between user permissions and the potential association between user permissions and behavioral risks, resulting in some security risks that are difficult to identify. For example, in Technique 2, security logs containing IP, port, URL addresses, etc. are obtained from multiple data sources, and Technique 3 also only collects host log data.
[0086] In addition to the above disadvantages, the inventor believes that there are also problems in the existing technologies such as poor audit timeliness, lagging risk warnings, certain limitations and subjective biases, and it is difficult to discover relatively complex potential risks. The specific analysis is as follows:
[0087] Disadvantage 4: The existing technologies have poor timeliness for IT audits and have a certain lag in risk warnings. Technique 2 involves a large amount of data preprocessing processes, and the graph database and graph neural network adopted are computationally complex, making it difficult to achieve real-time training and prediction. On the one hand, Technique 2 involves a large number of operations such as cleaning, deduplication of log data, and screening out standardized logs according to preset screening strategies, and the preprocessing process is long. The deduplication operation involved often requires full traversal or partial traversal of the data set, which is a time-consuming and resource-consuming process in itself. Especially when the data volume is too large, the calculation time will increase significantly. On the other hand, in Technique 2, the graph database and the graph deep neural network store and calculate data in the form of a graph, including nodes and edges. This data structure requires complex traversal operations during query. As the volume of log data to be processed increases, the time complexity of traversal will increase significantly, and at the same time, the I / O overhead and memory usage will also increase significantly.
[0088] Disadvantage 5: The existing technologies generally rely on manual experience to extract audit elements, construct audit features, and select data sets, which have certain limitations and subjective biases, and it is difficult to discover relatively complex potential risks. In Technology 2, the manually selected log key fields are information such as target IP, port, and log time, and the selection of other elements is missing, which cannot reflect the complete information and chronological information of the log data; in the selection of the data set, only the standardized logs with specific characteristic key values are retained, and there are also biases in the sample selection.
[0089] Based on the deficiencies of the existing technologies, after research, the inventors proposed a targeted solution for the IT audit scenario and the practical needs in this scenario, and provided a log audit monitoring method and a log audit monitoring system. The log audit monitoring method also operates based on the log audit monitoring system. To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0090] See Figure 1 , which is an architecture diagram of a log audit monitoring system provided by an embodiment of this application. As Figure 1 shown, in the architecture of this log audit monitoring system, it includes: an agent program and a log audit monitoring model. In the technical solution of this application, the above agent program needs to be deployed on multiple key log production nodes. For example, it is deployed on a production operation and maintenance platform, a bastion host, a jump server, a host, a MySQL database, etc.
[0091] In an example scenario, if it is necessary to audit and monitor the logs of each branch, the agent program can specifically be deployed on the above multiple key log production nodes of each branch. In another example scenario, if it is necessary to audit and monitor the logs of each subsidiary company, the agent program can specifically be deployed on the above multiple key log production nodes of each subsidiary company. It can be understood that for different actual audit scenarios, as key log production nodes, the configured quantity and type may vary, and the quantity and type of the multiple key log production nodes are not specifically limited here. Figure 1 In the system architecture shown in
[0092] The aforementioned agent program can be expressed in English as Agent, and these agent programs can monitor and capture the log streams generated by the key nodes of log production in real time. By deploying the agent programs on the above-mentioned multiple key nodes of log production, it is convenient to collect the original log data generated by the production changes of operation and maintenance personnel in real time. The data collected by the agent programs on the key nodes of log production can also be called buried point data.
[0093] The log audit detection model in the log audit monitoring system provided by the embodiments of the present application is a deep learning model based on the DeepLog architecture and is trained by an unsupervised learning method. Different from the models of the conventional DeepLog architecture, the log audit monitoring model in the present application has dual functions. On the one hand, the model has the function of monitoring abnormal behaviors for logs, and on the other hand, it also has the function of monitoring personnel authorization. The operation principle of the model will be introduced in detail in the following embodiments.
[0094] Figure 2 It is a flowchart of a log audit monitoring method provided by the embodiments of the present application. As Figure 2 shown, the method includes:
[0095] S201. Collect the original log data generated at the key nodes of log production through the agent program.
[0096] The agent program can directly send the collected original log data to the log audit monitoring model in a predetermined format and protocol. In other possible implementation manners, it can also be that the agent program sends the collected original log data to the data processing module, and the data processing module provides it to the log audit monitoring model after preprocessing. Based on actual requirements, the functions of the data processing module may be diverse, and the preprocessing method is not limited here.
[0097] In a possible implementation manner, when providing the original log data, the agent program can also provide the device identifier of the key node of log production, or other identifiers used to refer to the data source. Through the identifier, the data source can be distinguished.
[0098] S202. Parse the data based on the original log data through the log audit monitoring model to obtain the log key sequence and the basic parameter value vector sequence corresponding to the original log data.
[0099] Although the log data is unstructured text data, it is generated by a program with strict logic and control flow, so it is essentially similar to natural language. Therefore, when the log audit monitoring model receives the original log data (or the data generated after the original log data is processed), it can parse the data and convert the log data into a structured representation.
[0100] In this application, when parsing log data, the timestamp of the log is combined to construct a sequence that can serve the purpose of audit monitoring.
[0101] Specifically, the log audit monitoring model can parse each log in the original log data into two parts: a log key and a parameter value vector. For the convenience of distinguishing from the concepts introduced later, the parameter value vector in the two directly parsed parts is herein referred to as the basic parameter value vector. Among them, the log key represents the category or type of the log, and the log key is the constant part of the log content. The basic parameter value vector, on the other hand, characterizes the specific parameter information of the log, such as the timestamp, IP address, process ID, version number, etc., and the basic parameter value vector is the variable part of the log content. Parsing the log into a log key and a parameter value vector can be achieved through regular expressions, that is, each piece of the original log data is parsed into a data format recognizable by the log audit monitoring model through regular expressions. For the convenience of understanding, Table 1 shows the effect of several pieces of original log data before parsing, and Table 2 shows the effect of several pieces of original log data after parsing to obtain the log key and the basic parameter value vector.
[0102] In Table 1, taking the first log as an example, [2023-04-01 09:05:00] represents the timestamp, and INFO:Starting new application version 2.3.1. represents the log information.
[0103] Table 1
[0104]
[0105]
[0106] In the log information reflected in Table 2, the timestamps of this log are simply represented by t1, t2, etc. Taking the first log shown in Table 2 as an example, Starting new application version is the constant part of the log, which is represented by the log key log key = k1, and the set of log keys is finite. Taking the first log shown in Table 2 as an example, 2.3.1. is the variable part of the log and needs to be represented by the basic parameter value vector. In addition, the basic parameter value vector also stores the time difference between this log and the previous log, such as t1 - t0, where t1 is the timestamp of this log and t0 represents the timestamp of the previous log.
[0107] Table 2
[0108]
[0109] Based on the results of the above parsing, in combination with the timestamp of the log, the log keys of each log can be arranged in sequence to form a sequence. In addition, in combination with the timestamp of the log, the basic parameter value vectors of each log can be arranged in sequence to form a sequence. In the embodiments of the present application, the log key sequence refers to the sequence formed by arranging the log keys of each log in the original log data in the order of the log timestamps; the basic parameter value vector sequence refers to the sequence formed by arranging the basic parameter value vectors of each log in the original log data in the order of the log timestamps.
[0110] S203. The log audit monitoring model matches based on the operator user information of each log in the original log data and the existing personnel information data set to obtain the personnel information parameters of each log in the original log data.
[0111] In the technical solution of the present application, the proxy program obtains the operator user information related to the log while collecting the original log data. By correlating with the existing personnel information data set, the personnel information parameters corresponding to each log are determined. In an alternative implementation, the existing personnel information data set includes multiple personnel information entries; each personnel information entry includes employee information, user information, and permission information; the employee information at least includes: the employee status, and the employee status is one of in-service, separated, transferred, or retired; the permission information includes: user role permissions and authorization period information.
[0112] For example, the personnel information data set (which can be in the form of a data table) of a certain subsidiary company includes information such as username, user role permissions, employee status, and the department. If the operator user is not in the personnel information data set of the subsidiary company, the user cannot be found in the personnel information data set. In addition, by using the personnel information data set for matching, specific information such as the status of the employee and whether the permissions have expired can also be clarified, so as to facilitate the model to identify whether there are permission anomalies in the personnel information corresponding to the log.
[0113] In the specific implementation of this step, the log audit monitoring model matches based on the operator user information of each log in the original log data to determine whether there is a personnel information entry corresponding to the user information in the existing personnel information data set; if it is determined from the existing personnel information data set that there is a personnel information entry corresponding to the user information, some or all of the information in the personnel information entry is extracted as the personnel information parameters of the corresponding log in the original log data; among them, the partial information at least includes employee status and permission information; if it is determined from the existing personnel information data set that there is no personnel information entry corresponding to the user information, the corresponding personnel information parameters in the original log data are set to be empty.
[0114] Table 2 shows the effect of parsing several original log data to obtain log keys and new parameter value vectors. The first column of Table 3, different from the first column of Table 2, shows the personnel information parameters of each log in the original log data obtained through this step. It should be noted that the step of obtaining personnel information parameters described in this step can be implemented by the log audit monitoring model or can be completed by the proxy program in combination with the personnel information data set. Generally speaking, by identifying the personnel information parameters, data enhancement of the inherently concerned log information is realized, so that the personnel information parameters can be parsed during parsing, and a new parameter value vector different from the basic parameter value vector can be constructed. In Table 3, only "personnel information parameters" are used as a reference. In fact, the personnel information parameters incorporated into the log information and the new parameter value vector may include employee ID, employee name, department, employee status (employed, separated, transferred, retired), user name, permission information, etc. Table 3 is only a form example and is not a limitation of the specific content.
[0115] Table 3
[0116]
[0117] S204. The log audit monitoring model fuses the basic parameter value vector sequence and the personnel information parameters of each log to obtain a new parameter value vector sequence.
[0118] In an optional implementation manner, this step includes: the log audit monitoring model splices the basic parameter value vector and the personnel information parameters of each log one by one based on each basic parameter value vector in the basic parameter value vector sequence and the obtained personnel information parameters of each log to obtain a new parameter value vector for each log in the original log data; the obtained new parameter value vectors of each log are arranged in the order of the log timestamps to form a new parameter value vector sequence.
[0119] Combining the differences between the third column of Table 3 and the third column of Table 2, it can be seen that the difference between the new parameter value vector shown in the third column of Table 3 and the basic parameter value vector shown in the third column of Table 2 lies in the personnel information parameters included in the vector. In this application, based on each basic parameter value vector in the basic parameter value vector sequence obtained after step S202 is executed, combined with the personnel information parameters of each log, the new parameter value vector obtained after parsing each log can finally be obtained. Based on the log timestamps, a new parameter value vector sequence can still be constructed in the order of the timestamps from the earliest to the latest.
[0120] S205. The log audit monitoring model constructs a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence.
[0121] In this application, by identifying the constant part and variable part in the logs, a log template representing a class of system time can be constructed. Here, the constant part can be understood as the log key, and the variable part can be understood as a new parameter value vector, where the basic parameter values and personnel information parameters belong to the variables. Since the log audit monitoring model is a model based on DeepLog architecture and the model requires the input to be a sequence of logs, in this step, it is necessary to construct a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence, which is convenient for subsequent use as the input of the log audit monitoring model.
[0122] It can be understood that the log template sequence corresponding to the original log data in this application includes the log templates of each log arranged in chronological order from first to last.
[0123] S206. Analyze the log template sequence through the log audit monitoring model to obtain the first anomaly monitoring result and the second anomaly monitoring result.
[0124] Figure 3 This is an architecture diagram for realizing dual anomaly monitoring through a log audit monitoring model provided by an embodiment of this application. In Figure 3 the left end of the log audit monitoring model shows the process of gradually constructing the log template sequence from the original log data. In the embodiment of this application, after receiving the log template sequence input into the model, the log audit monitoring model will perform the following analysis on the log template sequence, aiming to finally obtain the first anomaly monitoring result and the second anomaly monitoring result.
[0125] Among them, the first anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern. The second anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal permissions.
[0126] In the embodiment of this application, the log audit monitoring model used is obtained through learning and training with normal log templates. This means that the log audit monitoring model can identify isolated data that deviate from normal behavior, thereby detecting abnormal behavior or abnormal permission situations.
[0127] Among them, abnormal behaviors include frequent operation behaviors, operation behaviors in abnormal time periods, malicious attack behaviors, etc. The following will be introduced in three example cases.
[0128] (1) For frequent operation behaviors, the model pre - models the log template sequence through a long short - term memory network (LSTM) to learn the time intervals and frequency distributions of normal operation and maintenance operations. Therefore, when a new log stream is input, the model can calculate the operation frequency of the current operation and maintenance personnel for this log stream and compare it with the learned normal log pattern. If it is identified that the operation frequency and interval deviate from the normal pattern, frequent operations can be determined, and then a risk warning can be further triggered. At this time, the first anomaly monitoring result indicates that the log is triggered by an abnormal behavior pattern and can specifically reveal that the abnormal behavior is a frequent operation behavior. In this way, the log audit monitoring model can achieve real - time monitoring of the operation and maintenance personnel's operation behaviors. When it is detected that a certain operation and maintenance personnel has a large number of overly frequent operations in a short period of time, a warning is triggered and can be further displayed and fed back to the auditors.
[0129] (2) For operation behaviors during abnormal time periods, the model can pre - learn the time patterns of normal operation and maintenance operations and even learn the time windows that conform to the normal operation and maintenance operation patterns. In addition, it can also learn the operation habits of each operation and maintenance personnel in different time periods. Then, when a new log stream is input into the model, the model can determine whether the current operation deviates from the normal time period. If the determination result is yes, it is identified as an operation behavior during an abnormal time period and a warning is triggered. In this way, the model can identify the operation and maintenance operations carried out by operation and maintenance personnel during non - working time periods (such as late at night, holidays, etc.). Generally, such time periods are considered abnormal time periods, but combined with the differences in the company's business situation, the span range of abnormal time periods can also be set separately. When such operations are detected, the model can issue a warning to indicate that there may be a risk of unauthorized operations currently.
[0130] (3) For malicious attack behaviors, the model learns the characteristics and patterns of normal log data. When new log stream data is input, the model compares it with the learned normal pattern. If there are characteristics or patterns in the current log data summary that are significantly different from the normal pattern (such as abnormal network requests, abnormal database operations), it is determined as a malicious attack behavior and a warning is triggered. In this way, the model can identify potential malicious attack behaviors in the log, such as SQL injection, DDoS attacks, etc. These malicious attack behaviors usually easily lead to serious security incidents such as system anomalies or data leakage. And through the log audit monitoring model introduced in this solution, such problems can be effectively warned.
[0131] For the monitoring of personnel permissions, in the technical solution of this application, by combining personnel information parameters with operation and maintenance data, the expression of data is enhanced. Thus, in the model training stage, by constructing the above-mentioned data-enhanced training data, the model can learn relevant information about the authorization of operation and maintenance operators. By comparing the logs of the normal operation and maintenance operation behaviors of users within the normal permissions, the model can identify the user operations with abnormal permissions that deviate from the normal mode. In this way, the log audit monitoring model has the risk warning function in terms of user authorization. Through the log audit monitoring model, it is possible to identify the unauthorized operation behaviors of personnel in states such as leaving the company, retiring, or transferring positions who are still performing operation and maintenance operations. In addition, it is also possible to identify the abuse of permissions such as an employee frequently modifying the permission settings or performing sensitive operations within a short period of time, or abnormal situations such as a user who has not logged in for a long time suddenly becoming active.
[0132] The log audit monitoring method provided by the technical solution of this application is applied to a log audit monitoring system. This system has a real-time embedded end-to-end architecture, which at least includes agent programs deployed at multiple key log production nodes, and a log audit monitoring model that has both abnormal behavior monitoring functions and personnel authorization monitoring functions for logs; the model is a deep learning model based on the DeepLog architecture. Since this method uses the agent programs deployed at the key log production nodes to collect buried point data and directly processes the collected data through the model, the link of exporting log data or manually sorting data is avoided. The collected data is used as the source data for model processing, reducing the risk of log data being tampered with in the audit operation and making the audit monitoring results more authentic and effective. Since the model is a deep learning model based on the DeepLog architecture, it is trained in an unsupervised learning manner and can identify anomalies without a large amount of annotation work, greatly reducing the manual annotation cost. At the same time, it also avoids the problems of model overfitting or poor generalization ability caused by sample annotation, which in turn leads to poor accuracy of audit monitoring.
[0133] In the method, the model is used to parse data based on the original log data to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the model is used to match the operator user information of each log in the original log data with the existing personnel information data set to obtain the personnel information parameters of each log in the original log data; the model is used to fuse the basic parameter value vector sequence and the personnel information parameters of each log to obtain a new parameter value vector sequence; the model is used to construct a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence; the model is used to analyze the log template sequence to obtain a first anomaly monitoring result and a second anomaly monitoring result. Since in the technical solution of this application, the personnel information parameters of the log are used to implement data enhancement on the basis of the original basic parameter values, and a new parameter value vector sequence is fused and constructed, so as to assist the log audit monitoring model to also have the function of personnel authorization monitoring while having the function of anomaly behavior monitoring. It can be seen that the log audit monitoring model adopted in the technical solution of this application breaks the independent audit monitoring barrier between the monitoring and operation and maintenance platform and the personnel information system, can detect the log operations of users with permission violations, and can more effectively audit and monitor relevant security risks.
[0134] To facilitate the understanding of the working mode of the log audit monitoring model, the DeepLog architecture is briefly introduced below. Figure 4 As shown in the figure for the DeepLog architecture diagram, this architecture includes an input layer, an Embedding layer, an LSTM layer, a fully connected layer, and an output layer that are connected in sequence from the input to the output direction.
[0135] DeepLog is actually a conditional probability model. In the technical solution of this application, the log audit monitoring model based on the DeepLog architecture can output the probability distribution of the prediction result based on the input log template sequence. Among them, the prediction result with the highest probability represents the most likely log template corresponding to it. During the training process of the model, the weights of the model are continuously adjusted through the backpropagation algorithm to make the predicted value as close as possible to the true value. Figure 4 in represents the predicted value, y represents the true value, P represents the probability distribution, and the content in [] represents the input sequence. The subscript of T distinguishes different times. As Figure 4 shown in the architecture, the role of the Embedding layer is to convert each log template sequence from a scalar to a vector. The LSTM layer learns the long-term dependencies in the log data through the gating system. After passing through the softmax activation function of the fully connected layer, the model obtains the final prediction result.
[0136] In an optional implementation, the log audit monitoring model is pre-trained using a sample data set of normal log template sequences in an unsupervised learning manner; the sample data set of normal log template sequences includes multiple normal log template sequence samples; each normal log template sequence sample includes multiple log template samples of historical normal logs with a timestamp chronological relationship; the normal log template sequence samples are constructed based on multiple log key sequence samples of historical normal logs with a timestamp chronological relationship and new parameter value vector sequence samples; the new parameter value vector sequence samples are obtained by fusing a basic parameter value vector sequence based on multiple historical normal logs with a timestamp chronological relationship and the personnel information parameters of each historical normal log.
[0137] In the training phase, as a possible implementation method, cross entropy is used as the loss function for log key anomaly monitoring, and mean square error (MSE) is used as the loss function for parameter value anomaly monitoring. By minimizing these loss functions, the log audit monitoring model gradually learns the normal log pattern, thus having the dual anomaly monitoring function to achieve effective and accurate audit monitoring of logs.
[0138] To facilitate understanding of the implementation of step S206 in the above embodiment, please refer to Figure 5 , which shows an example process of implementing dual anomaly monitoring through a log audit detection model in an embodiment of the present application. Figure 5 As shown in , in an optional implementation, analyzing the log template sequence through the log audit monitoring model to obtain the first abnormal monitoring result and the second abnormal monitoring result may include the following steps:
[0139] S2061. The log audit monitoring model extracts the log key sequence to be analyzed from the log template sequence based on the preset log key window length h. t-h ,…,m t-2 ,m t-1}, based on the log key sequence to be analyzed w = {m t-h ,…,m t-2 ,m t-1}Analyze the next log key m t Check if there is an exception and get the log key m t The log key exception analysis result of the corresponding log.
[0140] Still see Figure 3, in the technical solution of this application, the log audit detection model needs to analyze the abnormal conditions of the log key sequence of the current log stream, and the analysis result is called the log key abnormal analysis result. In the analysis mechanism, if it is necessary to identify the abnormal conditions of the log key of a certain log, it is necessary to analyze the log keys of the previous h logs of this log, that is, the log key sequence w = {m t-h ,…,m t-2 ,m t-1} referred to in S2061.
[0141] Specifically, the log audit monitoring model predicts the conditional probability distribution of the next log key m t-h ,…,m t-2 ,m t-1} based on the log key sequence w = {m t to be analyzed. This conditional probability distribution is used to characterize the possibilities of various candidate situations of the next log key m t . If the actual value of the next log key m t is within the first g candidate situations in the conditional probability distribution, it is determined that the log key abnormal analysis result of this log is normal; if the actual value of the next log key m t is not within the first g candidate situations in the conditional probability distribution, it is determined that the log key abnormal analysis result of this log is abnormal. Here, the first g candidate situations correspond to normal candidate situations. In this application, g is a positive integer.
[0142] In the embodiments of this application, when analyzing whether there are abnormal operation behaviors, it is not only based on the log key abnormal analysis result, but also combined with the abnormal analysis result of the parameter value. In some cases, the log key abnormal analysis result indicates normal, but the abnormal analysis result of the parameter value indicates a problem, then the risk also needs to be concerned. The following S2062 will specifically introduce the analysis process of the abnormal conditions of the parameter value based on the new parameter value vector.
[0143] S2062: The log audit monitoring model extracts the new parameter value vector corresponding to each log key from the log template sequence, and based on the new parameter value vector corresponding to each log key, predicts whether there is an abnormality in the parameter value of the corresponding next log, and obtains the parameter value abnormal analysis result of each log.
[0144] The analysis result of parameter value anomalies is used to indicate whether there are anomalies in the basic parameter values and whether there are anomalies in the personnel permission information. As mentioned before, data enhancement was performed on the basis of the basic parameter value vector in combination with the personnel information parameters to form a new parameter value vector. When analyzing this vector, two levels of analysis results can be obtained. One is whether there are anomalies in the basic parameter values (such as version numbers, etc.), and the other is whether there are anomalies in the personnel permission information. The former level is combined with the auxiliary and log key anomaly analysis results to further determine the first anomaly monitoring result; the latter level is used to determine the second anomaly monitoring result. In the specific implementation of this step, it can be:
[0145] First, based on the new parameter value vector corresponding to each log key, predict the first prediction distribution and the second prediction distribution of the next corresponding log. Among them, the first prediction distribution is the prediction distribution of the basic parameter values for abnormal behavior monitoring, and the second prediction distribution is the prediction distribution of the personnel information parameters for personnel authorization monitoring.
[0146] Next, calculate the first confidence interval of the basic parameter values of the next log according to the first prediction distribution. If the actual value of the basic parameter values of the next log is within the first confidence interval, the analysis result of the parameter value anomalies of the next log indicates that the basic parameter values are normal; if the actual value of the basic parameter values of the next log is not within the first confidence interval, the analysis result of the parameter value anomalies of the next log indicates that the basic parameter values are abnormal.
[0147] Among them, the first confidence interval refers to the normal range of the basic parameter values.
[0148] In addition, calculate the second confidence interval of the personnel information parameters of the next log according to the second prediction distribution. If the actual value of the personnel information parameters of the next log is within the second confidence interval, the analysis result of the parameter value anomalies of the next log indicates that the generation of the next log is triggered by a person with normal permissions; if the actual value of the personnel information parameters of the next log is not within the second confidence interval, the analysis result of the parameter value anomalies of the next log indicates that the generation of the next log is triggered by a person with abnormal permissions.
[0149] Among them, the second confidence interval refers to the normal range of the personnel information parameters.
[0150] S2063. If the log key anomaly analysis result of the same log indicates an anomaly and / or the parameter value anomaly analysis result indicates that there are anomalies in the basic parameter values, obtain the first anomaly monitoring result indicating that the generation of this log is triggered by an abnormal behavior pattern; if the log key anomaly analysis result of the same log indicates no anomaly, and the parameter value anomaly analysis result indicates that there are no anomalies in the basic parameter values, obtain the first anomaly monitoring result indicating that the generation of this log is triggered by a normal behavior pattern.
[0151] S2064. If the analysis result of the parameter value anomaly indicates that there is an anomaly in the personnel permission information, obtain a second anomaly monitoring result indicating that the generation of this log is triggered by a person with abnormal permissions; if the analysis result of the parameter value anomaly indicates that there is no anomaly in the personnel permission information, obtain a second anomaly monitoring result indicating that the generation of this log is triggered by a person with normal permissions.
[0152] In an optional implementation, the log audit monitoring system is also configured with an information display and feedback mechanism. Figure 6 This is a flowchart of another log audit monitoring method provided by the embodiments of the present application. As Figure 6 shown, different from Figure 2 the corresponding embodiment, in this embodiment, after S206 analyzes the log template sequence through the log audit monitoring model to obtain the first anomaly monitoring result and the second anomaly monitoring result, the log audit monitoring method further includes:
[0153] S207. If the first anomaly monitoring result indicates an anomaly and / or the second anomaly monitoring result indicates an anomaly, generate a risk warning message and push the risk warning message to the audit end.
[0154] In this way, the user of the audit end or the audit-related system can receive the risk warning message of the abnormal situation in a timely manner. The risk warning message can not only indicate whether it is the first anomaly monitoring result that indicates an anomaly or the second anomaly monitoring result that indicates an anomaly. It can also further display the specific anomaly behavior type predicted by the model in the risk warning message. For example, the first anomaly monitoring result indicates an anomaly, and it is given that the model infers one of the abnormal behaviors of frequent operation behavior, operation behavior in an abnormal time period, and malicious attack behavior. And the risk warning message can also give the relevant log information, and the log information carries the information of the operating user, so as to facilitate the audit end to conduct further analysis and risk confirmation.
[0155] The way for the model to push the risk warning message to the audit end can include email push, SMS push, API interface push, etc. In addition, the push method can also be customized by the user, so as to facilitate the audit end user to timely understand the risk warning situation and take corresponding countermeasures in a timely manner, achieving the purpose of effective risk monitoring.
[0156] S208. If the risk misjudgment information fed back by the audit end based on the risk warning message is received, use the risk misjudgment information to continue learning and optimizing the log audit monitoring model.
[0157] In practical applications, users or related systems at the audit end can further analyze the risk warning information pushed by the model. For example, the audit end can analyze whether the model inference result is incorrect. In case of incorrect results, the audit end needs to feedback the risk misjudgment information to the log audit monitoring system. These risk misjudgment information can be effectively utilized. For example, if a normal log entry is incorrectly classified as abnormal, the log audit monitoring model can dynamically adjust the model weights based on the feedback of such risk misjudgment information, or re-input it into the model as a training sample. After learning and optimizing the model, the probability of misjudging normal as abnormal in the future can be reduced. At the same time, it also avoids the interference of incorrect push conclusions on the work efficiency of the audit end. That is to say, the parameter weights of the log audit monitoring model applied in the technical solution of this application can be adjusted and updated in real time based on actual feedback, thus ensuring the actual monitoring experience of users.
[0158] In an alternative implementation, the log audit monitoring system further includes: a data integration module and a data storage module; a Kafka stream database is configured in the data storage module; among them, the Kafka stream database is selected based on the real-time requirements of the log audit monitoring system. As a distributed stream processing platform, Kafka can receive, store, and forward data streams in real time, and can also act as a buffer for data streams. The stream database based on Kafka is convenient for model training and immediate prediction. The Kafka stream database can not only provide training data for the log audit monitoring model, but also serve to provide real-time log streams to be monitored to the log audit monitoring model for performing substantial monitoring.
[0159] Kafka has the characteristics of high performance, high throughput, and low latency, and is well suitable for being used as the access layer of real-time data sources. As mentioned before "Shortcoming 4": The existing technology has poor timeliness for IT audits and has a certain lag in risk warning. Regarding the real-time requirement, through the Kafka stream database, this lag problem is well solved, and the configuration of the data integration module and the data storage module makes the structure of the log audit monitoring system more complete and the function more powerful.
[0160] In the technical solution of this application, after collecting the original log data generated at the key log production nodes through the proxy program, the method further includes:
[0161] The data integration module uses ETL technology to automatically extract (Extract), clean, filter, transform, and standardize (Transform) log data from the original log data in the audit scenario application of the log audit monitoring system, and load (Load) the processed data into the Kafka stream database; the Kafka stream database provides the log data to be analyzed processed by ETL technology to the log audit monitoring model.
[0162] Based on the log audit monitoring method introduced above, correspondingly, this application also provides a log audit monitoring system. The specific implementation of this system will be introduced below with reference to the accompanying drawings.
[0163] As Figure 7 shown, the log audit monitoring system in this application includes: a data collection module 71 and a model prediction module 72. This log audit monitoring system has a real-time embedded end-to-end architecture.
[0164] The data collection module 71 is implemented by proxy programs deployed on multiple key log production nodes;
[0165] The model prediction module 72 is implemented by a log audit monitoring model that has both abnormal behavior monitoring function and personnel authorization monitoring function for logs; the log audit monitoring model is a deep learning model based on the DeepLog architecture;
[0166] The data collection module 71 is used to collect the original log data generated at the key log production nodes through the proxy programs;
[0167] The model prediction module 72 is used to perform data parsing on the original log data through the log audit monitoring model to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the log key sequence is a sequence formed by arranging the log keys of each log in the original log data in the order of the log timestamps; the basic parameter value vector sequence is a sequence formed by arranging the basic parameter value vectors of each log in the original log data in the order of the log timestamps;
[0168] The model prediction module 72 is also used to match the operation user information of each log in the original log data with the existing personnel information data set through the log audit monitoring model to obtain the personnel information parameters of each log in the original log data;
[0169] The model prediction module 72 is also used to fuse a new parameter value vector sequence based on the basic parameter value vector sequence and the personnel information parameters of each log through the log audit monitoring model;
[0170] The model prediction module 72 is also used to construct a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence through the log audit monitoring model;
[0171] The model prediction module 72 is also used to analyze the log template sequence through log auditing monitoring to obtain a first anomaly monitoring result and a second anomaly monitoring result; the first anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second anomaly monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal permissions.
[0172] Figure 8 shows another structural effect of the log auditing monitoring system. As Figure 8 shown, different from Figure 7 the way shown, in an alternative implementation, the log auditing monitoring system is further configured with a display feedback module 73, a data integration module 74, and a data storage module 75.
[0173] The display feedback module 73 is used to implement an information display and feedback mechanism; the display feedback module 73 is connected to the model prediction module 72;
[0174] The display feedback module 73 is specifically used for:
[0175] If the first anomaly monitoring result indicates an anomaly and / or the second anomaly monitoring result indicates an anomaly, it receives the risk warning information generated by the model prediction module 72 and pushes the risk warning information to the audit end for display. As Figure 8 can be seen from the words "risk warning display" in the display feedback module 73, which represents that this module has this display function.
[0176] If it receives the risk misjudgment information fed back by the audit end based on the risk warning information, it feeds back the risk misjudgment information to the model prediction module 72 so that the model prediction module 72 can use the risk misjudgment information to continue learning and optimizing the log auditing monitoring model. As Figure 8 can be seen from the words "risk misjudgment feedback display" in the display feedback module 73, which represents that this module also has the display function of the feedback information on whether the result identified by the model is confirmed as an anomaly, thus facilitating the user to view.
[0177] As Figure 8 , in an alternative implementation, the system further includes: a data integration module 74 and a data storage module 75. The data storage module 75 is configured with a Kafka stream database; among them, the Kafka stream database is selected based on the real-time requirements of the log auditing monitoring system.
[0178] The data integration module 74 is used to adopt ETL technology to automatically extract, clean, filter, transform, and standardize log data from the original log data in the audit scenario application of the log auditing monitoring system, and load the processed data into the Kafka stream database;
[0179] A data storage module 75, configured to provide log data to be analyzed processed by ETL technology to a log audit monitoring model through a Kafka stream database.
[0180] It should be noted that, in the embodiment of the present application, the data collected by the data collection module 71 enters the data integration module 74 in the form of a message queue. After the data integration module 74 completes the ETL processing of the data, it is provided to the data storage module 75. The data storage module 75 can store data through online storage or a Kafka stream database to provide data support to the model prediction module 72. If the model prediction module 72 detects an anomaly, the relevant data can be sent to the data storage module 75 to build a risk warning database, providing a data basis for subsequent risk analysis and event tracing. In addition, the model prediction module 72 sends the monitoring results to the display feedback module 73 for interaction with the audit side.
[0181] In the technical solution of the present application, compared with the prior art, it has the following advantages in multiple aspects:
[0182] (1) For the problem that the data submitted by the audited personnel may be tampered with and distorted, the present application directly embeds it into the business system. By deploying an Agent at the log production end to collect buried point data in real time and inputting the collected stream data into an optimized DeepLog model for real-time prediction, the tampering risk that the audit data may face in the export stage and the manual collation stage is effectively avoided; at the same time, the entire process is imperceptible to the business system, thus realizing efficient and secure risk monitoring and early warning.
[0183] (2) Aiming at the disadvantage of poor timeliness of the prior art, the present application adopts the method of real-time data collection and online model training and prediction to effectively improve the timeliness of the audit system, and can give early warning and prevention of risks in a timely manner. The present application collects log data in real time through buried points and performs quasi-real-time processing through online storage and a Kafaka stream database. The DeepLog model supports the input of stream data and supports online incremental update of the model, thus realizing rapid risk monitoring and detection.
[0184] (3) Aiming at the limitations and subjective biases in the selection of audit samples and audit features, the present application can process a large amount of data based on a big data platform, perform real-time processing and prediction on the full amount of log data of the business system, thereby realizing full-sample audit; the present application performs unified log template parsing on the original log data of various systems, which not only shields the heterogeneity of the logs, but also retains as much of the original log information as possible, without the need to manually construct features. By learning the original log information, it helps the model learn the complex association relationships and temporal relationships between log entries, thereby realizing more comprehensive risk monitoring and effectively identifying potential risk information.
[0185] (4)Regarding the problem of relying on labeled data, this application learns the characteristics and patterns of normal log data. When new log data streams are input, the model compares them with the learned normal patterns, and determines data that deviates from the normal log patterns as risks and anomalies. Through this unsupervised method, learning can be carried out on a dataset without labeled samples, which can effectively reduce the cost and difficulty of data annotation. It can well solve the problems in traditional practices, such as the limited quantity and quality of manually or model-labeled samples, resulting in too small training data or high sample homogeneity, overfitting in model training, and the difficulty of the supervised model in identifying abnormal situations outside the training samples.
[0186] (5)Regarding the problem of extremely unbalanced distribution of positive and negative samples in log data, this application adopts the DeepLog unsupervised model. The core of this method is that it does not require a large number of labeled samples (i.e., positive and negative samples) for training, but only needs to train the model based on a small amount of normal logs. It builds a model by learning the log patterns during normal execution, and can automatically detect anomalies when the logs deviate from the normal patterns. This training method avoids a series of problems caused by the unbalanced distribution of positive and negative samples in log data.
[0187] (6)Different from the conventional use of traditional machine learning and deep learning models, although the DeepLog algorithm adopted in this application is an algorithm for processing log sequences based on the LSTM neural network, it has been specifically optimized by combining the characteristics of system logs and IT audit requirements, such as log template extraction, abnormal pattern detection, workflow model, online incremental training, etc., and only needs to be trained on a small number of normal log training samples. In summary, the model in this application has higher accuracy and efficiency in the risk detection of system logs.
[0188] (7)This application takes personnel information as an important input, associates it with the corresponding log data, and realizes data augmentation of log information. The augmented data is input into the model to identify illegal over-authorization, privilege abuse and other behaviors, such as former employees or transferred employees still performing operation and maintenance operations, or an employee's account frequently modifying permission settings or performing sensitive operations within a short period of time.
[0189] The model application layer of this patent application is not limited to a specific model like DeepLog. It can be a model such as LogAnomaly which is also based on LSTM and supports real-time log anomaly prediction for streaming data. Through algorithm principles and comparison of model effects, DeepLog is superior to some extent. It is not excluded that better algorithm models will emerge in the model application layer now or in the future. The focus of this application is on the targeted system architecture design and implementation in the IT audit scenario from processes such as data collection, data integration, data storage, model application, and display feedback. And it focuses on comparing the problems existing in the risk monitoring model in the prior art, and proposes to adopt the DeepLog model in the model application layer to achieve higher accuracy and better timeliness. Data augmentation is also done in combination with personnel authorization information to optimize the model effect.
[0190] The log audit and monitoring system provided by this application combines big data technology for pre-event risk warning, improving the real-time performance and accuracy of auditing. Compared with traditional expert rule auditing, it can better discover potential risks and cover a more comprehensive range of detected risk points. Mainstream machine learning algorithms and deep learning algorithms have different degrees of limitations when dealing with unstructured data such as log information; in addition, unsupervised algorithm models such as association rules and clustering have low accuracy; supervised algorithms require a large amount of manual effort in sample annotation; although large models have powerful computing capabilities, the training process requires a large amount of computing resources and large-scale training samples. After comprehensive comparative analysis, DeepLog has great advantages in log data anomaly detection. The technical selection comparison results are shown in Table 4.
[0191] Table 4
[0192]
[0193] The log audit and monitoring system of the technical solution of this application deploys an Agent at the log production end to collect buried point data in real time, and inputs the collected streaming data into the improved DeepLog model for real-time prediction, so as to achieve the role of risk monitoring. This architecture design improves the timeliness of early warning on the one hand and greatly improves the transparency of the audit process on the other hand. By directly embedding it into the business system, the risk of tampering that the audit data may face in the export stage is effectively avoided; at the same time, the entire process is imperceptible to the business system, thus realizing efficient and secure risk monitoring and early warning.
[0194] At the same time, based on the architecture solution that combines deep learning and big data platforms, on the one hand, it can endow this system with the ability to process massive data, thus breaking through the limitations of traditional sampling audits and moving towards the era of full-sample audits. On the other hand, the automated audit system transforms the traditional "project-based" regular audit mode into a composite mechanism of "monitoring-based" daily supervision and real-time warning. At the same time, through the online monitoring system, it can help the audit operation shift more from "on-site audit" to "remote audit". The above-mentioned transformation can not only improve the efficiency and value of audit work, but also further strengthen the enterprise's risk resistance and prevention and control capabilities.
[0195] It should be noted that each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0196] As described above, it is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A log audit monitoring method, characterized in that: The method is applied to a log audit monitoring system; The log audit monitoring system has a real-time embedded end-to-end architecture, and the log audit monitoring system at least includes: an agent program deployed at multiple key log production nodes, and a log audit monitoring model with both abnormal behavior monitoring functions and personnel authorization monitoring functions for logs; the log audit monitoring model is a deep learning model based on the DeepLog architecture; The method comprises: Collecting the original log data generated at the key log production node through the agent program; The log audit monitoring model is used to perform data parsing based on the original log data to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the log key sequence is a sequence in which the log key of each log in the original log data is arranged in the order of the log timestamps; the basic parameter value vector sequence is a sequence in which the basic parameter value vector of each log in the original log data is arranged in the order of the log timestamps; The log audit monitoring model matches the operating user information of each log in the original log data with the existing personnel information data set to obtain the personnel information parameters of each log in the original log data; The log audit monitoring model obtains a new parameter value vector sequence by fusing the basic parameter value vector sequence and the personnel information parameter of each log; The log audit monitoring model constructs a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence; The log template sequence is analyzed through the log audit monitoring model to obtain a first abnormal monitoring result and a second abnormal monitoring result; the first abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal authority.
2. The method according to claim 1, characterized in that The existing personnel information data set includes multiple personnel information items; each personnel information item includes employee information, user information and authority information; the employee information at least includes: employee status, which is one of being employed, leaving, transferring or retiring; the authority information includes: user role authority and authorization period information; The log audit monitoring model matches the operating user information of each log in the original log data with the existing personnel information data set to obtain the personnel information parameters of each log in the original log data, including: The log audit monitoring model matches, based on the operating user information of each log in the original log data, whether there is a personnel information entry corresponding to the user information in the existing personnel information data set; If it is determined from the existing personnel information data set that there is a personnel information entry corresponding to the user information, extract part or all of the information in the personnel information entry as the personnel information parameter of the corresponding log in the original log data; wherein the part of the information at least includes the employee status and the authority information; If it is determined from the matching of the existing personnel information data set that there is no personnel information entry corresponding to the user information, the corresponding personnel information parameter in the original log data is set to empty.
3. The method according to claim 1, characterized in that The log audit monitoring model obtains a new parameter value vector sequence by fusing the basic parameter value vector sequence and the personnel information parameter of each log, including: The log audit monitoring model performs concatenation of the basic parameter value vectors and the personnel information parameters of each log based on each basic parameter value vector of the basic parameter value vector sequence and the obtained personnel information parameters of each log, and obtains a new parameter value vector for each log in the original log data; The new parameter value vectors of each log obtained are arranged in the order of the log timestamps to form a new parameter value vector sequence.
4. The method according to claim 1, characterized in that The analyzing the log template sequence by the log audit monitoring model to obtain the first abnormal monitoring result and the second abnormal monitoring result includes: The log audit monitoring model extracts the log key sequence to be analyzed from the log template sequence based on the preset log key window length h. t-h ,…,m t-2 ,m t-1 }, based on the log key sequence to be analyzed w={m t-h ,…,m t-2 ,m t-1 }Analyze the next log key m t Check if there is an exception, and get the log key m t The log key anomaly analysis result of the corresponding log; The log audit monitoring model extracts a new parameter value vector corresponding to each log key from the log template sequence, and based on the new parameter value vector corresponding to each log key, predicts whether the parameter value of the corresponding next log is abnormal, and obtains the parameter value abnormality analysis result of each log; the parameter value abnormality analysis result is used to indicate whether the basic parameter value is abnormal and whether the personnel authority information is abnormal; If the log key anomaly analysis result of the same log indicates an anomaly and / or the parameter value anomaly analysis result indicates that the basic parameter value is anomaly, a first anomaly monitoring result indicating that the generation of the log is triggered by an abnormal behavior pattern is obtained; if the log key anomaly analysis result of the same log indicates that there is no anomaly, and the parameter value anomaly analysis result indicates that the basic parameter value is not anomaly, a first anomaly monitoring result indicating that the generation of the log is triggered by a normal behavior pattern is obtained; If the parameter value abnormal analysis result indicates that there is an abnormality in the personnel authority information, a second abnormal monitoring result is obtained, indicating that the generation of the log is triggered by a person with abnormal authority; if the parameter value abnormal analysis result indicates that there is no abnormality in the personnel authority information, a second abnormal monitoring result is obtained, indicating that the generation of the log is triggered by a person with normal authority.
5. The method according to claim 4, characterized in that The log key sequence w={m t-h ,…,m t-2 ,m t-1 }Analyze the next log key m t Check if there is an exception, and get the log key m t The corresponding log key anomaly analysis results for the log include: Based on the log key sequence to be analyzed w={m t-h ,…,m t-2 ,m t-1 }Predict the next log key m t The conditional probability distribution of the next log key m t the possibility of multiple candidate scenarios; If the next log key m t If the actual value of is within the first g candidate cases in the conditional probability distribution, the log key abnormality analysis result of the log is determined to be normal; if the next log key m t If the actual value of is not within the first g candidate cases in the conditional probability distribution, the log key anomaly analysis result of the log is determined to be abnormal.
6. The method according to claim 4, characterized in that The method of predicting whether the parameter value of the corresponding next log is abnormal based on the new parameter value vector corresponding to each log key, and obtaining the abnormal parameter value analysis result of each log, includes: Based on the new parameter value vector corresponding to each log key, predict the first predicted distribution and the second predicted distribution of the corresponding next log; the first predicted distribution is the predicted distribution of the basic parameter value for abnormal behavior monitoring, and the second predicted distribution is the predicted distribution of the personnel information parameter for personnel authorization monitoring; Calculating a first confidence interval of the basic parameter value of the next log according to the first predicted distribution; if the actual value of the basic parameter value of the next log is within the first confidence interval, the parameter value abnormality analysis result of the next log indicates that the basic parameter value is normal; if the actual value of the basic parameter value of the next log is not within the first confidence interval, the parameter value abnormality analysis result of the next log indicates that the basic parameter value is abnormal; A second confidence interval of the personnel information parameter of the next log is calculated according to the second predicted distribution. If the actual value of the personnel information parameter of the next log is within the second confidence interval, the parameter value abnormal analysis result of the next log indicates that the generation of the next log is triggered by a person with normal authority; if the actual value of the personnel information parameter of the next log is not within the second confidence interval, the parameter value abnormal analysis result of the next log indicates that the generation of the next log is triggered by a person with abnormal authority.
7. The method according to claim 1, characterized in that The log audit monitoring model is pre-trained by a sample data set of a normal log template sequence in an unsupervised learning manner; The sample data set of the normal log template sequence includes a plurality of normal log template sequence samples; Each of the normal log template sequence samples includes a plurality of log template samples of historical normal logs having a chronological relationship of timestamps; The normal log template sequence samples are constructed based on the log key sequence samples of the multiple historical normal logs with a chronological relationship of timestamps and the new parameter value vector sequence samples; The new parameter value vector sequence sample is obtained by fusing the basic parameter value vector sequence of the multiple historical normal logs with a chronological relationship of timestamps and the personnel information parameter of each historical normal log.
8. The method according to claim 1, characterized in that The log audit monitoring system is also equipped with an information display and feedback mechanism; After analyzing the log template sequence by the log audit monitoring model to obtain the first abnormal monitoring result and the second abnormal monitoring result, the method further includes: If the first abnormal monitoring result indicates an abnormality and / or the second abnormal monitoring result indicates an abnormality, risk warning information is generated and pushed to the audit end; If the risk misjudgment information fed back by the audit end based on the risk warning information is received, the log audit monitoring model is further studied and optimized using the risk misjudgment information.
9. The method according to claim 1, characterized in that: The log audit monitoring system further includes: a data integration module and a data storage module; a Kafka stream database is configured in the data storage module; wherein the Kafka stream database is selected based on the real-time requirements of the log audit monitoring system; After collecting the original log data generated at the key log production node through the agent program, the method further includes: The data integration module adopts ETL technology to automatically extract, clean, filter, convert and standardize the original log data in the audit scenario application of the log audit monitoring system, and loads the processed data into the Kafka stream database; the Kafka stream database provides the log audit monitoring model with the log data to be analyzed that has been processed by the ETL technology.
10. A log audit monitoring system, characterized in that: The log audit monitoring system has a real-time embedded end-to-end architecture, and the log audit monitoring system at least includes: a data acquisition module and a model prediction module; The data collection module is implemented by an agent program deployed at multiple key log production nodes; The model prediction module is implemented by a log audit monitoring model that has both abnormal behavior monitoring functions and personnel authorization monitoring functions for logs; the log audit monitoring model is a deep learning model based on the DeepLog architecture; The data collection module is used to collect the original log data generated at the key log production node through the agent program; The model prediction module is used to perform data analysis based on the original log data through the log audit monitoring model to obtain a log key sequence and a basic parameter value vector sequence corresponding to the original log data; the log key sequence is a sequence in which the log key of each log in the original log data is arranged in the order of log timestamps; the basic parameter value vector sequence is a sequence in which the basic parameter value vector of each log in the original log data is arranged in the order of log timestamps; The model prediction module is further used for the log audit monitoring model to match the operating user information of each log in the original log data with the existing personnel information data set to obtain the personnel information parameters of each log in the original log data; The model prediction module is further used for obtaining a new parameter value vector sequence by fusing the log audit monitoring model based on the basic parameter value vector sequence and the personnel information parameter of each log; The model prediction module is further used to construct a log template sequence corresponding to the original log data based on the log key sequence and the new parameter value vector sequence by the log audit monitoring model; The model prediction module is also used to analyze the log template sequence through the log audit monitoring model to obtain a first abnormal monitoring result and a second abnormal monitoring result; the first abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by an abnormal behavior pattern; the second abnormal monitoring result is used to indicate whether the generation of a certain log in the original log data is triggered by a person with abnormal authority.
11. The system according to claim 10, characterized in that The system is also configured with a display feedback module, which is used to implement information display and feedback mechanism; the display feedback module is connected to the model prediction module; The display feedback module is specifically used for: If the first abnormal monitoring result indicates an abnormality and / or the second abnormal monitoring result indicates an abnormality, receiving the risk warning information generated by the model prediction module, and pushing the risk warning information to the audit end for display; If the audit end receives risk misjudgment information based on the risk warning information, the risk misjudgment information is fed back to the model prediction module so that the model prediction module can continue to learn and optimize the log audit monitoring model using the risk misjudgment information.
12. The system according to claim 10 or 11, characterized in that The system further comprises: a data integration module and a data storage module; a Kafka stream database is configured in the data storage module; wherein the Kafka stream database is selected based on the real-time requirements of the log audit monitoring system; The data integration module is used to use ETL technology to automatically extract, clean, filter, convert and standardize log data from the original log data in the audit scenario application of the log audit monitoring system, and load the processed data into the Kafka stream database; The data storage module is used to provide the log data to be analyzed that has been processed by the ETL technology to the log audit monitoring model through the Kafka stream database.
Citation Information
Patent Citations
Log audit monitoring method and device
CN116318786A
Client information management system based on enterprise information security monitoring
CN119128900A
System for integrally analyzing and auditing heterogeneous personal information protection products
KR1020180075279A