Distributed database anomaly detection method based on multivariable logs
Through the combination of algorithms such as LFA, RoBERTa, Transformer and VAE, the complexity of abnormal detection in distributed databases is solved, efficient and accurate abnormal detection is achieved, operation and maintenance costs are reduced, user experience is improved, and diverse application scenarios are adapted to.
Patent Information
- Application Number
- CN202510436488.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
When the existing distributed database anomaly detection model processes data stored across multiple nodes and is distributed unevenly, it has problems such as poor detection effect, insufficient real-time, high communication costs and poor scalability, making it difficult to adapt to rapidly changing data characteristics and threats.
The LFA algorithm is used to parse log data and group it, and the deep semantic features are extracted in combination with the improved RoBERTa model. The multi-head self-attention mechanism of the Transformer architecture is used to analyze the log event sequence dependencies, and noise is removed through VAE and anomaly detection is performed using a random forest clustering classifier.
It realizes efficient and accurate abnormal detection, reduces operation and maintenance costs, improves user experience, supports decision-making, adapts to diverse application scenarios, and has the ability to continuously learn and self-optimize.
Smart Images

Figure CN120371688A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control, and particularly to a method for detecting anomalies in a distributed database based on multivariate logs. Background Art
[0002] The models for detecting anomalies in distributed databases face many complex challenges. These challenges not only affect the performance and accuracy of the models but also increase the difficulty of system design and implementation. Most existing anomaly detection algorithms and models are designed based on a single-node environment. When faced with data stored across multiple nodes and unevenly distributed, there may be problems with poor detection effects for certain types of anomalies. For anomalies that require quick responses, existing models may not be able to well meet the real-time requirements. Especially when dealing with large-scale data sets, the processing delay will become a problem. To perform effective anomaly detection in a distributed system, nodes may need to frequently exchange information, which will lead to an increase in communication costs and affect the overall performance of the system. With the development of business and the progress of technology, data patterns and access patterns are also constantly evolving. This requires that the anomaly detection model must have good scalability and adaptability so that it can quickly adjust to cope with newly emerging threats or changing data characteristics.
[0003] All these factors combined make it complex and challenging to achieve efficient and accurate anomaly detection in a distributed database environment. Summary of the Invention
[0004] The present invention provides a method for detecting anomalies in a distributed database based on multivariate logs to solve the technical problems mentioned in the background art.
[0005] The method for detecting anomalies in a distributed database based on multivariate logs includes:
[0006] S1. In a distributed system, each node generates unstructured log data, and randomly divides the unstructured log data into training logs and test logs;
[0007] S2. Use the LFA (Log Forest Algorithm) algorithm to parse the log data and divide it into log event classifications and log event groups;
[0008] S3. Introduce the improved pre-trained language model RoBERTa to extract deep semantic features for the log events of each distributed database;
[0009] S4. To further capture the dependency relationships of the log event sequences, use the multi-head self-attention mechanism in the Transformer architecture for in-depth analysis;
[0010] S5. Based on the deep semantic features of the log events obtained in S3 and S4 and the sequence dependence relationship, we independently estimate and calculate a probability for the database log of each node, which reflects the possibility of a specific log event occurring or belonging to an abnormal pattern. We use an encoder (VAE) to remove the noise in the features and enhance the consistency of the feature representation, and standardize the length of the probability output by each node. The VAE maps the input features to the latent space through the encoder, thereby removing the noise;
[0011] S6. The standardized feature vectors are input into a random forest clustering classifier for anomaly detection.
[0012] Preferably, in step S2, the steps of parsing the log data using the LFA (Log Forest Algorithm) algorithm include:
[0013] LFA uses a tree structure to quickly extract log templates, constructs one or more log trees, and each node of the tree represents a log template; each log entry finds the most similar template through matching with the existing templates and uses a similarity calculation function to classify and group it. The formula is as follows:
[0014] ;
[0015] where, is the existing log template, is the current log entry, represents the similarity calculation function. Through this method, the unstructured log data is efficiently parsed into a structured format, that is, it is divided into log event classifications and log event groups.
[0016] Preferably, in step S3, each log event is mapped to a high-dimensional vector space to generate a semantic feature vector , combined with the time characteristic and the event frequency , through a feature fusion method, a count vector is generated. The formula is as follows:
[0017] ;
[0018] where, is the feature fusion function, is the semantic feature obtained from the RoBERTa model, is the time characteristic vector, is the frequency information. In this way, the semantic information is effectively combined with the time and frequency characteristics, and the generated count vector contains more comprehensive log event features.
[0019] Preferably, in step S4, the multi-head self-attention mechanism calculates the attention weights through the following formula:
[0020] ;
[0021] where, , , , , are trainable parameter matrices, is the key vector dimension, is the input feature matrix, and the multi-head mechanism is further extended to:
[0022] ;
[0023] where , this mechanism can simultaneously focus on different parts of the log sequence, capture dependencies within a long time span, and significantly improve the context information capture ability and processing efficiency.
[0024] Preferably, in step S5, the encoding process is:
[0025] ;
[0026] where, is the latent variable, is the original feature, and are the mean and standard deviation of the encoder output respectively. During the decoding process, the reconstructed feature is generated through the following formula:
[0027] ;
[0028] Joint optimization is performed through the reconstruction loss and divergence:
[0029] ;
[0030] where, is the objective function or loss function of the entire model, used to optimize the model parameters; represents the expected value of z, where z is sampled from the conditional distribution . Here, is a posterior probability distribution controlled by the parameter , which describes the distribution of the latent variable z given the input x; KL is the divergence.
[0031] This processing method makes the feature vectors more stable and consistent, thus providing more reliable data input for subsequent anomaly detection.
[0032] Preferably, in step S6, the standardized feature vector is input into a random forest clustering classifier for anomaly detection. The random forest consists of multiple decision trees, and each tree is trained based on a feature subset. The classification process is as follows:
[0033] ;
[0034] wherein, is the number of decision trees, is the prediction result of the th tree. When the classification result indicates that the current feature vector belongs to the abnormal category, the system immediately triggers an alarm and records the details of the anomaly. Otherwise, it is marked as normal.
[0035] Beneficial effects achieved by the present invention:
[0036] Reduce operation and maintenance costs: Through automated anomaly detection and diagnosis processes, the need for manual intervention is reduced. The operation and maintenance team can focus more on system optimization and other key tasks, thereby reducing labor costs and improving overall operational efficiency.
[0037] Enhance user experience: Precise anomaly detection helps quickly locate the root cause of problems and shorten the fault recovery time.
[0038] Support decision-making: The combination of deep learning and traditional statistical methods provides rich analysis results of abnormal behaviors. This information can help management better understand the system operation status, identify potential risk points, and provide a basis for strategic planning and technology investment.
[0039] Promote technological innovation: The various advanced algorithms and technologies used in the invention, such as RoBERTa, Transformer, and VAE, drive technological progress in related fields. This is not only beneficial to the implementation of the current project but also lays a solid foundation for future research and development.
[0040] Adapt to diverse application scenarios: The solution is designed flexibly and is applicable to different types and scales of distributed database environments. Whether it is a small startup or a large multinational enterprise, it can be customized and deployed according to its own needs to ensure the best performance.
[0041] Continuous learning and self-optimization: The machine learning-based method enables the system to have the ability of self-optimization. As new data continuously flows in, the model can automatically adjust parameters to adapt to changing workload patterns and maintain long-term effectiveness. Description of the Drawings
[0042] Figure 1 is a flowchart of the distributed database anomaly detection method based on multivariate logs. Detailed Embodiments
[0043] The technical solution of the present invention will be described in detail below in conjunction with specific drawings.
[0044] Please refer to Figure 1 , the embodiment of the present invention provides a distributed database anomaly detection method based on multivariate logs, and the method includes:
[0045] S1. In a distributed system, each node generates a large amount of unstructured log data. To facilitate the training and testing of the model, these unstructured log data are randomly divided into training logs and test logs;
[0046] S2. Use the LFA (Log Forest Algorithm) algorithm to parse the log data and divide it into log event classification and log event grouping;
[0047] S3. Introduce the improved pre-trained language model RoBERTa to extract deep semantic features of log events for each distributed database;
[0048] S4. In order to further capture the dependency relationship of the log event sequence, use the multi-head self-attention mechanism in the Transformer architecture for in-depth analysis;
[0049] S5. Based on the deep semantic features of the log events and the sequence dependency relationship obtained in S3 and S4, we independently estimate and calculate a probability for the database log of each node, which reflects the possibility of a specific log event occurring or belonging to a certain abnormal pattern. Use the encoder (VAE) to remove the noise in the features and improve the consistency of the feature representation, standardize the length of the probability output by each node, and the VAE maps the input features to the latent space through the encoder to remove the noise;
[0050] S6. The standardized feature vectors are input into a random forest clustering classifier for anomaly detection.
[0051] In step S2 of this embodiment, the steps of parsing the log data using the LFA (Log Forest Algorithm) algorithm include:
[0052] LFA uses a tree structure to quickly extract log templates, constructs one or more log trees, and each node of the tree represents a log template; each log entry is matched with the existing templates, and the most similar template is found using a similarity calculation function and classified and grouped, and the formula is as follows:
[0053] ;
[0054] Wherein, is the existing log template, is the current log entry, represents the similarity calculation function. In this way, the unstructured log data is efficiently parsed into a structured format, that is, it is divided into log event classification and log event grouping.
[0055] In step S3 of this embodiment, each log event is mapped to a high-dimensional vector space to generate a semantic feature vector , combined with the time characteristics and event frequency , through the feature fusion method, a count vector is generated , the formula is as follows:
[0056] ;
[0057] Among them, is the feature fusion function, is the semantic feature obtained from the RoBERTa model, is the time characteristic vector, is the frequency information. In this way, the semantic information is effectively combined with the time and frequency characteristics, and the generated count vector contains more comprehensive log event characteristics.
[0058] In step S4 of this embodiment, the multi-head self-attention mechanism calculates the attention weights through the following formula:
[0059] ;
[0060] Among them, , , , , are trainable parameter matrices, is the key vector dimension, is the input feature matrix, and the multi-head mechanism is further extended to:
[0061] ;
[0062] Among them , this mechanism can simultaneously focus on different parts of the log sequence, capture the dependencies within a long time span, and significantly improve the context information capture ability and processing efficiency.
[0063] In step S5 of this embodiment, the encoding process is:
[0064] ;
[0065] Among them, is the latent variable, is the original feature, and are the mean and standard deviation of the encoder output respectively. During the decoding process, the reconstructed feature is generated through the following formula:
[0066] ;
[0067] Joint optimization is performed through the reconstruction loss and divergence:
[0068] ;
[0069] where, is the objective function or loss function of the entire model, which is used to optimize the model parameters; represents the expected value of z, where z is sampled from the conditional distribution . Here, is a posterior probability distribution controlled by the parameter , which describes the distribution of the latent variable z given the input x; KL is the divergence.
[0070] This processing method makes the feature vector more stable and consistent, thus providing a more reliable data input for subsequent anomaly detection.
[0071] In step S6 of this embodiment, the standardized feature vector is input into a random forest clustering classifier for anomaly detection. The random forest consists of multiple decision trees, and each tree is trained based on a feature subset. The classification process is as follows:
[0072] ;
[0073] where, is the number of decision trees, is the prediction result of the th tree. When the classification result indicates that the current feature vector belongs to the anomaly category, the system immediately triggers an alarm and records the anomaly details, otherwise it is marked as normal.
[0074] The innovation of the present invention is reflected in the following aspects:
[0075] (1)Efficient Log Parsing Method: The present invention adopts the LFA (Log Forest Algorithm) log template mining method. Through in-depth analysis of the log format, the LFA method can quickly extract different types of log information, thus greatly reducing the parsing time. When dealing with large-scale log data, the efficiency of LFA is particularly prominent. It can quickly complete the parsing task of massive logs while maintaining accuracy, ensuring the timeliness and reliability of data processing. The log events processed by LFA are effectively classified and grouped, making the originally chaotic log data well-organized, facilitating subsequent time sorting and logical analysis.
[0076] (2)Deep Semantic Mining: In terms of semantic understanding, the present invention introduces the RoBERTa model (a pre-trained language model based on BERT), and combines temporal characteristics with semantic features to generate richer count vectors. These count vectors can capture deeper semantic information of log data, providing a solid foundation for subsequent analysis and decision-making. By introducing deep learning methods, the model can extract more potential and valuable information from a large amount of unstructured data, further improving the effectiveness of log parsing and anomaly detection.
[0077] (3)Transformer Architecture: The present invention is based on the Transformer architecture and uses the multi-head self-attention mechanism (Multi-head Attention) to analyze long-term dependencies in log data. The Transformer architecture can effectively handle dependencies between different time points and is particularly suitable for processing long sequence data. The multi-head self-attention mechanism captures context information at different levels through parallel computing, improving the computational efficiency of the model and enhancing the model's ability to understand the context of log content, thus showing higher accuracy and robustness in log parsing and anomaly detection.
[0078] (4)Enhanced Stability: To improve the stability of the system and the accuracy of anomaly detection, the present invention uses a variational autoencoder (VAE) for denoising and normalizing feature vectors. The VAE reduces the impact of data noise on model training by denoising the input data, thus enhancing the stability of the model. In addition, combining multiple anomaly detection algorithms such as random forests can more accurately identify abnormal patterns, effectively improving the detection accuracy and precision of the system. The combination of these methods enables the system to maintain efficient and stable performance in a complex and changing log data environment.
[0079] It should be noted that in this text, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes such element.
[0080] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present invention.
Claims
1. A method for anomaly detection in a distributed database based on multivariate logs, characterized in that The method includes: S1. In a distributed system, each node generates unstructured log data, and randomly divides the unstructured log data into training logs and test logs; S2. Use the LFA algorithm to parse the log data; and divide it into log event classification and log event grouping; S3. Introduce the improved pre-trained language model RoBERTa to extract deep semantic features of log events for each distributed database; S4. To further capture the dependency relationship of the log event sequence, use the multi-head self-attention mechanism in the Transformer architecture for in-depth analysis; S5. Based on the deep semantic features of the log events and the sequence dependency obtained in S3 and S4, independently estimate and calculate a probability for the database log of each node, which reflects the possibility of a specific log event occurring or belonging to an abnormal pattern, use an encoder to remove noise in the features and improve the consistency of the feature representation, standardize the length of the probability output by each node, and the VAE maps the input features to the latent space through the encoder to remove noise; S6. The standardized feature vectors are input into a random forest clustering classifier for anomaly detection.
2. The distributed database anomaly detection method based on multivariate logs according to claim 1, wherein In step S2, the steps of using the LFA algorithm to parse the log data include: LFA uses a tree structure to quickly extract log templates, constructing one or more log trees, where each node of the tree represents a log template; each log entry is matched with existing templates and the similarity calculation function is used to find the most similar template and classify and group it. The formula is as follows: ; Among them, is an existing log template, is the current log entry, represents a similarity calculation function. Unstructured log data is efficiently parsed into a structured format, that is, it is divided into log event classifications and log event groups.
3. The distributed database anomaly detection method based on multivariate logs according to claim 1, wherein In step S3, each log event is mapped to a high-dimensional vector space to generate a semantic feature vector , combined with the time characteristics and the event frequency , through a feature fusion method, a count vector is generated , and the formula is as follows: ; Among them, is the feature fusion function, is the semantic feature obtained from the RoBERTa model, is the time characteristic vector, is the frequency information.
4. The distributed database anomaly detection method based on multivariate logs according to claim 1, characterized in that, In step S4, the multi-head self-attention mechanism calculates the attention weights through the following formula: ; Among them, , , , , are trainable parameter matrices, is the key vector dimension, is the input feature matrix, and the multi-head mechanism is further extended to: ; Among them 。 5. The distributed database anomaly detection method based on multivariate logs according to claim 1, wherein, In step S5, the encoding process is: ; Among them, is a latent variable, is an original feature, and are the mean and standard deviation of the encoder output respectively. During the decoding process, the reconstructed features are generated through the following formula: ; Joint optimization through reconstruction loss and divergence: 。 Among them, is the objective function or loss function of the entire model, which is used to optimize the model parameters; represents the expected value of z, where z is sampled from the conditional distribution and is a posterior probability distribution controlled by the parameter which describes the distribution of the latent variable z given the input x; KL is the divergence.
6. The distributed database anomaly detection method based on multivariate logs according to claim 1, wherein In step S6, the standardized feature vectors are input into a random forest clustering classifier for anomaly detection. The random forest consists of multiple decision trees, and each tree is trained based on a subset of features. The classification process is: ; Among them, is the number of decision trees, is the prediction result of the th tree. When the classification result indicates that the current feature vector belongs to the abnormal category, the system immediately triggers an alarm and records the details of the abnormality. Otherwise, it is marked as normal.
Citation Information
Cited By
Log anomaly detection method based on log classification and deep learning
CN120744921A
A log anomaly detection method based on log classification and deep learning
CN120744921B