Deep learning-driven system for detecting performance anomalies
A deep learning-based system addresses the challenge of accurate anomaly detection in complex systems by using RNNs and LSTMs to preprocess and adaptively learn system behavior, enhancing detection accuracy and reducing downtime.
Patent Information
- Application Number
- DE202025102432
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-05-04
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2035-05-31
AI Technical Summary
Traditional performance monitoring tools struggle with accurate anomaly detection in complex systems, leading to false positives and missed detections, and there is a need for a more intelligent, scalable, and adaptive approach that can learn normal behavior and adapt to system changes.
A deep learning-based system utilizing RNNs and LSTMs analyzes historical and real-time performance data to detect anomalies, preprocesses data, and provides automated insights, with adaptive learning to continuously improve.
Enables real-time anomaly detection, reduces false alarms, and supports proactive alerts, optimizing system resources and minimizing downtime.
Smart Images

Figure 00000005_0000
Abstract
Description
The present invention relates to a system for detecting performance anomalies in computer systems, applications or networks using deep learning techniques. The system utilizes an advanced deep learning architecture to monitor system performance metrics and identify patterns that deviate from the expected behavior, thus allowing early detection of potential system failures, inefficiencies, or vulnerabilities.As systems and networks become more and more complex, conventional performance monitoring tools have difficulty accurately detecting anomalies. Most conventional systems rely on predefined thresholds or rule-based techniques to detect issues, which may result in false positives or missed detections. There is a growing need for a more intelligent, scalable, and adaptive anomaly detection approach that can self-learn the normal behavior of systems and adapt to changes over time.Deep learning models, in particular recurrent neural networks (RNNs) and long-term memory networks (LSTMs), have proven to be particularly effective in predicting time series and detecting anomalies. However, their use for system-wide real-time detection of anomalies in performance metrics still has to be completely realized. The present invention utilizes deep learning models to analyze historical and real-time performance data, thus allowing early detection of performance degradations or failures, minimizing downtime, and ensuring optimum system health.An object of the present disclosure is to enable real-time detection of performance anomalies, to reduce system downtime, and to improve reliability.Another object of the present disclosure is to adapt to evolving system behavior by continuously retraining the deep learning model with new data.Another object of the present disclosure is to provide automated, data-controlled insights into the causes of performance problems and thus to simplify the diagnosis.Another object of the present disclosure is to reduce false alarms as compared to conventional threshold based anomaly detection methods.Another object of the present disclosure is seamless integration into existing monitoring tools, where data from multiple heterogeneous sources is merged.Another object of the present disclosure is to provide a scalable solution capable of monitoring complex systems with large amounts of performance data.Another object of the present disclosure is to assist in optimizing system resources by early detection of inefficiencies and performance bottle necks.Another object of the present disclosure is to provide proactive alerts that enable a more rapid solution to problems and minimize impact on end users.Another embodiment of the present invention is the data collection module that collects system performance data such as CPU usage, memory consumption, hard disk I / O, network bandwidth, and application specific metrics. The data may be retrieved from various sources, including system protocols, monitoring tools, and network traffic analyzers.Another embodiment of the present invention is that the collected raw performance data is preprocessed to remove noise and irrelevant information. This module converts the data to a format suitable for input to the deep learning model, which may include normalization of values, aggregation of data over time slots, and extraction of relevant features such as moving averages, percentiles, or rates of change.Another embodiment of the present invention is that the deep learning module comprises a deep learning architecture, such as an LSTM or autoencoder based model, trained to learn the normal behavior of system performance metrics. The model is trained from historical data to understand patterns of normal operation. After training, the model may detect deviations in real time, such as abnormal spikes or power drops, that indicate potential system failures or inefficiencies.Another embodiment of the present invention is that the deep learning module comprises a deep learning architecture, e.g., an LSTM or autoencoder based model trained to learn the normal behavior of system performance metrics. The model is trained from historical data to understand patterns of normal operation. Once trained, the model may detect deviations in real time, such as abnormal power spikes or dips, indicative of potential system failures or inefficiencies.Another embodiment of the present invention is that the deep learning model processes incoming real-time data and compares it to learned patterns. When the model detects an anomaly, e.g., a performance metric that exceeds predefined thresholds or behaves unexpectedly, it marks the event and triggers an alarm. The alerts may be sent via email, text message, or by integration into an IT operation dashboard so that the system administrators may respond promptly.If an anomaly is detected, the decision support module helps diagnose the root cause by analyzing the affected system components. It generates reports with insights into the performance metrics and potential factors contributing to it, and helps administrators to quickly identify problems and take corrective action.Another embodiment of the present invention is that the system continuously adjusts to changes in system behavior over time by updating the deep learning model. New performance data, including feedback from corrected anomalies, is used to retraining the model, improve its accuracy, and ensure that it remains effective in evolving system configurations.The present invention relates to a deep learning based system for detecting performance anomalies that automatically learns the normal patterns of system performance data, compares running performance metrics to the learned patterns, and detects anomalies in real time. The system is capable of detecting deviations from expected performance, such as spikes in resource usage, abnormal response times or network latency issues, and generating alerts in time.The invention is explained again below with reference to the figure. The following shows: FIG. 1 shows the deep learning-based system for detecting performance anomalies ( 100).FIG. 1 illustrates the deep learning based system for detecting performance anomalies ( 100). The system consists of the following modules:A data acquisition moduleThe data acquisition module retrieves a wide range of system performance metrics. These metrics may originate from hardware sensors, system protocols, network traffic, database performance monitors, application protocols, and other monitoring tools. The data may be collected at various intervals (e.g., minute, hour, etc.), depending on system requirements.Preprocessing ModuleThe collected data is preprocessed to remove noise and outliers. To improve the quality of the input data, various methods are used for feature extraction, such as moving average smoothing, time window, and outlier detection. Feature extraction helps reduce dimensionality of the data and draw model attention to the most important variables that affect system performance.Deep Learning ModelThe heart of the invention is the deep learning model, which is responsible for learning and predicting the normal system behavior. A recurrent neural network (RNN), in particular an LSTM network, is well suited for processing sequential data such as time-series performance metrics. The model may predict future performance values based on historical trends and detect anomalies when the real-time data deviates significantly from the prediction.An alternative deep learning architecture that may be used is the autoencoder network that learns a compact representation of the data and detects anomalies when the reconstruction error exceeds a threshold. The system is flexible and supports various deep learning architectures depending on the specific performance data and system requirements.Anomaly Detection ModuleAfter the deep learning model is trained, it monitors the performance data in real time. When an anomaly is detected, the system marks the event and triggers a warning. Anomalies may include unusual spikes in CPU usage, memory consumption, application response times, or network latency. The warnings contain relevant information such as the type of anomaly, the system components involved, and the possible effects, which enables a quick remedy.Decision Support ModuleThe decision support and diagnostic module uses the results of the deep learning model to generate detailed reports of performance problems. These reports contain insight into the specific metrics causing deviations, correlations with other system events, and suggestions for the next steps. The module may also provide visualization of the performance trend to assist administrators in diagnosing the problem.Adaptive Learning ModuleThe adaptive learning and retraining module ensures that the system gets better and better over time. The model is retraining periodically from the latest performance data, including feedback from previous anomaly detections. The transhull process ensures that the model continues to respond to new patterns of system behavior while minimizing the number of false alarms.
Claims
A deep learning controlled system for detecting performance anomalies (100), comprising: a. a data acquisition module configured to acquire real-time and historical performance data from one or more system components including, but not limited to, CPU utilization, memory consumption, hard disk I / O, network traffic, and application specific metrics; b. a pre-processing module configured to process and clean the collected performance data by performing operations such as noise cancellation, normalization, feature extraction, and data transformation to a suitable format for input to a deep learning model; c. a deep learning model comprising a trained deep learning model selected from a group consisting of recurrent neural networks (RNNs), long short term memory networks (LSTMs), and auto-encoders, wherein the deep learning model is trained on historical performance data to learn the normal system performance and is configured to predict expected system performance metrics based on previous data; d. an anomaly detection module configured to compare real-time system performance data to predictions generated by the deep learning model and identify deviations from the predicted behavior and mark such deviations as performance anomalies; e. a warning module configured to trigger a warning when an anomaly is detected, the warning providing relevant context information about the anomaly, including the affected system component, the type of anomaly, and a severity indicator; f. a decision support module configured to generate a diagnostic report that includes insight into the root cause of the detected anomaly, correlations with other system metrics, and recommended remedial actions; AN adaptive learning module configured to periodically retraining the deep learning model using updated system performance data, improve the accuracy of the model over time, and adapt to changes in system behavior.The system (100) of claim 1, wherein the deep learning model is an LSTM network specifically configured to process time-series data and predict future performance values based on historical system performance patterns.The system (100) of claim 1, wherein the anomaly detection module is further configured to classify detected anomalies into different types based on their effects on system performance, such as CPU congestion, memory leaks, or network latency.The system (100) of claim 1, wherein the alerting module is configured to send notifications over multiple communication channels including email, text message, and integration with third party monitoring dashboards.The system (100) of claim 1, wherein the preprocessing module includes a feature extraction component that calculates derived features such as moving averages, percentiles, and rate of change metrics to improve the accuracy of the deep learning model.The system (100) of claim 1, wherein the decision support module further includes a visualization tool that provides graphical representations of performance trends and anomalies that have been detected over time to assist administrators in root cause analysis.The system (100) of claim 1, wherein the adaptive learning module is configured to incrementally update the model with new data whereby the system can continuously refine its anomaly detection capabilities without requiring retraining from reason to reason.The system (100) of claim 1, wherein the data collection module is configured to collect performance data from a plurality of heterogeneous sources including server protocols, network traffic analyzers, and application performance monitoring tools.
Citation Information
Cited By
Adaptive load regulation switching power supply control system and method
CN120566860A
Device pressure testing method, device, storage medium and program product
CN120892271A