Intelligent method and device for system performance prediction and anomaly detection

Through deep learning and time series analysis technology, combined with ARIMA and LSTM models, disassembly and predict system time series, the problem of difficult to deal with complex and variable system environments in the existing technology is solved, intelligent monitoring and optimization of system performance is achieved, and the accuracy and robustness of prediction are improved.

CN120066925AActive Publication Date: 2025-05-30UNIVERSAL UBIQUITOUS TECH CO LTD

Patent Information

Application Number
CN202510551850.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing system performance monitoring methods are difficult to effectively deal with complex and changeable system environments, and there are problems such as false alarms, missed alarms and difficulty in updating and maintaining rules.

Method used

The deep learning method is adopted to disassemble the system's time series into trend, seasonal components and residual components through time series analysis, and combine ARIMA and LSTM models for prediction and feature interaction to achieve intelligent monitoring and optimization of system performance.

Benefits of technology

It improves the accuracy and robustness of system performance prediction, reduces sensitivity to data noise and abnormalities, and maintains stable and reliable prediction results in a variety of operating environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066925A_ABST
    Figure CN120066925A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent method and device for system performance prediction and anomaly detection, and aims at the complexity and variability of system data, a deep learning method is adopted, so that a model can capture the complex mode and time dependence of the data more comprehensively. According to the method, the accuracy of system performance prediction is effectively improved through multi-layer analysis and multi-angle feature extraction. And through feature selection and data enhancement technologies, the model has good adaptive capacity when facing data with different scales and different properties. The robustness reduces the sensitivity of the model to data noise and abnormality, and ensures that the prediction result is still stable and reliable in various operation environments. According to the method, the capability of monitoring complex data can be improved, and intelligent monitoring and optimization of system performance are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to an intelligent method and device for system performance prediction and anomaly detection. Background Art

[0002] In modern information systems, real-time monitoring of system performance and timely detection of anomalies are crucial for maintaining system stability. Traditional monitoring methods usually rely on fixed thresholds and simple statistical methods, making it difficult to cope with complex and changing system environments. Therefore, introducing deep learning and time series analysis technologies to achieve performance prediction and anomaly detection through historical data analysis has important theoretical significance and application value.

[0003] Existing technical solutions mainly include the following three types: threshold-based monitoring systems, statistical analysis methods, and rule-based anomaly detection.

[0004] 1. Threshold-based Monitoring System Fixed-threshold monitoring is one of the most commonly used methods in traditional IT systems. This solution sets one or more fixed thresholds for each monitoring metric (such as CPU usage, memory occupancy, network latency, etc.). When any metric exceeds its threshold, the system will trigger an alarm to remind the operation and maintenance personnel to conduct inspections and handling. This method is simple to implement, easy to deploy and understand. The system alarm rules can be adjusted according to experience and historical data, meeting the monitoring requirements in some simple and stable environments, but it lacks flexibility, with serious false alarm and missed alarm phenomena and high maintenance costs.

[0005] 2. Anomaly Detection Based on Statistical Analysis This solution detects and identifies abnormal behaviors in data by applying simple statistical methods. The core technologies include mean, standard deviation analysis, time sliding window technology, etc. These methods attempt to determine whether there are abnormal situations currently through historical baselines and statistical distributions. Compared with fixed thresholds, statistical analysis can capture some subtle changes that are not easily noticed and perform quick detection through simple calculations, being applicable to certain stable metrics, but there will be problems of limited accuracy and sensitivity, difficulty in dealing with multi-dimensional data, and poor adaptability to environmental changes.

[0006] 3. Rule-based Anomaly Pattern Recognition This solution relies on experience accumulation and manually defined rule libraries to identify abnormal conditions in the system. The rules are usually formulated based on historical fault data and the experience of senior professionals and can effectively identify specific fault types. The rules are highly effective when applied to specific and known problems and can quickly identify preset abnormal patterns. For faults familiar to the operation and maintenance team, the rule method is highly efficient, but there will be problems such as difficult rule update and maintenance, limited coverage, and easy rule conflicts.

[0007] In summary, although the existing three technical solutions each have their own advantages, they have obvious deficiencies in dealing with complexity, variability, and unknown pattern recognition. This creates a strong demand for the development and introduction of new technologies to improve the robustness and intelligence level of the system. Summary of the Invention

[0008] In view of the problems in the prior art, the present application provides an intelligent method and device for system performance prediction and anomaly detection to enhance the monitoring ability of complex data and achieve intelligent monitoring and optimization of system performance.

[0009] To solve at least one of the above problems, the present application provides the following technical solutions: In a first aspect, the present application provides an intelligent method for system performance prediction and anomaly detection, the method comprising: Obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into a first ARIMA model to generate a trend component prediction value and a seasonal component prediction value; Calculate the differences between the trend component prediction value and the seasonal component prediction value and the actual values respectively to obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into an LSTM model to generate a predicted residual; Input the time series into a second ARIMA model to generate a first predicted value; add the first predicted value and the predicted residual to obtain a second predicted value; perform feature interaction on the second predicted value, the first predicted value, the initial residual, and the predicted residual, and use a regularization algorithm and an LSTM layer to optimize and obtain target interaction features; Integrate the target interaction features, the initial residual, and the predicted residual into a feature data set, and perform standardization processing on it; input the standardized feature data set into a prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold.

[0010] Further, the training method of the first ARIMA model comprises: Obtain the time series from system monitoring and generate input-output pairs using a sliding window; Use the ADF test to check the stationarity of the time series; if the stationarity of the time series is lower than a preset value, eliminate the trend component through first-order differencing; Construct an ARIMA initialization model, input the time series and model parameters p, d, q, and use the fit() method to train the model to obtain a fitting result; Analyze the model residuals and adjust the model parameters according to the model residuals until the model residuals conform to white noise.

[0011] Further, the steps of analyzing the model residuals and adjusting the model parameters according to the model residuals until the model residuals conform to white noise include: Plot the autocorrelation function (ACF) graph of the model residuals; if in the ACF graph, all autocorrelation values at different lags are within the confidence interval, then determine that the model residuals are white noise; if the autocorrelation values at a lag exceed the confidence interval, then adjust the model parameters until all autocorrelation values at different lags are within the confidence interval. It also includes normality verification. The method of normality verification is: check the degree of approximate normal distribution of the model residuals through a QQ plot; if the points of the model residuals are close to the 45-degree line of the QQ plot, it means that the model residuals conform to the normal distribution; if the deviation or bending amplitude of the points of the model residuals from the 45-degree line of the QQ plot exceeds a preset value, then adjust the model parameters until all the points of the model residuals are close to the 45-degree line of the QQ plot. It also includes randomness verification. The method of randomness verification is: use the Ljung-Box test to statistically evaluate whether the autocorrelations at multiple lags are significantly zero; if the p-value of the test is greater than the preset value, then determine that the model residuals are an independent random sequence; if the p-value of the test is less than the preset value, then adjust the model parameters until the p-value of the test is greater than the preset value.

[0012] Further, the LSTM model is built by stacking multiple layers of LSTM, and the LSTM is a bidirectional LSTM, and multiple units are set in each layer of LSTM. In the LSTM model, Optuna is used for automated hyperparameter optimization to determine the number of units; each layer of LSTM contains 64 to 256 units. In the LSTM model, a Dropout layer is added after each layer of LSTM, and the Dropout rate of the Dropout layer is set to 0.2 - 0.5. In the LSTM model, multiple fully connected layers are added after all LSTMs; the LSTM layer is used to extract the features of the initial residuals of the trend component prediction value / the initial residuals of the seasonal component prediction value, and the fully connected layer is used to construct the mapping from the features to the output of the LSTM; each layer of the fully connected layer contains 32 to 128 neurons, and the data before and after the fully connected layer are processed using batch normalization. In the LSTM model, an attention layer is added after the LSTM and before the fully connected layer; for each time step, the attention layer uses a separate neural network layer to calculate the weight scores of its outputs, and performs a dot product operation on the outputs of the LSTM and the calculated weight scores to generate the weighted outputs for each time step, and then sums these weighted outputs to form the output of the attention layer.

[0013] Furthermore, the training method of the LSTM model includes: obtaining a time series from system monitoring, performing normalization processing on the time series using Z-score, and generating input-output pairs using a sliding window; Using the mean squared error MSE as the loss function of the LSTM model, the gradient descent optimizer selects Adam, the learning rate is set to 0.001, and a learning rate adjustment strategy is used during training, and the batch size is selected from 32 to 64; The training method of the LSTM model also includes: deploying the LSTM model in a containerized environment; the method of deploying the LSTM model in a containerized environment includes: Using the tf.saved_model module of TensorFlow to save the trained LSTM model in a portable format; Encapsulating the LSTM model into an API service, integrating a monitoring tool to monitor the model performance and system health status, and setting up logging for debugging; Creating a Docker container, deploying the model file and the API service into the Docker container, and configuring multiple services through DockerCompose; the services include the API service, the monitoring tool, and the database; Using the API endpoint to request the LSTM model for prediction, and using a specified web interface to monitor the performance metrics of the API service; The Docker container uses CI / CD tools for automated deployment and update.

[0014] Furthermore, after the step of inputting the two initial residuals into the LSTM model to generate the predicted residuals, the following steps are also included: Randomly sampling from the historical data of the trend component and the seasonal component through bootstrap sampling, and using this as the input to train multiple groups of the ARIMA model and the LSTM model; Using a weighted voting system to aggregate the predicted values of multiple groups of the ARIMA model and the LSTM model to generate the predicted residuals; the weights of the weighted voting system are adjusted according to the performance evaluation results of the ARIMA model and the LSTM model.

[0015] Further, the step of performing feature interaction on the second predicted value with the first predicted value, the initial residual, and the predicted residual, and optimizing using a regularization algorithm and an LSTM layer to obtain the target interaction feature includes: Generate polynomial interaction features using Polynomial Features; Use the joint Lasso algorithm to select the first interaction features with information content exceeding a preset value from the interaction features; Further optimize the first interaction features using an LSTM layer to enable them to have temporal information and dynamic change patterns, obtaining the target interaction features; The step of determining whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold includes: Calculate the first difference between the historical target predicted value and the actual value; construct an error distribution graph based on the first difference, and set at least one percentile as the anomaly threshold; Calculate the second difference between the new target predicted value and the actual value, and compare the size of the second difference with the anomaly threshold: if the second difference is greater than the anomaly threshold, it is determined that the system performance is abnormal; if the second difference is less than the anomaly threshold, it is determined that the system performance is normal.

[0016] In a second aspect, the present application provides an intelligent device for system performance prediction and anomaly detection, including: A first prediction module, configured to obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into a first ARIMA model to generate a trend component predicted value and a seasonal component predicted value; A second prediction module, configured to calculate the differences between the trend component predicted value and the seasonal component predicted value and the actual value respectively, to obtain the initial residual of the trend component predicted value and the initial residual of the seasonal component predicted value; input the two initial residuals into an LSTM model to generate a predicted residual; A feature interaction module, configured to input the time series into a second ARIMA model to generate a first predicted value; add the first predicted value and the predicted residual to obtain a second predicted value; perform feature interaction on the second predicted value with the first predicted value, the initial residual, and the predicted residual, and optimize using a regularization algorithm and an LSTM layer to obtain the target interaction feature; A result output module, configured to integrate the target interaction feature, the initial residual, and the predicted residual into a feature data set, and perform standardization processing on it; input the standardized feature data set into a prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold.

[0017] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the intelligent method for system performance prediction and anomaly detection are implemented.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the intelligent method for system performance prediction and anomaly detection are implemented.

[0019] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the intelligent method for system performance prediction and anomaly detection are implemented.

[0020] As can be seen from the above technical solutions, the present application provides an intelligent method and device for system performance prediction and anomaly detection. In view of the complexity and variability of system data, by adopting the method of deep learning, the model can capture the complex patterns and time dependencies of data more comprehensively. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time series data with non-linear and variable characteristics.

[0021] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency, but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flowchart of the intelligent method for system performance prediction and anomaly detection in the embodiments of the present application; Figure 2It is a structural diagram of an intelligent device for system performance prediction and anomaly detection in the embodiments of the present application; Figure 3 It is a schematic structural diagram of an electronic device in the embodiments of the present application.

[0024] Reference numerals: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed implementation manners

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0026] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0027] Considering the problems existing in the prior art, the present application provides an intelligent method and device for system performance prediction and anomaly detection. In view of the complexity and variability of system data, by adopting the method of deep learning, the model can capture the complex patterns and time dependencies of data more comprehensively. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time-series data with non-linear and variable characteristics.

[0028] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency, but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0029] In order to improve the monitoring ability of complex data and achieve accurate prediction and monitoring of system performance, an embodiment of an intelligent method for system performance prediction and anomaly detection is provided in this application. Refer to Figure 1 The intelligent method for system performance prediction and anomaly detection specifically includes the following content: Step S101: Obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into a first ARIMA model to generate a predicted value of the trend component and a predicted value of the seasonal component.

[0030] In this embodiment, the time series data of the system includes logs such as CPU usage, memory, and network traffic, and may also include the disk usage percentage, disk read / write speed, and read / write request count for reflecting the disk usage situation; the number of processes and services, the memory usage of each process, the service response time, and the number of unresponsive times for reflecting processes and services; the query response time, query execution count, and query error count for reflecting database metrics; the CPU context switch count, system call count, and read / write execution rate for reflecting operating system resources; the response time, the number of processed transactions, and the number of errors in the application program or API for reflecting application performance metrics; the network latency time, the number of lost network packets, the number of TCP connections, and the number of packets entering and leaving the firewall for reflecting network security; the number of startups, shutdowns, crashes, the number of driver errors, the SMART status of the hard disk, temperature monitoring, fan speed, and the number of alerts and notifications sent through the monitoring system for reflecting system events and warning situations, etc. Exemplarily, monitoring platforms such as Datadog and New Relic can be selected to implement the monitoring of the system.

[0031] In this embodiment, the ARIMA model (Autoregressive Integrated Moving Average model) is used as the time series model to effectively process time series data and capture the dynamic change trend of the data.

[0032] Optionally, in this embodiment, STL (Seasonal-Trend decomposition using LOESS) is used to decompose the time series of the system into a trend component, a seasonal component, and a residual component. By splitting, the input dimension of the model is reduced, thereby simplifying the model structure, reducing the computing cost, and reducing the risk of overfitting. At the same time, in this embodiment, the trend component and the seasonal component are selected as the inputs of the first ARIMA model, while the residual component is discarded. By only focusing on the trend component and the seasonal component, the first ARIMA model can focus on obvious patterns and cycles, rather than trying to explain complex short-term or random fluctuations, which can improve the stability of the model and make it easier to interpret.

[0033] Optionally, the training method of the first ARIMA model includes: Data preparation: Obtain the time series from system monitoring and understand the data characteristics by plotting the time series graph; Use the sliding window technique to segment the time series data, determine the window size and sliding step length to generate input-output pairs for better extraction of local time features; Data stationarity processing: Use the ADF (Augmented Dickey-Fuller) test to check the stationarity of the time series; If the stationarity of the time series is lower than the preset value, eliminate the trend component through first-order differencing to achieve data stationarity, and process the non-stationary variance through logarithmic transformation, power transformation, etc.; Select ARIMA parameters: Use the number of differences d to determine the stationarity of the data, plot the autocorrelation function ACF and partial autocorrelation function PACF graphs, and observe the lag order to determine p and q; Model construction and training: Construct the statsmodels.tsa.arima.model.ARIMA initialization model, input the time series and model parameters p, d, q, and use the fit() method to train the model to obtain the fitting result; Model diagnosis: Analyze the model residuals and adjust the model parameters according to the model residuals until the model residuals conform to white noise; Prediction and evaluation: Use the trained ARIMA model to perform future step prediction, calculate the prediction error index, evaluate the prediction performance of the model and check its accuracy, and make necessary model adjustments according to the evaluation results.

[0034] Further optionally, the steps of analyzing the model residuals and adjusting the model parameters according to the model residuals until the model residuals conform to white noise include: Plot the autocorrelation function ACF graph of the model residuals and check whether there is significant lag autocorrelation in the residual sequence. If there is no significant autocorrelation, it means that the model has fully captured the data characteristics and the residuals are white noise. Specifically, if all the autocorrelation values of the lag periods are within the confidence interval in the autocorrelation function ACF graph, it is determined that the model residuals are white noise; if the autocorrelation values of the lag periods exceed the confidence interval, adjust the model parameters until all the autocorrelation values of the lag periods are within the confidence interval; It also includes normality verification, and the method of this normality verification is: Check the degree of approximate normal distribution of the model residuals through the QQ plot; If the points of the model residuals are close to the 45-degree line of the QQ plot, it means that the model residuals conform to the normal distribution; If the deviation or bending amplitude of the points of the model residuals from the 45-degree line of the QQ plot exceeds the preset value, adjust the model parameters until all the points of the model residuals are close to the 45-degree line of the QQ plot; It also includes randomness verification, and the method of randomness verification is as follows: Use the Ljung-Box test to statistically evaluate whether the autocorrelation at multiple lags is significantly zero; if the p-value of the test is greater than the preset value, it is determined that the model residuals are an independent random sequence; if the p-value of the test is less than the preset value, adjust the model parameters until the p-value of the test is greater than the preset value.

[0035] Optionally, the training method of the first ARIMA model further includes: Dataset division: Divide the time series data into a training set and a test set, construct the model with the training set, and evaluate the prediction performance of the model with the test set; Calculate the prediction error: Make predictions on the test dataset, calculate the mean squared error and the mean absolute error, and these metrics are used to quantify the error between the true value and the predicted value; Visualization of prediction results: Plot a comparison chart of the predicted value and the actual observed value. On the chart, the predicted value should closely follow the actual value and there should be no systematic bias.

[0036] Step S102: Calculate the differences between the predicted values of the trend component and the seasonal component and the actual values respectively, to obtain the initial residuals of the predicted value of the trend component and the initial residuals of the predicted value of the seasonal component; input the two initial residuals into the LSTM model to generate predicted residuals.

[0037] In this embodiment, when constructing a system performance prediction and anomaly detection model based on deep learning, we use LSTM (Long Short-Term Memory Network) as the core component. This model is particularly suitable for time series data and can capture long-term dependencies.

[0038] Optionally, the LSTM model is built by stacking multiple LSTMs, and the LSTM is a bidirectional LSTM. Use Bidirectional LSTM to capture the forward and backward features of the input sequence, and apply the Dropout regularization technique after each LSTM layer. The Dropout rate is controlled between 0.2 and 0.5 to prevent overfitting; use Optuna for automated hyperparameter optimization to determine the optimal number of units, with each layer containing 64 to 256 units to better extract features. For example, a multi-layer LSTM model can be built with 3 layers stacked and 128 units in each layer.

[0039] Optionally, in the LSTM model, add multiple fully connected layers after all LSTMs; the LSTM layer is used to extract the features of the initial residuals of the predicted value of the trend component / the initial residuals of the predicted value of the seasonal component, and the fully connected layer is used to construct the mapping from the features to the output of the LSTM; each fully connected layer contains 32 to 128 neurons, and batch normalization is used before and after each fully connected layer to stabilize and accelerate the training process.

[0040] Optionally, in this LSTM model, an attention layer is added after the LSTM and before the fully connected layer, so that the model can focus on important features at specific time points. For each time step, the attention layer uses a separate neural network layer to calculate the weight scores of its output, and performs a dot product operation on the output of the LSTM and the calculated weight scores to generate the weighted output for each time step. Then, these weighted outputs are summed up to form the output of the attention layer. After the model is trained, the attention weights can be analyzed to understand which input features or time points the model has focused on, and the model architecture can be further optimized accordingly.

[0041] Optionally, the training method of this LSTM model includes: obtaining time series from system monitoring, independently calculating the mean μ and standard deviation σ of each input feature using Z-score, and applying the formula z = (x - μ) / σ to each input x to eliminate the dimensional difference.

[0042] Optionally, this LSTM model uses the sliding window technique to generate input-output pairs. The size of the window determines the length of each data segment. A larger window can capture more time dependencies and long-term cycle information, while a smaller window can more sensitively reflect short-term changes. Since the performance data has obvious daily periodicity, our window size is set to 96 (one data point every 15 minutes), and the window size is set to 20 steps. The time series is sliced through the sliding window technique to obtain sufficient training samples.

[0043] Optionally, during the training process of this LSTM model, the mean squared error (MSE) is used as the loss function, and the Adam gradient descent optimizer is selected. Due to its adaptability and efficiency, the learning rate is set to 0.001, and a learning rate adjustment strategy is used during the training process. The batch size is selected from 32 to 64 to balance the training efficiency and memory consumption.

[0044] Optionally, during the training process of this LSTM model, the training dataset is traversed several times, and the learning rate is adjusted each time. For example, by gradually reducing the learning rate, the model can perform more detailed optimization when approaching convergence, improving stability and final performance. Then, the training and validation losses and other performance metrics are recorded, and the hyperparameters are adjusted according to the evaluation results.

[0045] Through the above design, this LSTM model has good generalization ability while ensuring good fitting ability, and can adapt to complex system performance prediction and anomaly detection tasks.

[0046] Optionally, the LSTM model can be deployed in a containerized environment to support dynamic updates and real-time predictions. The methods for deploying the LSTM model in a containerized environment include: using the tf.saved_model module of TensorFlow to save the trained LSTM model in a portable format; encapsulating the LSTM model into an API service, integrating a monitoring tool to monitor the model performance and system health status, and setting up logging for debugging; creating a Docker container, deploying the model file and the API service into the Docker container, and configuring multiple services through Docker Compose; the services include the API service, the monitoring tool, and the database; using the API endpoint to request the LSTM model for prediction, and using a specified web interface to monitor the performance metrics of the API service; the Docker container uses CI / CD tools for automated deployment and updates to ensure that the container can be automatically rebuilt and published after the code or model is updated.

[0047] Optionally, after the step of simultaneously inputting the two initial residuals, that is, the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value, into the LSTM model to generate a prediction residual, the following steps are further included: Randomly sampling samples from the historical data of the trend component and the seasonal component through bootstrap sampling, and using these as inputs to train multiple groups of ARIMA models and LSTM models; using a weighted voting system to aggregate the prediction values of multiple groups of ARIMA models and LSTM models to generate a prediction residual; the weights of the weighted voting system are adjusted according to the performance evaluation results of the ARIMA models and the LSTM models. This design can effectively use the bootstrap method to train multiple time series models and enhance the accuracy and robustness of the prediction by combining the prediction results of multiple models.

[0048] Step S103: Input the time series into the second ARIMA model to generate a first prediction value; add the first prediction value to the prediction residual to obtain a second prediction value; perform feature interaction on the second prediction value, the first prediction value, the initial residual, and the prediction residual, and use a regularization algorithm and an LSTM layer to optimize to obtain the target interaction feature.

[0049] Optionally, this step uses Polynomial Features to generate polynomial interaction features; uses the Joint Lasso algorithm to select the first interaction features with information content exceeding a preset value from the interaction features; further optimizes the first interaction features using an LSTM layer to enable them to have time series information and dynamic change rules, and obtains the target interaction feature.

[0050] Optionally, in the step of using Polynomial Features to generate polynomial interaction features, set the first predicted value y provided by the second ARIMA model and the predicted residual e adjusted by the LSTM model to generate an adjusted second predicted value: ŷadjusted = ŷARIMA + êLSTM. Use the adjusted second predicted value and the outputs of each model to generate these features for polynomial interaction. For simple quadratic interactions, construct features in the following form: self-squares: (ŷARIMA)², (êLSTM)², (ŷadjusted)², (ê'LSTM)² of the initial residual of the predicted value of the trend component; (ê''LSTM)² of the initial residual of the predicted value of the seasonal component, cross-terms: ŷARIMA × êLSTM, ŷARIMA × ŷadjusted, êLSTM × ŷadjusted, ŷARIMA × ê'LSTM, ŷadjusted × ê'LSTM, ŷARIMA × ê''LSTM, ŷadjusted × ê''LSTM, ê'LSTM × ê''LSTM.

[0051] Optionally, in the step of using the joint Lasso algorithm to select the first interaction features with information content exceeding a preset value from the interaction features, Lasso calculates the coefficients of each interaction feature through lasso.coef_. Interaction features with larger absolute values of the coefficients are more important in the model. We can filter out the interaction features with coefficients greater than the threshold (such as 0.1) according to the preset threshold. The strength of Lasso regularization can be controlled by the regularization parameter alpha. Larger values will increase the penalty on the feature coefficients, thus compressing the coefficients of more features to zero.

[0052] Optionally, in the step of further optimizing the first interaction feature using an LSTM layer to enable it to have temporal information and dynamic change patterns and obtain the target interaction feature, first, the first interaction feature is converted into data suitable for the input format of the LSTM layer, that is, the first interaction feature is organized into batches, and the time step and the number of features are specified in each batch to make it conform to the shape of the data required for input to the LSTM layer. Exemplarily, the first interaction feature can be reshaped into a three-dimensional array, and the time step is set to 1 to meet this requirement. Then, a simple LSTM model is constructed, which includes an LSTM layer and an output layer. The number of neurons in the LSTM layer is set to 50, and the activation function is ReLU. The output layer is a fully connected layer that outputs a continuous predicted value; finally, the mean squared error (MSE) is used as the loss function for model training, and predictions are made on the test set. The hyperparameters of the simple LSTM model are optimized according to the prediction results to obtain better model performance. In this step, the first interaction feature is learned and optimized through the LSTM network, emphasizing the influence of important information in the input interaction feature while weakening or ignoring the influence of unimportant or redundant features, thereby improving the prediction ability of the prediction model.

[0053] In this embodiment, feature interaction is used to capture the non-linear relationship between the original time series data, learn complex patterns that cannot be expressed in the original data, and a regularization technique (such as Lasso) is used to select the most informative interaction features, which can further help control the model complexity and avoid overfitting.

[0054] Step S104: Integrate the target interaction feature, the initial residual, and the predicted residual into a feature dataset, and perform normalization processing on it; input the normalized feature dataset into the prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and the preset threshold.

[0055] Optionally, in this embodiment, a gradient boosting machine is used as the prediction model. Without using time series data, predictions are made only through the target interaction feature, the initial residual, and the predicted residual generated in the previous steps, ensuring avoidance of overfitting, ensuring that the model has good generalization ability, and further improving the performance of the model using the residual and interaction features. This prediction method combines the simplicity of the model with the powerful prediction ability of the gradient boosting machine, and can effectively improve the prediction accuracy and prediction efficiency of complex time series data.

[0056] Optionally, the step of determining whether the system performance is abnormal according to the comparison result between the target prediction value and the preset threshold includes: calculating a first difference between the historical target prediction value and the actual value; constructing an error distribution graph based on the first difference, and setting at least one percentile (such as 95% and / or 99%) as the abnormality threshold; calculating a second difference between the new target prediction value and the actual value, and comparing the size of the second difference with the abnormality threshold: if the second difference is greater than the abnormality threshold, it is determined that the system performance is abnormal; if the second difference is less than the abnormality threshold, it is determined that the system performance is normal.

[0057] As can be seen from the above description, the intelligent method for system performance prediction and anomaly detection provided by the embodiments of the present application can, in view of the complexity and variability of system data, adopt a deep learning method to enable the model to more comprehensively capture the complex patterns and time dependencies of the data. Through multi-layer analysis and multi-angle feature extraction, this method effectively improves the accuracy of system performance prediction, especially when dealing with time-series data with non-linear and variable characteristics.

[0058] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency, but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0059] In order to improve the monitoring ability of complex data and achieve accurate prediction and monitoring of system performance, the present application provides an embodiment of an intelligent device for system performance prediction and anomaly detection that implements all or part of the content of the intelligent method for system performance prediction and anomaly detection, see Figure 2 The intelligent device for system performance prediction and anomaly detection specifically includes the following content: The first prediction module 10 is used to obtain the time series of the system, and decompose the time series into a trend component, a seasonal component and a residual component; input the trend component and the seasonal component into the first ARIMA model to generate a trend component prediction value and a seasonal component prediction value; The second prediction module 20 is used to calculate the differences between the trend component prediction value and the seasonal component prediction value and the actual value respectively, to obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into the LSTM model to generate a prediction residual; The feature interaction module 30 is configured to input the time series into the second ARIMA model to generate a first predicted value; add the first predicted value to the prediction residual to obtain a second predicted value; perform feature interaction on the second predicted value, the first predicted value, the initial residual, and the prediction residual, and use a regularization algorithm and an LSTM layer for optimization to obtain target interaction features; The result output module 40 is configured to integrate the target interaction features, the initial residual, and the prediction residual into a feature data set and perform normalization processing on it; input the normalized feature data set into a prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold.

[0060] As can be seen from the above description, the intelligent device for system performance prediction and anomaly detection provided by the embodiments of the present application can, in view of the complexity and variability of system data, adopt a deep learning method to enable the model to more comprehensively capture the complex patterns and time dependencies of the data. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time series data with non-linear and variable characteristics.

[0061] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0062] From the hardware level, in order to improve the monitoring ability of complex data and achieve accurate prediction and monitoring of system performance, the present application provides an embodiment of an electronic device for implementing all or part of the content in the intelligent method for system performance prediction and anomaly detection. The electronic device specifically includes the following contents: A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between an intelligent device for system performance prediction and anomaly detection and related devices such as a core business system, a user terminal, and a related database, etc.; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the intelligent method for system performance prediction and anomaly detection, and the embodiments of the intelligent device for system performance prediction and anomaly detection, the content of which is incorporated herein, and the repeated parts will not be elaborated again.

[0063] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0064] In practical applications, part of the intelligent method for system performance prediction and anomaly detection can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation thereto. If all operations are completed in the client device, the client device may further include a processor.

[0065] The above-mentioned client device may have a communication module (i.e., a communication unit), and can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and may also include a server on an intermediate platform in other implementation scenarios, such as a server on a third-party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.

[0066] Figure 3 This is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 3 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 3 is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0067] In one embodiment, the function of the intelligent method for system performance prediction and anomaly detection can be integrated into the central processing unit 9100. Among them, the central processing unit 9100 can be configured to perform the following controls: Step S101: Obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into the first ARIMA model to generate a trend component prediction value and a seasonal component prediction value; Step S102: Calculate the differences between the trend component prediction value and the seasonal component prediction value and the actual values respectively to obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into the LSTM model to generate a prediction residual; Step S103: Input the time series into the second ARIMA model to generate a first prediction value; add the first prediction value and the prediction residual to obtain a second prediction value; perform feature interaction on the second prediction value, the first prediction value, the initial residual, and the prediction residual, and use a regularization algorithm and an LSTM layer to optimize and obtain target interaction features; Step S104: Integrate the target interaction features, the initial residual, and the prediction residual into a feature data set, and perform standardization processing on it; input the standardized feature data set into the prediction model to generate a target prediction value; determine whether the system performance is abnormal according to the comparison result between the target prediction value and a preset threshold.

[0068] As can be seen from the above description, for the complexity and variability of system data, the electronic device provided by the embodiments of the present application enables the model to more comprehensively capture the complex patterns and time dependencies of the data by adopting a deep learning method. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time series data with non-linear and variable characteristics.

[0069] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency, but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0070] In another embodiment, the intelligent device for system performance prediction and anomaly detection can be separately configured from the central processing unit 9100. For example, the intelligent device for system performance prediction and anomaly detection can be configured as a chip connected to the central processing unit 9100, and the intelligent method functions of system performance prediction and anomaly detection are realized through the control of the central processing unit.

[0071] As Figure 3 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 3 all the components shown in Figure 3 ; in addition, the electronic device 9600 may further include

[0072] As Figure 3 shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.

[0073] Among them, the memory 9140 can be, for example, one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above-mentioned information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the programs stored in the memory 9140 to implement information storage or processing, etc.

[0074] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display can be, for example, an LCD display, but is not limited thereto.

[0075] The memory 9140 can be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It can also be such a memory that stores information even when powered off, can be selectively erased and has more data. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.

[0076] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device such as a messaging application, an address book application, etc.

[0077] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide an input signal and receive an output signal, which may be the same as in the case of a conventional mobile communication terminal.

[0078] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0079] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the intelligent method for system performance prediction and anomaly detection in the above embodiments where the execution subject is a server or a client. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements all steps of the intelligent method for system performance prediction and anomaly detection in the above embodiments where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented: Step S101: Obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into a first ARIMA model to generate a trend component prediction value and a seasonal component prediction value; Step S102: Calculate the differences between the trend component prediction value and the seasonal component prediction value and the actual values respectively to obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into an LSTM model to generate a prediction residual; Step S103: Input the time series into the second ARIMA model to generate a first predicted value; add the first predicted value to the prediction residual to obtain a second predicted value; perform feature interaction on the second predicted value, the first predicted value, the initial residual, and the prediction residual, and use a regularization algorithm and an LSTM layer for optimization to obtain target interaction features; Step S104: Integrate the target interaction features, the initial residual, and the prediction residual into a feature dataset, and perform standardization processing on it; input the standardized feature dataset into a prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold.

[0080] As can be seen from the above description, for the complexity and variability of system data, the computer-readable storage medium provided by the embodiments of the present application enables the model to more comprehensively capture the complex patterns and time dependencies of the data by adopting a deep learning method. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time series data with non-linear and variable characteristics.

[0081] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0082] The embodiments of the present application also provide a computer program product that can implement all the steps of the intelligent method for system performance prediction and anomaly detection with the execution subject being a server or a client in the above embodiments. When the computer program / instructions are executed by a processor, the steps of the intelligent method for system performance prediction and anomaly detection are implemented. For example, the computer program / instructions implement the following steps: Step S101: Obtain the time series of the system, and decompose the time series into a trend component, a seasonal component, and a residual component; input the trend component and the seasonal component into the first ARIMA model to generate a trend component predicted value and a seasonal component predicted value; Step S102: Calculate the differences between the predicted values of the trend component and the seasonal component and the actual values respectively to obtain the initial residuals of the trend component predicted value and the initial residuals of the seasonal component predicted value; input the two initial residuals into the LSTM model to generate predicted residuals; Step S103: Input the time series into the second ARIMA model to generate a first predicted value; add the first predicted value and the predicted residuals to obtain a second predicted value; perform feature interaction on the second predicted value, the first predicted value, the initial residuals, and the predicted residuals, and use a regularization algorithm and an LSTM layer for optimization to obtain target interaction features; Step S104: Integrate the target interaction features, the initial residuals, and the predicted residuals into a feature dataset and perform standardization processing on it; input the standardized feature dataset into the prediction model to generate a target predicted value; determine whether the system performance is abnormal according to the comparison result between the target predicted value and a preset threshold.

[0083] As can be seen from the above description, for the complexity and variability of system data, the computer program product provided by the embodiments of the present application enables the model to capture the complex patterns and time dependencies of the data more comprehensively by adopting a deep learning method. This method effectively improves the accuracy of system performance prediction through multi-layer analysis and multi-angle feature extraction, especially when dealing with time series data with non-linear and variable characteristics.

[0084] Moreover, through feature selection and data augmentation techniques, the model has good adaptability when facing data of different scales and natures. This robustness reduces the sensitivity of the model to data noise and anomalies, ensuring that the prediction results are still stable and reliable in diverse operating environments. At the same time, through reasonable parameter settings and model structure design, the solution not only optimizes the execution efficiency but also shows good scalability when dealing with real-time data and large-scale multi-dimensional data. This enables the solution to be quickly deployed and play a role in actual business, supporting the continuous optimization and upgrade of the system.

[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can adopt the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.

[0089] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An intelligent method for system performance prediction and anomaly detection, characterized in that: The method comprises: Obtaining a time series of the system, and decomposing the time series into a trend component, a seasonal component, and a residual component; inputting the trend component and the seasonal component into a first ARIMA model to generate a trend component forecast value and a seasonal component forecast value; Calculate the difference between the trend component prediction value and the seasonal component prediction value and the actual value respectively, and obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into the LSTM model to generate the prediction residual; Input the time series into the second ARIMA model to generate a first prediction value; add the first prediction value to the prediction residual to obtain a second prediction value; perform feature interaction on the second prediction value, the first prediction value, the initial residual, and the prediction residual, and use a regularization algorithm and LSTM layer optimization to obtain a target interaction feature; The target interaction feature, the initial residual and the predicted residual are integrated into a feature data set, and the data set is standardized; the standardized feature data set is input into the prediction model to generate a target prediction value; and whether the system performance is abnormal is determined based on a comparison result between the target prediction value and a preset threshold.

2. The intelligent method for system performance prediction and anomaly detection according to claim 1, characterized in that: The training method of the first ARIMA model includes: Get time series from system monitoring and use sliding windows to generate input-output pairs; Use the ADF test to check the stationarity of the time series; if the stationarity of the time series is lower than the preset value, eliminate the trend component through the first-order difference; Build the ARIMA initialization model, pass in the time series and model parameters p, d, q, use the fit() method to train the model, and obtain the fitting results; The model residuals are analyzed, and the model parameters are adjusted according to the model residuals until the model residuals conform to white noise.

3. The intelligent method for system performance prediction and anomaly detection according to claim 2, characterized in that: The step of analyzing the model residual and adjusting the model parameters according to the model residual until the model residual conforms to white noise includes: Draw an autocorrelation function ACF graph of the model residual; if in the autocorrelation function ACF graph, all lag autocorrelation values ​​are within the confidence interval, then the model residual is determined to be white noise; if the lag autocorrelation value exceeds the confidence interval, then adjust the model parameters until all lag autocorrelation values ​​are within the confidence interval; It also includes normality verification, and the method of normality verification is: checking the degree to which the model residual approximates the normal distribution through the QQ graph; if the point of the model residual is close to the 45-degree line of the QQ graph, it means that the model residual conforms to the normal distribution; if the deviation or bending amplitude of the point of the model residual compared to the 45-degree line of the QQ graph exceeds a preset value, the model parameters are adjusted until all the points of the model residual are close to the 45-degree line of the QQ graph; It also includes randomness verification, and the method of randomness verification is: using the Ljung-Box test to statistically evaluate whether the autocorrelation under multiple lags is significantly zero; if the p-value of the test is greater than the preset value, it is determined that the model residuals are independent random sequences; if the p-value of the test is less than the preset value, the model parameters are adjusted until the p-value of the test is greater than the preset value.

4. The intelligent method for system performance prediction and anomaly detection according to claim 1, characterized in that: The LSTM model is constructed by stacking multiple layers of LSTM, and the LSTM is a bidirectional LSTM, and each layer of LSTM is provided with multiple units; In the LSTM model, Optuna was used for automated hyperparameter optimization to determine the number of units; each LSTM layer contained 64 to 256 units; In the LSTM model, a Dropout layer is added after each LSTM layer, and the Dropout rate of the Dropout layer is set to 0.2-0.5; In the LSTM model, multiple fully connected layers are added after all LSTMs; the LSTM layer is used to extract the features of the initial residual of the trend component prediction value / the initial residual of the seasonal component prediction value, and the fully connected layer is used to construct a mapping from the features to the output of the LSTM; each of the fully connected layers contains 32 to 128 neurons, and the data before and after the fully connected layer are processed using batch normalization; In the LSTM model, an attention layer is added after the LSTM and before the fully connected layer; for each time step, the attention layer uses a separate neural network layer to calculate the weight score of its output, and performs a dot multiplication operation on the output of the LSTM and the calculated weight score to generate a weighted output for each time step, and then these weighted outputs are added together to form the output of the attention layer.

5. The intelligent method for system performance prediction and anomaly detection according to claim 1, characterized in that: The training method of the LSTM model includes: obtaining a time series from system monitoring, standardizing the time series using a Z-score, and generating input-output pairs using a sliding window; The mean square error (MSE) is used as the loss function of the LSTM model, the gradient descent optimizer is Adam, the learning rate is set to 0.001, and the learning rate adjustment strategy is used during training. The batch size is selected from 32 to 64. The training method of the LSTM model also includes: deploying the LSTM model in a containerized environment; the method of deploying the LSTM model in a containerized environment includes: Use TensorFlow's tf.saved_model module to save the trained LSTM model into a portable format; Encapsulate the LSTM model into an API service, integrate monitoring tools to monitor model performance and system health, and set up logging for debugging; Create a Docker container, deploy the model file and API service into the Docker container, and configure multiple services through DockerCompose; the services include API service, monitoring tools and database; Use the API endpoint to request the LSTM model to make predictions and use the specified web interface to monitor the performance indicators of the API service; The Docker container uses CI / CD tools for automated deployment and updates.

6. The intelligent method for system performance prediction and anomaly detection according to claim 1, characterized in that: After the step of inputting the two initial residuals into the LSTM model to generate the prediction residuals, the following step is further included: Randomly extract samples from the historical data of the trend component and the seasonal component through self-service sampling, and use the samples as input to train multiple groups of the ARIMA model and the LSTM model; A weighted voting system is used to aggregate multiple groups of predicted values ​​of the ARIMA model and the LSTM model to generate the predicted residual; the weight of the weighted voting system is adjusted according to the performance evaluation results of the ARIMA model and the LSTM model.

7. The intelligent method for system performance prediction and anomaly detection according to claim 1, characterized in that: The step of performing feature interaction between the second predicted value, the first predicted value, the initial residual, and the predicted residual, and obtaining target interaction features by using a regularization algorithm and an LSTM layer optimization comprises: Use Polynomial Features to generate polynomial interaction features; Selecting a first interactive feature whose information amount exceeds a preset value from the interactive features using a joint Lasso algorithm; The first interaction feature is further optimized using an LSTM layer so that it has time series information and dynamic change rules, thereby obtaining the target interaction feature; The step of determining whether the system performance is abnormal according to the comparison result between the target prediction value and the preset threshold value comprises: Calculate a first difference between a historical target predicted value and an actual value; construct an error distribution graph based on the first difference, and set at least one percentile as an abnormal threshold; Calculate the second difference between the new target prediction value and the actual value, and compare the second difference with the abnormal threshold: if the second difference is greater than the abnormal threshold, determine that the system performance is abnormal; if the second difference is less than the abnormal threshold, determine that the system performance is normal.

8. An intelligent device for system performance prediction and anomaly detection, characterized in that: The device comprises: The first prediction module is used to obtain the time series of the system and decompose the time series into a trend component, a seasonal component and a residual component; the trend component and the seasonal component are input into the first ARIMA model to generate a trend component prediction value and a seasonal component prediction value; The second prediction module is used to calculate the difference between the trend component prediction value and the seasonal component prediction value and the actual value, respectively, to obtain the initial residual of the trend component prediction value and the initial residual of the seasonal component prediction value; input the two initial residuals into the LSTM model to generate prediction residuals; A feature interaction module is used to input the time series into the second ARIMA model to generate a first prediction value; add the first prediction value to the prediction residual to obtain a second prediction value; perform feature interaction on the second prediction value with the first prediction value, the initial residual, and the prediction residual, and use a regularization algorithm and LSTM layer optimization to obtain a target interaction feature; The result output module is used to integrate the target interaction feature, the initial residual and the predicted residual into a feature data set and perform standardization on it; input the standardized feature data set into the prediction model to generate a target prediction value; and determine whether the system performance is abnormal based on the comparison result between the target prediction value and a preset threshold.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the intelligent method for system performance prediction and anomaly detection described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent method for system performance prediction and anomaly detection described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Hybrid cloud scene-oriented time series data anomaly prediction method based on ensemble learning technology

    CN112131212A

  • Air quality prediction method based on seasonal recurrent neural network

    CN113240170A

  • AI-based distribution network 10KV line technology loss reduction intelligent adjustment method and device

    CN117913987A

  • Short-term load prediction method and system based on ARIMA and CNN-LSTM combined model

    CN118644096A

  • Data prediction method and device, electronic equipment and storage medium

    CN119226948A

Cited By

  • Energy consumption anomaly detection method and device based on divide-and-conquer fusion architecture

    CN120822165A

  • Automobile manufacturing parameter optimization method and device

    CN121188894A